-0.22
Confirmed task success · rank 43 of 46
- Unit
- IPS effect estimate, higher is better
- Range
- -0.25 to -0.19
- Sample
- 24482 observations
- Configuration
- Inkling
- Measured
- 28 Sep 2026
- Not shown
- Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
-0.22
Praise over complaint · rank 44 of 46
- Unit
- IPS effect estimate, higher is better
- Range
- -0.25 to -0.20
- Sample
- 9597 observations
- Configuration
- Inkling
- Measured
- 28 Sep 2026
- Not shown
- Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
-0.12
Steerability · rank 42 of 46
- Unit
- IPS effect estimate, higher is better
- Range
- -0.14 to -0.10
- Sample
- 31566 observations
- Configuration
- Inkling
- Measured
- 28 Sep 2026
- Not shown
- Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
-0.00
Tool grounding · rank 29 of 46
- Unit
- IPS effect estimate, higher is better
- Range
- -0.00 to 0.00
- Sample
- 1538933 observations
- Configuration
- Inkling
- Measured
- 28 Sep 2026
- Not shown
- Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
1,452
Overall · rank 46 of 177
- Unit
- Arena rating, higher is better
- Range
- 1,448 to 1,455
- Sample
- 32469 votes
- Configuration
- Inkling
- Measured
- 25 Sep 2026
- Not shown
- Not an error rate: claims that cannot be checked on the web are skipped, and preference still carries most of the weight.
1,444
Business, management and finance · rank 59 of 402
- Unit
- Arena rating, higher is better
- Range
- 1,436 to 1,452
- Sample
- 6290 votes
- Configuration
- Inkling
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,384
Creative writing · rank 99 of 407
- Unit
- Arena rating, higher is better
- Range
- 1,376 to 1,392
- Sample
- 6718 votes
- Configuration
- Inkling
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,474
Expert prompts · rank 42 of 359
- Unit
- Arena rating, higher is better
- Range
- 1,464 to 1,484
- Sample
- 3981 votes
- Configuration
- Inkling
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,427
Instruction following · rank 76 of 409
- Unit
- Arena rating, higher is better
- Range
- 1,420 to 1,433
- Sample
- 11891 votes
- Configuration
- Inkling
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,442
Overall · rank 73 of 409
- Unit
- Arena rating, higher is better
- Range
- 1,438 to 1,447
- Sample
- 32778 votes
- Configuration
- Inkling
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,403
Writing, literature and language · rank 91 of 408
- Unit
- Arena rating, higher is better
- Range
- 1,396 to 1,411
- Sample
- 8708 votes
- Configuration
- Inkling
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
11.3%
Consistency (pass^4) · rank 14 of 21
- Unit
- % of tasks, higher is better
- Configuration
- inkling
- Measured
- 24 Jul 2026
- Not shown
- Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.
25.0%
Task success (pass^1) · rank 16 of 21
- Unit
- % of tasks, higher is better
- Configuration
- inkling
- Measured
- 24 Jul 2026
- Not shown
- Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.