1,451
Overall · rank 18 of 44
- Unit
- Arena rating, higher is better
- Range
- 1,442 to 1,460
- Sample
- 10769 votes
- Configuration
- gemini-3-pro
- Measured
- 13 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,201
Overall · rank 5 of 34
- Unit
- Arena rating, higher is better
- Range
- 1,196 to 1,207
- Sample
- 37024 votes
- Configuration
- gemini-3-pro
- Measured
- 24 Aug 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,481
Overall · rank 4 of 177
- Unit
- Arena rating, higher is better
- Range
- 1,478 to 1,485
- Sample
- 40987 votes
- Configuration
- gemini-3-pro
- Measured
- 25 Sep 2026
- Not shown
- Not an error rate: claims that cannot be checked on the web are skipped, and preference still carries most of the weight.
1,474
Business, management and finance · rank 11 of 402
- Unit
- Arena rating, higher is better
- Range
- 1,467 to 1,481
- Sample
- 7727 votes
- Configuration
- gemini-3-pro
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,484
Creative writing · rank 3 of 407
- Unit
- Arena rating, higher is better
- Range
- 1,476 to 1,492
- Sample
- 6510 votes
- Configuration
- gemini-3-pro
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,498
Expert prompts · rank 14 of 359
- Unit
- Arena rating, higher is better
- Range
- 1,487 to 1,510
- Sample
- 2814 votes
- Configuration
- gemini-3-pro
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,473
Instruction following · rank 15 of 409
- Unit
- Arena rating, higher is better
- Range
- 1,466 to 1,479
- Sample
- 11653 votes
- Configuration
- gemini-3-pro
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,485
Overall · rank 8 of 409
- Unit
- Arena rating, higher is better
- Range
- 1,482 to 1,489
- Sample
- 41921 votes
- Configuration
- gemini-3-pro
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,481
Writing, literature and language · rank 5 of 408
- Unit
- Arena rating, higher is better
- Range
- 1,474 to 1,488
- Sample
- 9527 votes
- Configuration
- gemini-3-pro
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
49.0%
Average · rank 19 of 34
- Unit
- % resolved, higher is better
- Configuration
- gemini-3-pro
- Measured
- 30 Jun 2026
- Not shown
- Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
36.8%
Front end (SWE-Bench Multimodal) · rank 11 of 34
- Unit
- % resolved, higher is better
- Configuration
- gemini-3-pro
- Measured
- 30 Jun 2026
- Not shown
- Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
25.0%
Greenfield (Commit0) · rank 14 of 34
- Unit
- % resolved, higher is better
- Configuration
- gemini-3-pro
- Measured
- 30 Jun 2026
- Not shown
- Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
44.2%
Information gathering (GAIA) · rank 25 of 34
- Unit
- % resolved, higher is better
- Configuration
- gemini-3-pro
- Measured
- 30 Jun 2026
- Not shown
- Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
70.6%
Issue resolution (SWE-Bench) · rank 25 of 34
- Unit
- % resolved, higher is better
- Configuration
- gemini-3-pro
- Measured
- 30 Jun 2026
- Not shown
- Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
68.6%
Testing (SWT-Bench) · rank 17 of 34
- Unit
- % resolved, higher is better
- Configuration
- gemini-3-pro
- Measured
- 30 Jun 2026
- Not shown
- Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
66.0%
Consistency (pass^4) · rank 6 of 8
- Unit
- % of tasks, higher is better
- Configuration
- gemini-3-pro
- Measured
- 2 Mar 2026
- Not shown
- Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.
80.5%
Task success (pass^1) · rank 6 of 8
- Unit
- % of tasks, higher is better
- Configuration
- gemini-3-pro
- Measured
- 2 Mar 2026
- Not shown
- Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.
4.1%
Consistency (pass^4) · rank 3 of 3
- Unit
- % of tasks, higher is better
- Configuration
- gemini-3-pro
- Measured
- 2 Mar 2026
- Not shown
- Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.
18.0%
Task success (pass^1) · rank 3 of 3
- Unit
- % of tasks, higher is better
- Configuration
- gemini-3-pro
- Measured
- 2 Mar 2026
- Not shown
- Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.
47.4%
Consistency (pass^4) · rank 5 of 8
- Unit
- % of tasks, higher is better
- Configuration
- gemini-3-pro
- Measured
- 2 Mar 2026
- Not shown
- Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.
75.9%
Task success (pass^1) · rank 5 of 8
- Unit
- % of tasks, higher is better
- Configuration
- gemini-3-pro
- Measured
- 2 Mar 2026
- Not shown
- Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.
74.6%
Consistency (pass^4) · rank 3 of 8
- Unit
- % of tasks, higher is better
- Configuration
- gemini-3-pro
- Measured
- 2 Mar 2026
- Not shown
- Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.
91.0%
Task success (pass^1) · rank 4 of 8
- Unit
- % of tasks, higher is better
- Configuration
- gemini-3-pro
- Measured
- 2 Mar 2026
- Not shown
- Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.