-0.14
Confirmed task success · rank 39 of 46
- Unit
- IPS effect estimate, higher is better
- Range
- -0.17 to -0.12
- Sample
- 27787 observations
- Configuration
- MiniMax M3
- Measured
- 28 Sep 2026
- Not shown
- Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
-0.13
Praise over complaint · rank 35 of 46
- Unit
- IPS effect estimate, higher is better
- Range
- -0.16 to -0.11
- Sample
- 11282 observations
- Configuration
- MiniMax M3
- Measured
- 28 Sep 2026
- Not shown
- Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
-0.07
Steerability · rank 35 of 46
- Unit
- IPS effect estimate, higher is better
- Range
- -0.09 to -0.05
- Sample
- 38615 observations
- Configuration
- MiniMax M3
- Measured
- 28 Sep 2026
- Not shown
- Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
-0.00
Tool grounding · rank 29 of 46
- Unit
- IPS effect estimate, higher is better
- Range
- -0.00 to 0.00
- Sample
- 2922906 observations
- Configuration
- MiniMax M3
- Measured
- 28 Sep 2026
- Not shown
- Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
1,433
Overall · rank 34 of 44
- Unit
- Arena rating, higher is better
- Range
- 1,425 to 1,441
- Sample
- 6314 votes
- Configuration
- MiniMax M3
- Measured
- 13 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,434
Overall · rank 84 of 177
- Unit
- Arena rating, higher is better
- Range
- 1,431 to 1,437
- Sample
- 56345 votes
- Configuration
- MiniMax M3
- Measured
- 25 Sep 2026
- Not shown
- Not an error rate: claims that cannot be checked on the web are skipped, and preference still carries most of the weight.
1,450
Business, management and finance · rank 47 of 402
- Unit
- Arena rating, higher is better
- Range
- 1,443 to 1,456
- Sample
- 11108 votes
- Configuration
- MiniMax M3
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,406
Creative writing · rank 76 of 407
- Unit
- Arena rating, higher is better
- Range
- 1,399 to 1,413
- Sample
- 10169 votes
- Configuration
- MiniMax M3
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,471
Expert prompts · rank 51 of 359
- Unit
- Arena rating, higher is better
- Range
- 1,462 to 1,479
- Sample
- 6168 votes
- Configuration
- MiniMax M3
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,435
Instruction following · rank 63 of 409
- Unit
- Arena rating, higher is better
- Range
- 1,429 to 1,440
- Sample
- 19768 votes
- Configuration
- MiniMax M3
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,440
Overall · rank 77 of 409
- Unit
- Arena rating, higher is better
- Range
- 1,436 to 1,444
- Sample
- 56646 votes
- Configuration
- MiniMax M3
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,422
Writing, literature and language · rank 70 of 408
- Unit
- Arena rating, higher is better
- Range
- 1,416 to 1,428
- Sample
- 14261 votes
- Configuration
- MiniMax M3
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
57.2%
Average · rank 13 of 34
- Unit
- % resolved, higher is better
- Configuration
- minimax-m3
- Measured
- 30 Jun 2026
- Not shown
- Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
36.8%
Front end (SWE-Bench Multimodal) · rank 11 of 34
- Unit
- % resolved, higher is better
- Configuration
- minimax-m3
- Measured
- 30 Jun 2026
- Not shown
- Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
25.0%
Greenfield (Commit0) · rank 14 of 34
- Unit
- % resolved, higher is better
- Configuration
- minimax-m3
- Measured
- 30 Jun 2026
- Not shown
- Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
66.7%
Information gathering (GAIA) · rank 15 of 34
- Unit
- % resolved, higher is better
- Configuration
- minimax-m3
- Measured
- 30 Jun 2026
- Not shown
- Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
76.4%
Issue resolution (SWE-Bench) · rank 9 of 34
- Unit
- % resolved, higher is better
- Configuration
- minimax-m3
- Measured
- 30 Jun 2026
- Not shown
- Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
81.1%
Testing (SWT-Bench) · rank 4 of 34
- Unit
- % resolved, higher is better
- Configuration
- minimax-m3
- Measured
- 30 Jun 2026
- Not shown
- Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.