-0.09
Confirmed task success · rank 34 of 46
- Unit
- IPS effect estimate, higher is better
- Range
- -0.11 to -0.07
- Sample
- 29712 observations
- Configuration
- Qwen3.7 Max
- Measured
- 28 Sep 2026
- Not shown
- Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
-0.08
Praise over complaint · rank 32 of 46
- Unit
- IPS effect estimate, higher is better
- Range
- -0.11 to -0.06
- Sample
- 12178 observations
- Configuration
- Qwen3.7 Max
- Measured
- 28 Sep 2026
- Not shown
- Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
-0.03
Steerability · rank 24 of 46
- Unit
- IPS effect estimate, higher is better
- Range
- -0.05 to -0.02
- Sample
- 46375 observations
- Configuration
- Qwen3.7 Max
- Measured
- 28 Sep 2026
- Not shown
- Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
-0.00
Tool grounding · rank 30 of 46
- Unit
- IPS effect estimate, higher is better
- Range
- -0.00 to -0.00
- Sample
- 1821444 observations
- Configuration
- Qwen3.7 Max
- Measured
- 28 Sep 2026
- Not shown
- Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.