Benchmarks / OpenHands Index
Reported by OpenHands Index
OpenHands Index
Building a library from its specification until its test suite passes.
- Results dated
- 30 Jun 2026
- Models
- 34
- Unit
- % resolved
- Licence
- Apache License 2.0
Full results
| # | Model | OpenHands Index: greenfield (Commit0) % resolved, higher is better |
|---|---|---|
| 1 | Claude Fable 5 (openhands)Anthropic |
62.5%
|
| 1 | Claude Opus 4.8 (openhands)Anthropic |
62.5%
|
| 3 | Claude Opus 4.6 (openhands)Anthropic |
56.2%
|
| 3 | Claude Opus 4.7 (openhands)Anthropic |
56.2%
|
| 3 | GPT-5.4 (openhands)OpenAI |
56.2%
|
| 6 | Claude Sonnet 4.6 (openhands)Anthropic |
50.0%
|
| 6 | GPT-5.2 (openhands)OpenAI |
50.0%
|
| 8 | GPT-5.2-Codex (openhands)OpenAI |
43.8%
|
| 8 | GPT-5.5 (openhands)OpenAI |
43.8%
|
| 10 | Claude Opus 4.5 (openhands)Anthropic |
37.5%
|
| 10 | Gemini 3.5 Flash (openhands)Google |
37.5%
|
| 10 | GLM 5.1 (openhands)Z.ai |
37.5%
|
| 13 | GLM 5 (openhands)Z.ai |
31.2%
|
| 14 | Qwen3.6 Plus (openhands)Alibaba |
25.0%
|
| 14 | Qwen3 Coder Next (openhands)Alibaba |
25.0%
|
| 14 | DeepSeek V3.2 (thinking reasoning, openhands)DeepSeek |
25.0%
|
| 14 | Gemini 3.1 Pro Preview (openhands)Google |
25.0%
|
| 14 | gemini-3-pro (openhands)Google |
25.0%
|
| 14 | MiniMax M3 (openhands)MiniMax |
25.0%
|
| 14 | Kimi K2.6 (openhands)Moonshot AI |
25.0%
|
| 21 | gemini-3-flash (openhands)Google |
18.8%
|
| 21 | MiniMax M2.1 (openhands)MiniMax |
18.8%
|
| 21 | MiniMax M2.7 (openhands)MiniMax |
18.8%
|
| 21 | Kimi K2.5 (openhands)Moonshot AI |
18.8%
|
| 25 | Qwen3.5-Flash (openhands)Alibaba |
12.5%
|
| 25 | Claude Sonnet 4.5 (openhands)Anthropic |
12.5%
|
| 25 | Trinity Large Thinking (openhands)Arcee Ai |
12.5%
|
| 25 | DeepSeek V4 Pro 0423 (openhands)DeepSeek |
12.5%
|
| 25 | MiniMax M2.5 (openhands)MiniMax |
12.5%
|
| 25 | Kimi K2 Thinking (openhands)Moonshot AI |
12.5%
|
| 25 | Nemotron 3 Super (openhands)NVIDIA |
12.5%
|
| 25 | GLM 4.7 (openhands)Z.ai |
12.5%
|
| 33 | Nemotron 3 Nano 30B A3B (openhands)NVIDIA |
6.2%
|
| 34 | Qwen3 Coder 480B A35B (openhands)Alibaba |
0.0%
|
Results as published by OpenHands Index; we do not re-run them.
What it measures
Building a library from its specification until its test suite passes.
What it does not measure
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Source
OpenHands Index by the OpenHands contributors, Apache License 2.0.