Benchmarks / OpenHands Index

Reported by OpenHands Index

OpenHands Index

The equally weighted average of the five category scores below.

Results dated
30 Jun 2026
Models
34
Unit
% resolved
Licence
Apache License 2.0
OpenHands Index: average, % resolved, higher is better
#ModelOpenHands Index: average
% resolved, higher is better
1 Claude Fable 5Anthropic · claude-fable-5
81.0%
2 Claude Opus 4.8Anthropic · claude-opus-4.8
71.9%
3 Claude Opus 4.7Anthropic · claude-opus-4.7
69.7%
4 Claude Opus 4.6Anthropic · claude-opus-4.6
66.7%
5 GPT-5.5OpenAI · gpt-5.5
65.9%
6 GPT-5.4OpenAI · gpt-5.4
64.3%
7 Gemini 3.5 FlashGoogle · gemini-3.5-flash
62.6%
8 Claude Opus 4.5Anthropic · claude-opus-4.5
60.6%
9 Gemini 3.1 Pro PreviewGoogle · gemini-3.1-pro-preview
60.6%
10 GPT-5.2OpenAI · gpt-5.2
58.8%
11 GPT-5.2-CodexOpenAI · gpt-5.2-codex
58.3%
12 GLM 5.1Z.ai · glm-5.1
58.2%
13 MiniMax M3MiniMax · minimax-m3
57.2%
14 Kimi K2.6Moonshot AI · kimi-k2.6
57.1%
15 Claude Sonnet 4.5Anthropic · claude-sonnet-4.5
53.0%
16 Qwen3.6 PlusAlibaba · qwen3.6-plus
52.9%
17 GLM 5Z.ai · glm-5
49.4%
18 Kimi K2.5Moonshot AI · kimi-k2.5
49.2%
19 gemini-3-proGoogle
49.0%
20 gemini-3-flashGoogle
49.0%
21 DeepSeek V3.2DeepSeek · deepseek-v3.2
45.7%
22 MiniMax M2.5MiniMax · minimax-m2.5
45.2%
23 Claude Sonnet 4.6Anthropic · claude-sonnet-4.6
44.5%
24 Qwen3 Coder NextAlibaba · qwen3-coder-next
43.8%
25 MiniMax M2.7MiniMax · minimax-m2.7
43.4%
26 GLM 4.7Z.ai · glm-4.7
42.3%
27 MiniMax M2.1MiniMax · minimax-m2.1
41.2%
28 Kimi K2 ThinkingMoonshot AI · kimi-k2-thinking
41.0%
29 DeepSeek V4 Pro 0423DeepSeek · deepseek-v4-pro
40.7%
30 Qwen3.5-FlashAlibaba · qwen3.5-flash-02-23
38.1%
31 Nemotron 3 SuperNVIDIA · nemotron-3-super-120b-a12b
36.2%
32 Trinity Large ThinkingArcee Ai · trinity-large-thinking
32.1%
33 Qwen3 Coder 480B A35BAlibaba · qwen3-coder
30.9%
34 Nemotron 3 Nano 30B A3BNVIDIA · nemotron-3-nano-30b-a3b
15.5%

Results as published by OpenHands Index; we do not re-run them.

What it measures

The equally weighted average of the five category scores below.

What it does not measure

Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.

OpenHands Index by the OpenHands contributors, Apache License 2.0.