58.3%
Average · rank 11 of 34
- Unit
- % resolved, higher is better
- Configuration
- GPT-5.2-Codex (openhands)
- Measured
- 30 Jun 2026
- Not shown
- Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
35.9%
Front end (SWE-Bench Multimodal) · rank 16 of 34
- Unit
- % resolved, higher is better
- Configuration
- GPT-5.2-Codex (openhands)
- Measured
- 30 Jun 2026
- Not shown
- Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
43.8%
Greenfield (Commit0) · rank 8 of 34
- Unit
- % resolved, higher is better
- Configuration
- GPT-5.2-Codex (openhands)
- Measured
- 30 Jun 2026
- Not shown
- Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
70.9%
Information gathering (GAIA) · rank 12 of 34
- Unit
- % resolved, higher is better
- Configuration
- GPT-5.2-Codex (openhands)
- Measured
- 30 Jun 2026
- Not shown
- Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
73.8%
Issue resolution (SWE-Bench) · rank 19 of 34
- Unit
- % resolved, higher is better
- Configuration
- GPT-5.2-Codex (openhands)
- Measured
- 30 Jun 2026
- Not shown
- Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
67.0%
Testing (SWT-Bench) · rank 20 of 34
- Unit
- % resolved, higher is better
- Configuration
- GPT-5.2-Codex (openhands)
- Measured
- 30 Jun 2026
- Not shown
- Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.