Benchmarks / APEX-Agents
Reported by APEX-Agents
APEX-Agents
Share of investment banking, consulting and corporate law tasks an agent completes in a simulated workplace with files and apps, graded against expert criteria (one attempt).
- Results dated
- 1 Oct 2026
- Models
- 39
- Unit
- % of tasks
Full results
APEX-Agents: tasks passed, % of tasks, higher is better
| # | Model | Tasks passed % of tasks, higher is better |
| 1 |
Claude Sonnet 5.5 (max)Anthropic |
|
| 2 |
Claude Opus 5.5 (max reasoning)Anthropic |
|
| 3 |
Claude Fable 5.1Anthropic |
|
| 4 |
Gemini 3.7 FlashGoogle |
|
| 5 |
Claude Opus 5 (max)Anthropic |
|
| 6 |
Grok 4.6xAI |
|
| 7 |
GPT-6 AstraOpenAI |
|
| 8 |
Gemini 3.8 FlashGoogle |
|
| 9 |
Claude Fable 5Anthropic |
|
| 10 |
GPT-6.1 Sol (max)OpenAI |
|
| 11 |
Claude Fable 5.1 (high reasoning)Anthropic |
|
| 12 |
GPT-5.6 Terra (max)OpenAI |
|
| 13 |
Muse Spark 1.3Meta |
|
| 14 |
GLM 5.3Z.ai |
|
| 15 |
Grok 4.5xAI |
|
| 16 |
GPT-5.5OpenAI |
|
| 17 |
Claude Sonnet 5Anthropic |
|
| 18 |
GPT-6 Sol (max)OpenAI |
|
| 19 |
GLM 5.3 FlashZ.ai |
|
| 20 |
GPT-5.4OpenAI |
|
| 21 |
GPT-5.6 Sol Pro (max)OpenAI |
|
| 22 |
Kimi K3Moonshot AI |
|
| 23 |
Claude Opus 4.7 (max reasoning)Anthropic |
|
| 24 |
Claude Opus 4.8 (max reasoning)Anthropic |
|
| 25 |
DeepSeek V4 Pro 0813DeepSeek |
|
| 26 |
Gemini 3.6 FlashGoogle |
|
| 27 |
Claude Opus 4.6 (max reasoning)Anthropic |
|
| 28 |
Claude Sonnet 5.5 (medium)Anthropic |
|
| 29 |
GPT-6 Luna (max)OpenAI |
|
| 30 |
Claude Sonnet 4.6 (high reasoning)Anthropic |
|
| 31 |
GLM 5.1Z.ai |
|
| 32 |
MiniMax M3MiniMax |
|
| 33 |
Kimi K2.7 CodeMoonshot AI |
|
| 34 |
Muse Spark 1.2Meta |
|
| 35 |
Gemini 3.1 Pro PreviewGoogle |
|
| 36 |
InklingThinkingmachines |
|
| 37 |
Gemini 3.5 FlashGoogle |
|
| 38 |
Qwen3.5 397B A17BAlibaba |
|
| 39 |
gpt-oss-120bOpenAI |
|
Results as published by APEX-Agents; we do not re-run them.
What it measures
Share of investment banking, consulting and corporate law tasks an agent completes in a simulated workplace with files and apps, graded against expert criteria (one attempt).
What it does not measure
Not your firm's documents or tools; graded by rubric, not by a client.
Mercor; collected by Epoch AI. Licence: Published with permission (Mercor results via Epoch AI).