Benchmarks / APEX-Agents

Reported by APEX-Agents

APEX-Agents

Share of investment banking, consulting and corporate law tasks an agent completes in a simulated workplace with files and apps, graded against expert criteria (one attempt).

Results dated
1 Oct 2026
Models
39
Unit
% of tasks
Licence
Published with permission (Mercor results via Epoch AI)

Tasks passed: Claude Opus 5.5

Top 15 of 39 results · % of tasks, higher is better. Choose a model to highlight it.Clear highlight

  1. 1 Claude Sonnet 5.5 (max)Anthropic 75.5%
  2. 2 Claude Opus 5.5 (max reasoning)Anthropic 73.5%
  3. 3 Claude Fable 5.1Anthropic 68.6%
  4. 4 Gemini 3.7 FlashGoogle 67.8%
  5. 5 Claude Opus 5 (max)Anthropic 65.8%
  6. 6 Grok 4.6xAI 65.3%
  7. 7 GPT-6 AstraOpenAI 64.7%
  8. 8 Gemini 3.8 FlashGoogle 64.3%
  9. 9 Claude Fable 5Anthropic 63.6%
  10. 10 GPT-6.1 Sol (max)OpenAI 60.0%
  11. 11 Claude Fable 5.1 (high reasoning)Anthropic 59.7%
  12. 12 GPT-5.6 Terra (max)OpenAI 58.2%
  13. 13 Muse Spark 1.3Meta 57.8%
  14. 14 GLM 5.3Z.ai 56.6%
  15. 15 Grok 4.5xAI 56.2%

Full results

APEX-Agents: tasks passed, % of tasks, higher is better
#ModelTasks passed
% of tasks, higher is better
1 Claude Sonnet 5.5 (max)Anthropic
75.5%
2 Claude Opus 5.5 (max reasoning)Anthropic
73.5%
3 Claude Fable 5.1Anthropic
68.6%
4 Gemini 3.7 FlashGoogle
67.8%
5 Claude Opus 5 (max)Anthropic
65.8%
6 Grok 4.6xAI
65.3%
7 GPT-6 AstraOpenAI
64.7%
8 Gemini 3.8 FlashGoogle
64.3%
9 Claude Fable 5Anthropic
63.6%
10 GPT-6.1 Sol (max)OpenAI
60.0%
11 Claude Fable 5.1 (high reasoning)Anthropic
59.7%
12 GPT-5.6 Terra (max)OpenAI
58.2%
13 Muse Spark 1.3Meta
57.8%
14 GLM 5.3Z.ai
56.6%
15 Grok 4.5xAI
56.2%
16 GPT-5.5OpenAI
55.1%
17 Claude Sonnet 5Anthropic
54.5%
18 GPT-6 Sol (max)OpenAI
54.3%
19 GLM 5.3 FlashZ.ai
52.8%
20 GPT-5.4OpenAI
52.4%
21 GPT-5.6 Sol Pro (max)OpenAI
51.4%
22 Kimi K3Moonshot AI
50.6%
23 Claude Opus 4.7 (max reasoning)Anthropic
49.2%
24 Claude Opus 4.8 (max reasoning)Anthropic
48.9%
25 DeepSeek V4 Pro 0813DeepSeek
47.3%
26 Gemini 3.6 FlashGoogle
46.9%
27 Claude Opus 4.6 (max reasoning)Anthropic
46.3%
28 Claude Sonnet 5.5 (medium)Anthropic
44.6%
29 GPT-6 Luna (max)OpenAI
44.3%
30 Claude Sonnet 4.6 (high reasoning)Anthropic
43.0%
31 GLM 5.1Z.ai
40.9%
32 MiniMax M3MiniMax
37.7%
33 Kimi K2.7 CodeMoonshot AI
37.6%
34 Muse Spark 1.2Meta
36.4%
35 Gemini 3.1 Pro PreviewGoogle
35.3%
36 InklingThinkingmachines
33.8%
37 Gemini 3.5 FlashGoogle
27.5%
38 Qwen3.5 397B A17BAlibaba
24.9%
39 gpt-oss-120bOpenAI
4.4%

Results as published by APEX-Agents; we do not re-run them.

What it measures

Share of investment banking, consulting and corporate law tasks an agent completes in a simulated workplace with files and apps, graded against expert criteria (one attempt).

What it does not measure

Not your firm's documents or tools; graded by rubric, not by a client.

Mercor; collected by Epoch AI. Licence: Published with permission (Mercor results via Epoch AI).