Share of investment banking, consulting and corporate law tasks an agent completes in a simulated workplace with files and apps, graded against expert criteria (one attempt).
APEX-Agents: tasks passed, % of tasks, higher is better
#
Model
Tasks passed % of tasks, higher is better
1
Claude Sonnet 5.5 (max)Anthropic
75.5%
2
Claude Opus 5.5 (max reasoning)Anthropic
73.5%
3
Claude Fable 5.1Anthropic
68.6%
4
Gemini 3.7 FlashGoogle
67.8%
5
Claude Opus 5 (max)Anthropic
65.8%
6
Grok 4.6xAI
65.3%
7
GPT-6 AstraOpenAI
64.7%
8
Gemini 3.8 FlashGoogle
64.3%
9
Claude Fable 5Anthropic
63.6%
10
GPT-6.1 Sol (max)OpenAI
60.0%
11
Claude Fable 5.1 (high reasoning)Anthropic
59.7%
12
GPT-5.6 Terra (max)OpenAI
58.2%
13
Muse Spark 1.3Meta
57.8%
14
GLM 5.3Z.ai
56.6%
15
Grok 4.5xAI
56.2%
16
GPT-5.5OpenAI
55.1%
17
Claude Sonnet 5Anthropic
54.5%
18
GPT-6 Sol (max)OpenAI
54.3%
19
GLM 5.3 FlashZ.ai
52.8%
20
GPT-5.4OpenAI
52.4%
21
GPT-5.6 Sol Pro (max)OpenAI
51.4%
22
Kimi K3Moonshot AI
50.6%
23
Claude Opus 4.7 (max reasoning)Anthropic
49.2%
24
Claude Opus 4.8 (max reasoning)Anthropic
48.9%
25
DeepSeek V4 Pro 0813DeepSeek
47.3%
26
Gemini 3.6 FlashGoogle
46.9%
27
Claude Opus 4.6 (max reasoning)Anthropic
46.3%
28
Claude Sonnet 5.5 (medium)Anthropic
44.6%
29
GPT-6 Luna (max)OpenAI
44.3%
30
Claude Sonnet 4.6 (high reasoning)Anthropic
43.0%
31
GLM 5.1Z.ai
40.9%
32
MiniMax M3MiniMax
37.7%
33
Kimi K2.7 CodeMoonshot AI
37.6%
34
Muse Spark 1.2Meta
36.4%
35
Gemini 3.1 Pro PreviewGoogle
35.3%
36
InklingThinkingmachines
33.8%
37
Gemini 3.5 FlashGoogle
27.5%
38
Qwen3.5 397B A17BAlibaba
24.9%
39
gpt-oss-120bOpenAI
4.4%
Results as published by APEX-Agents; we do not re-run them.
What it measures
Share of investment banking, consulting and corporate law tasks an agent completes in a simulated workplace with files and apps, graded against expert criteria (one attempt).
What it does not measure
Not your firm's documents or tools; graded by rubric, not by a client.