Models / Claude Opus 4.8

Anthropic

Claude Opus 4.8

44 published results from 4 sources. Each card shows where the number comes from and what it does not measure. Results are never combined into one score.

Provider
Anthropic
Sources
4
Our benchmarks
0
Price
Not yet published

Reported by others

0.04
Confirmed task success · rank 10 of 46
Unit
IPS effect estimate, higher is better
Range
0.01 to 0.06
Sample
29607 observations
Configuration
Claude Opus 4.8
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
0.15
Praise over complaint · rank 3 of 46
Unit
IPS effect estimate, higher is better
Range
0.11 to 0.20
Sample
9301 observations
Configuration
Claude Opus 4.8
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
0.10
Steerability · rank 1 of 46
Unit
IPS effect estimate, higher is better
Range
0.07 to 0.12
Sample
30485 observations
Configuration
Claude Opus 4.8
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
-0.00
Tool grounding · rank 27 of 46
Unit
IPS effect estimate, higher is better
Range
-0.00 to 0.00
Sample
2072136 observations
Configuration
Claude Opus 4.8
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
1,475
Overall · rank 7 of 44
Unit
Arena rating, higher is better
Range
1,468 to 1,482
Sample
12128 votes
Configuration
Claude Opus 4.8
Measured
13 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,486
Overall · rank 1 of 44
Unit
Arena rating, higher is better
Range
1,479 to 1,493
Sample
11551 votes
Configuration
Claude Opus 4.8 (high)
Measured
13 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,195
Overall · rank 7 of 34
Unit
Arena rating, higher is better
Range
1,189 to 1,201
Sample
70998 votes
Configuration
Claude Opus 4.8
Measured
24 Aug 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,465
Overall · rank 25 of 177
Unit
Arena rating, higher is better
Range
1,461 to 1,468
Sample
62333 votes
Configuration
Claude Opus 4.8
Measured
25 Sep 2026
Not shown
Not an error rate: claims that cannot be checked on the web are skipped, and preference still carries most of the weight.
1,469
Overall · rank 18 of 177
Unit
Arena rating, higher is better
Range
1,466 to 1,472
Sample
61136 votes
Configuration
Claude Opus 4.8 (high)
Measured
25 Sep 2026
Not shown
Not an error rate: claims that cannot be checked on the web are skipped, and preference still carries most of the weight.
1,487
Business, management and finance · rank 6 of 402
Unit
Arena rating, higher is better
Range
1,481 to 1,493
Sample
12442 votes
Configuration
Claude Opus 4.8
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,490
Business, management and finance · rank 3 of 402
Unit
Arena rating, higher is better
Range
1,483 to 1,496
Sample
12373 votes
Configuration
Claude Opus 4.8 (high)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,462
Creative writing · rank 12 of 407
Unit
Arena rating, higher is better
Range
1,455 to 1,469
Sample
11977 votes
Configuration
Claude Opus 4.8
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,467
Creative writing · rank 9 of 407
Unit
Arena rating, higher is better
Range
1,461 to 1,474
Sample
11912 votes
Configuration
Claude Opus 4.8 (high)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,515
Expert prompts · rank 8 of 359
Unit
Arena rating, higher is better
Range
1,507 to 1,523
Sample
7157 votes
Configuration
Claude Opus 4.8
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,525
Expert prompts · rank 4 of 359
Unit
Arena rating, higher is better
Range
1,517 to 1,533
Sample
6826 votes
Configuration
Claude Opus 4.8 (high)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,478
Instruction following · rank 11 of 409
Unit
Arena rating, higher is better
Range
1,472 to 1,483
Sample
22526 votes
Configuration
Claude Opus 4.8
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,491
Instruction following · rank 4 of 409
Unit
Arena rating, higher is better
Range
1,486 to 1,496
Sample
21963 votes
Configuration
Claude Opus 4.8 (high)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,474
Overall · rank 21 of 409
Unit
Arena rating, higher is better
Range
1,470 to 1,478
Sample
62717 votes
Configuration
Claude Opus 4.8
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,480
Overall · rank 13 of 409
Unit
Arena rating, higher is better
Range
1,476 to 1,484
Sample
61560 votes
Configuration
Claude Opus 4.8 (high)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,467
Writing, literature and language · rank 14 of 408
Unit
Arena rating, higher is better
Range
1,461 to 1,473
Sample
16468 votes
Configuration
Claude Opus 4.8
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,474
Writing, literature and language · rank 8 of 408
Unit
Arena rating, higher is better
Range
1,468 to 1,480
Sample
16315 votes
Configuration
Claude Opus 4.8 (high)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
Reported by OpenHands Index
71.9%
Average · rank 2 of 34
Unit
% resolved, higher is better
Configuration
claude-opus-4.8
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by OpenHands Index
50.0%
Front end (SWE-Bench Multimodal) · rank 2 of 34
Unit
% resolved, higher is better
Configuration
claude-opus-4.8
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by OpenHands Index
62.5%
Greenfield (Commit0) · rank 1 of 34
Unit
% resolved, higher is better
Configuration
claude-opus-4.8
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by OpenHands Index
78.8%
Information gathering (GAIA) · rank 8 of 34
Unit
% resolved, higher is better
Configuration
claude-opus-4.8
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by OpenHands Index
83.8%
Issue resolution (SWE-Bench) · rank 2 of 34
Unit
% resolved, higher is better
Configuration
claude-opus-4.8
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by OpenHands Index
84.3%
Testing (SWT-Bench) · rank 2 of 34
Unit
% resolved, higher is better
Configuration
claude-opus-4.8
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by tau2-bench
22.7%
Consistency (pass^4) · rank 8 of 21
Unit
% of tasks, higher is better
Configuration
claude-opus-4.8
Measured
23 Jul 2026
Not shown
Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.
Reported by tau2-bench
39.7%
Task success (pass^1) · rank 8 of 21
Unit
% of tasks, higher is better
Configuration
claude-opus-4.8
Measured
23 Jul 2026
Not shown
Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.
Reported by UGI Leaderboard
9.0%
Requested-length error · rank 67 of 370
Unit
% off the requested word count, lower is better
Configuration
Claude Opus 4.8 (high)
Measured
29 May 2026
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
7.0%
Requested-length error · rank 48 of 370
Unit
% off the requested word count, lower is better
Configuration
claude-opus-4.8
Measured
29 May 2026
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
9.0%
Requested-length error · rank 67 of 370
Unit
% off the requested word count, lower is better
Configuration
claude-opus-4.8
Measured
29 May 2026
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
10.0%
Requested-length error · rank 81 of 370
Unit
% off the requested word count, lower is better
Configuration
claude-opus-4.8
Measured
29 May 2026
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
6.0%
Requested-length error · rank 41 of 370
Unit
% off the requested word count, lower is better
Configuration
claude-opus-4.8
Measured
29 May 2026
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
0.38
Style adherence · rank 62 of 370
Unit
score from 0 to 1, higher is better
Configuration
Claude Opus 4.8 (high)
Measured
29 May 2026
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
0.38
Style adherence · rank 52 of 370
Unit
score from 0 to 1, higher is better
Configuration
claude-opus-4.8
Measured
29 May 2026
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
0.40
Style adherence · rank 19 of 370
Unit
score from 0 to 1, higher is better
Configuration
claude-opus-4.8
Measured
29 May 2026
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
0.37
Style adherence · rank 97 of 370
Unit
score from 0 to 1, higher is better
Configuration
claude-opus-4.8
Measured
29 May 2026
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
0.39
Style adherence · rank 48 of 370
Unit
score from 0 to 1, higher is better
Configuration
claude-opus-4.8
Measured
29 May 2026
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
64.3
Writing score · rank 71 of 370
Unit
score out of 100, higher is better
Configuration
Claude Opus 4.8 (high)
Measured
29 May 2026
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.
Reported by UGI Leaderboard
64.7
Writing score · rank 68 of 370
Unit
score out of 100, higher is better
Configuration
claude-opus-4.8
Measured
29 May 2026
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.
Reported by UGI Leaderboard
65.9
Writing score · rank 56 of 370
Unit
score out of 100, higher is better
Configuration
claude-opus-4.8
Measured
29 May 2026
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.
Reported by UGI Leaderboard
63.5
Writing score · rank 78 of 370
Unit
score out of 100, higher is better
Configuration
claude-opus-4.8
Measured
29 May 2026
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.
Reported by UGI Leaderboard
65.6
Writing score · rank 59 of 370
Unit
score out of 100, higher is better
Configuration
claude-opus-4.8
Measured
29 May 2026
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.

Compare Claude Opus 4.8 with