Models / Claude Opus 5.5

Anthropic

Claude Opus 5.5

47 published results from 5 sources. Each card shows where the number comes from and what it does not measure. Results are never combined into one score.

Provider
Anthropic
Sources
5
Our benchmarks
3
Price
Not yet published

Measured by Spring Prompt

Measured by Spring Prompt · BulletBench
557
Ladder Elo · rank 2 of 16
Unit
ladder Elo, higher is better
Range
370 to 727
Sample
8 games
Configuration
Claude Opus 5.5, provider default reasoning
Measured
29 Sep 2026
Not shown
Not a business skill: chess against an engine on a real clock, used here as a speed-under-pressure signal.
Measured by Spring Prompt · BulletBench
7.5 s
Median move time · rank 12 of 16
Unit
milliseconds, lower is better
Configuration
Claude Opus 5.5, provider default reasoning
Measured
29 Sep 2026
Not shown
Not general response speed: replies are one short move.
Measured by Spring Prompt · CatalogBench
96.6%
Content quality · rank 8 of 18
Unit
% of checks, higher is better
Configuration
Claude Opus 5.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
0.0%
Failed outputs · rank 1 of 18
Unit
% of products, lower is better
Configuration
Claude Opus 5.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not a provider outage: provider errors are retried before a product counts as failed.
Measured by Spring Prompt · CatalogBench
3.6%
Publish-ready listings · rank 9 of 18
Unit
% of products, higher is better
Configuration
Claude Opus 5.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
0.0%
Reliably publish-ready · rank 10 of 18
Unit
% of products, higher is better
Configuration
Claude Opus 5.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
4.36
Unsupported claims · rank 10 of 18
Unit
claims per product, lower is better
Configuration
Claude Opus 5.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
100.0%
Channel compliance · rank 1 of 18
Unit
% of products, higher is better
Configuration
Claude Opus 5.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
100.0%
Conflicts caught · rank 1 of 18
Unit
% of conflicts, higher is better
Configuration
Claude Opus 5.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
97.2%
Content quality · rank 5 of 18
Unit
% of checks, higher is better
Configuration
Claude Opus 5.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
$0.0400
Cost per product · rank 16 of 18
Unit
US dollars, lower is better
Configuration
Claude Opus 5.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your cost: prices are those charged through OpenRouter on the run date.
Measured by Spring Prompt · CatalogBench
96.6%
Decision accuracy · rank 6 of 18
Unit
% of decisions, higher is better
Configuration
Claude Opus 5.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
0.0%
Failed outputs · rank 1 of 18
Unit
% of products, lower is better
Configuration
Claude Opus 5.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not a provider outage: provider errors are retried before a product counts as failed.
Measured by Spring Prompt · CatalogBench
92.2%
Field accuracy · rank 14 of 18
Unit
% of missing fields, higher is better
Configuration
Claude Opus 5.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
0.5%
Invented values · rank 4 of 18
Unit
% of filled values, lower is better
Configuration
Claude Opus 5.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
58.3%
Publish-ready listings · rank 6 of 18
Unit
% of products, higher is better
Configuration
Claude Opus 5.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
39.3%
Reliably publish-ready · rank 6 of 18
Unit
% of products, higher is better
Configuration
Claude Opus 5.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
0.83
Unsupported claims · rank 8 of 18
Unit
claims per product, lower is better
Configuration
Claude Opus 5.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · ROASBench
$1.06
Cost of a run · rank 15 of 17
Unit
US dollars, lower is better
Configuration
Claude Opus 5.5, provider default reasoning
Measured
29 Sep 2026
Not shown
Not your cost: prices are those charged through OpenRouter on the run date.
Measured by Spring Prompt · ROASBench
0
Months over budget · rank 1 of 17
Unit
months of 12, lower is better
Configuration
Claude Opus 5.5, provider default reasoning
Measured
29 Sep 2026
Not shown
Not real ad performance: the market is a deterministic simulation of one invented skincare brand.
Measured by Spring Prompt · ROASBench
51.6
Overall score · rank 5 of 17
Unit
score out of 100, higher is better
Configuration
Claude Opus 5.5, provider default reasoning
Measured
29 Sep 2026
Not shown
Not real ad performance: the market is a deterministic simulation of one invented skincare brand.
Measured by Spring Prompt · ROASBench
£0.69
Return on ad spend · rank 2 of 17
Unit
profit per £1 spent, higher is better
Configuration
Claude Opus 5.5, provider default reasoning
Measured
29 Sep 2026
Not shown
Not real ad performance: the market is a deterministic simulation of one invented skincare brand.

Reported by others

0.16
Confirmed task success · rank 1 of 46
Unit
IPS effect estimate, higher is better
Range
0.12 to 0.20
Sample
3385 observations
Configuration
Claude Opus 5.5
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
0.19
Praise over complaint · rank 1 of 46
Unit
IPS effect estimate, higher is better
Range
0.12 to 0.27
Sample
1149 observations
Configuration
Claude Opus 5.5
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
0.14
Steerability · rank 1 of 46
Unit
IPS effect estimate, higher is better
Range
0.10 to 0.18
Sample
3699 observations
Configuration
Claude Opus 5.5
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
0.00
Tool grounding · rank 1 of 46
Unit
IPS effect estimate, higher is better
Range
0.00 to 0.00
Sample
814080 observations
Configuration
Claude Opus 5.5
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
1,499
Business, management and finance · rank 1 of 402
Unit
Arena rating, higher is better
Range
1,471 to 1,528
Sample
436 votes
Configuration
Claude Opus 5.5 (high)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,521
Creative writing · rank 1 of 407
Unit
Arena rating, higher is better
Range
1,492 to 1,550
Sample
487 votes
Configuration
Claude Opus 5.5 (high)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,556
Expert prompts · rank 1 of 359
Unit
Arena rating, higher is better
Range
1,522 to 1,591
Sample
277 votes
Configuration
Claude Opus 5.5 (high)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,516
Instruction following · rank 1 of 409
Unit
Arena rating, higher is better
Range
1,496 to 1,537
Sample
812 votes
Configuration
Claude Opus 5.5 (high)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,509
Overall · rank 1 of 409
Unit
Arena rating, higher is better
Range
1,496 to 1,521
Sample
2307 votes
Configuration
Claude Opus 5.5 (high)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,520
Writing, literature and language · rank 1 of 408
Unit
Arena rating, higher is better
Range
1,496 to 1,544
Sample
655 votes
Configuration
Claude Opus 5.5 (high)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
Reported by UGI Leaderboard
5.0%
Requested-length error · rank 34 of 370
Unit
% off the requested word count, lower is better
Configuration
Claude Opus 5.5 (high)
Measured
23 Sep 2026
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
5.0%
Requested-length error · rank 34 of 370
Unit
% off the requested word count, lower is better
Configuration
claude-opus-5.5
Measured
23 Sep 2026
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
2.0%
Requested-length error · rank 7 of 370
Unit
% off the requested word count, lower is better
Configuration
claude-opus-5.5
Measured
23 Sep 2026
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
5.0%
Requested-length error · rank 34 of 370
Unit
% off the requested word count, lower is better
Configuration
claude-opus-5.5
Measured
23 Sep 2026
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
5.0%
Requested-length error · rank 34 of 370
Unit
% off the requested word count, lower is better
Configuration
claude-opus-5.5
Measured
23 Sep 2026
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
0.42
Style adherence · rank 4 of 370
Unit
score from 0 to 1, higher is better
Configuration
Claude Opus 5.5 (high)
Measured
23 Sep 2026
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
0.40
Style adherence · rank 22 of 370
Unit
score from 0 to 1, higher is better
Configuration
claude-opus-5.5
Measured
23 Sep 2026
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
0.43
Style adherence · rank 1 of 370
Unit
score from 0 to 1, higher is better
Configuration
claude-opus-5.5
Measured
23 Sep 2026
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
0.43
Style adherence · rank 2 of 370
Unit
score from 0 to 1, higher is better
Configuration
claude-opus-5.5
Measured
23 Sep 2026
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
0.42
Style adherence · rank 5 of 370
Unit
score from 0 to 1, higher is better
Configuration
claude-opus-5.5
Measured
23 Sep 2026
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
68.4
Writing score · rank 35 of 370
Unit
score out of 100, higher is better
Configuration
Claude Opus 5.5 (high)
Measured
23 Sep 2026
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.
Reported by UGI Leaderboard
70.0
Writing score · rank 22 of 370
Unit
score out of 100, higher is better
Configuration
claude-opus-5.5
Measured
23 Sep 2026
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.
Reported by UGI Leaderboard
68.9
Writing score · rank 31 of 370
Unit
score out of 100, higher is better
Configuration
claude-opus-5.5
Measured
23 Sep 2026
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.
Reported by UGI Leaderboard
69.8
Writing score · rank 24 of 370
Unit
score out of 100, higher is better
Configuration
claude-opus-5.5
Measured
23 Sep 2026
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.
Reported by UGI Leaderboard
69.4
Writing score · rank 28 of 370
Unit
score out of 100, higher is better
Configuration
claude-opus-5.5
Measured
23 Sep 2026
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.

Compare Claude Opus 5.5 with