Models / Claude Sonnet 5.5

Anthropic

Claude Sonnet 5.5

22 published results from 3 sources. Each card shows where the number comes from and what it does not measure. Results are never combined into one score.

Provider
Anthropic
Sources
3
Our benchmarks
3
Price
Not yet published

Measured by Spring Prompt

Measured by Spring Prompt · BulletBench
557
Ladder Elo · rank 2 of 16
Unit
ladder Elo, higher is better
Range
370 to 727
Sample
8 games
Configuration
Claude Sonnet 5.5, provider default reasoning
Measured
29 Sep 2026
Not shown
Not a business skill: chess against an engine on a real clock, used here as a speed-under-pressure signal.
Measured by Spring Prompt · BulletBench
8.1 s
Median move time · rank 13 of 16
Unit
milliseconds, lower is better
Configuration
Claude Sonnet 5.5, provider default reasoning
Measured
29 Sep 2026
Not shown
Not general response speed: replies are one short move.
Measured by Spring Prompt · CatalogBench
96.5%
Content quality · rank 10 of 18
Unit
% of checks, higher is better
Configuration
Claude Sonnet 5.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
0.0%
Failed outputs · rank 1 of 18
Unit
% of products, lower is better
Configuration
Claude Sonnet 5.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not a provider outage: provider errors are retried before a product counts as failed.
Measured by Spring Prompt · CatalogBench
8.9%
Publish-ready listings · rank 8 of 18
Unit
% of products, higher is better
Configuration
Claude Sonnet 5.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
1.8%
Reliably publish-ready · rank 8 of 18
Unit
% of products, higher is better
Configuration
Claude Sonnet 5.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
3.97
Unsupported claims · rank 9 of 18
Unit
claims per product, lower is better
Configuration
Claude Sonnet 5.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
99.4%
Channel compliance · rank 9 of 18
Unit
% of products, higher is better
Configuration
Claude Sonnet 5.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
100.0%
Conflicts caught · rank 1 of 18
Unit
% of conflicts, higher is better
Configuration
Claude Sonnet 5.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
96.2%
Content quality · rank 12 of 18
Unit
% of checks, higher is better
Configuration
Claude Sonnet 5.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
$0.0172
Cost per product · rank 11 of 18
Unit
US dollars, lower is better
Configuration
Claude Sonnet 5.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your cost: prices are those charged through OpenRouter on the run date.
Measured by Spring Prompt · CatalogBench
94.5%
Decision accuracy · rank 11 of 18
Unit
% of decisions, higher is better
Configuration
Claude Sonnet 5.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
0.0%
Failed outputs · rank 1 of 18
Unit
% of products, lower is better
Configuration
Claude Sonnet 5.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not a provider outage: provider errors are retried before a product counts as failed.
Measured by Spring Prompt · CatalogBench
91.5%
Field accuracy · rank 16 of 18
Unit
% of missing fields, higher is better
Configuration
Claude Sonnet 5.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
0.8%
Invented values · rank 7 of 18
Unit
% of filled values, lower is better
Configuration
Claude Sonnet 5.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
35.1%
Publish-ready listings · rank 11 of 18
Unit
% of products, higher is better
Configuration
Claude Sonnet 5.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
21.4%
Reliably publish-ready · rank 10 of 18
Unit
% of products, higher is better
Configuration
Claude Sonnet 5.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
1.59
Unsupported claims · rank 10 of 18
Unit
claims per product, lower is better
Configuration
Claude Sonnet 5.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · ROASBench
$0.42
Cost of a run · rank 10 of 17
Unit
US dollars, lower is better
Configuration
Claude Sonnet 5.5, provider default reasoning
Measured
29 Sep 2026
Not shown
Not your cost: prices are those charged through OpenRouter on the run date.
Measured by Spring Prompt · ROASBench
0
Months over budget · rank 1 of 17
Unit
months of 12, lower is better
Configuration
Claude Sonnet 5.5, provider default reasoning
Measured
29 Sep 2026
Not shown
Not real ad performance: the market is a deterministic simulation of one invented skincare brand.
Measured by Spring Prompt · ROASBench
52.0
Overall score · rank 4 of 17
Unit
score out of 100, higher is better
Configuration
Claude Sonnet 5.5, provider default reasoning
Measured
29 Sep 2026
Not shown
Not real ad performance: the market is a deterministic simulation of one invented skincare brand.
Measured by Spring Prompt · ROASBench
£0.70
Return on ad spend · rank 1 of 17
Unit
profit per £1 spent, higher is better
Configuration
Claude Sonnet 5.5, provider default reasoning
Measured
29 Sep 2026
Not shown
Not real ad performance: the market is a deterministic simulation of one invented skincare brand.

Compare Claude Sonnet 5.5 with