Models / Claude Haiku 4.5

Anthropic

Claude Haiku 4.5

54 published results from 7 sources. Each card shows where the number comes from and what it does not measure. Results are never combined into one score.

Provider
Anthropic
Sources
7
Our benchmarks
3
Price
Not yet published

Measured by Spring Prompt

Measured by Spring Prompt · BulletBench
491
Ladder Elo · rank 2 of 16
Unit
ladder Elo, higher is better
Range
299 to 659
Sample
8 games
Configuration
Claude Haiku 4.5, provider default reasoning
Measured
29 Sep 2026
Not shown
Not a business skill: chess against an engine on a real clock, used here as a speed-under-pressure signal.
Measured by Spring Prompt · BulletBench
1.3 s
Median move time · rank 3 of 16
Unit
milliseconds, lower is better
Configuration
Claude Haiku 4.5, provider default reasoning
Measured
29 Sep 2026
Not shown
Not general response speed: replies are one short move.
Measured by Spring Prompt · CatalogBench
94.4%
Content quality · rank 15 of 18
Unit
% of checks, higher is better
Configuration
Claude Haiku 4.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
0.0%
Failed outputs · rank 1 of 18
Unit
% of products, lower is better
Configuration
Claude Haiku 4.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not a provider outage: provider errors are retried before a product counts as failed.
Measured by Spring Prompt · CatalogBench
0.0%
Publish-ready listings · rank 15 of 18
Unit
% of products, higher is better
Configuration
Claude Haiku 4.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
0.0%
Reliably publish-ready · rank 10 of 18
Unit
% of products, higher is better
Configuration
Claude Haiku 4.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
14.3
Unsupported claims · rank 18 of 18
Unit
claims per product, lower is better
Configuration
Claude Haiku 4.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
76.8%
Channel compliance · rank 18 of 18
Unit
% of products, higher is better
Configuration
Claude Haiku 4.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
100.0%
Conflicts caught · rank 1 of 18
Unit
% of conflicts, higher is better
Configuration
Claude Haiku 4.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
94.6%
Content quality · rank 15 of 18
Unit
% of checks, higher is better
Configuration
Claude Haiku 4.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
$0.0061
Cost per product · rank 4 of 18
Unit
US dollars, lower is better
Configuration
Claude Haiku 4.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your cost: prices are those charged through OpenRouter on the run date.
Measured by Spring Prompt · CatalogBench
86.0%
Decision accuracy · rank 17 of 18
Unit
% of decisions, higher is better
Configuration
Claude Haiku 4.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
0.0%
Failed outputs · rank 1 of 18
Unit
% of products, lower is better
Configuration
Claude Haiku 4.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not a provider outage: provider errors are retried before a product counts as failed.
Measured by Spring Prompt · CatalogBench
93.7%
Field accuracy · rank 7 of 18
Unit
% of missing fields, higher is better
Configuration
Claude Haiku 4.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
6.2%
Invented values · rank 18 of 18
Unit
% of filled values, lower is better
Configuration
Claude Haiku 4.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
0.6%
Publish-ready listings · rank 18 of 18
Unit
% of products, higher is better
Configuration
Claude Haiku 4.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
0.0%
Reliably publish-ready · rank 17 of 18
Unit
% of products, higher is better
Configuration
Claude Haiku 4.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
7.92
Unsupported claims · rank 18 of 18
Unit
claims per product, lower is better
Configuration
Claude Haiku 4.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · ROASBench
$0.0298
Cost of a run · rank 2 of 17
Unit
US dollars, lower is better
Configuration
Claude Haiku 4.5, provider default reasoning
Measured
29 Sep 2026
Not shown
Not your cost: prices are those charged through OpenRouter on the run date.
Measured by Spring Prompt · ROASBench
2
Months over budget · rank 15 of 17
Unit
months of 12, lower is better
Configuration
Claude Haiku 4.5, provider default reasoning
Measured
29 Sep 2026
Not shown
Not real ad performance: the market is a deterministic simulation of one invented skincare brand.
Measured by Spring Prompt · ROASBench
17.5
Overall score · rank 16 of 17
Unit
score out of 100, higher is better
Configuration
Claude Haiku 4.5, provider default reasoning
Measured
29 Sep 2026
Not shown
Not real ad performance: the market is a deterministic simulation of one invented skincare brand.
Measured by Spring Prompt · ROASBench
−£0.21
Return on ad spend · rank 17 of 17
Unit
profit per £1 spent, higher is better
Configuration
Claude Haiku 4.5, provider default reasoning
Measured
29 Sep 2026
Not shown
Not real ad performance: the market is a deterministic simulation of one invented skincare brand.

Reported by others

1,419
Overall · rank 38 of 44
Unit
Arena rating, higher is better
Range
1,413 to 1,425
Sample
34677 votes
Configuration
Claude Haiku 4.5
Measured
13 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,418
Overall · rank 117 of 177
Unit
Arena rating, higher is better
Range
1,416 to 1,420
Sample
133816 votes
Configuration
Claude Haiku 4.5
Measured
25 Sep 2026
Not shown
Not an error rate: claims that cannot be checked on the web are skipped, and preference still carries most of the weight.
1,420
Business, management and finance · rank 106 of 402
Unit
Arena rating, higher is better
Range
1,415 to 1,424
Sample
27420 votes
Configuration
Claude Haiku 4.5
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,387
Creative writing · rank 99 of 407
Unit
Arena rating, higher is better
Range
1,382 to 1,392
Sample
23265 votes
Configuration
Claude Haiku 4.5
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,452
Expert prompts · rank 84 of 359
Unit
Arena rating, higher is better
Range
1,446 to 1,458
Sample
12831 votes
Configuration
Claude Haiku 4.5
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,415
Instruction following · rank 99 of 409
Unit
Arena rating, higher is better
Range
1,411 to 1,419
Sample
44952 votes
Configuration
Claude Haiku 4.5
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,414
Overall · rank 122 of 409
Unit
Arena rating, higher is better
Range
1,411 to 1,416
Sample
140791 votes
Configuration
Claude Haiku 4.5
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,399
Writing, literature and language · rank 96 of 408
Unit
Arena rating, higher is better
Range
1,395 to 1,404
Sample
33818 votes
Configuration
Claude Haiku 4.5
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
85.1%
Irrelevance detection · rank 32 of 109
Unit
% correct, higher is better
Configuration
claude-haiku-4.5
Measured
16 Dec 2025
Not shown
Not general refusal or safety behaviour.
95.3%
Irrelevance detection · rank 4 of 109
Unit
% correct, higher is better
Configuration
claude-haiku-4.5
Measured
16 Dec 2025
Not shown
Not general refusal or safety behaviour.
54.4%
Memory · rank 7 of 109
Unit
% correct, higher is better
Configuration
claude-haiku-4.5
Measured
16 Dec 2025
Not shown
Not long-term personal memory in a product: sessions are BFCL's scripted ones.
2.6%
Memory · rank 99 of 109
Unit
% correct, higher is better
Configuration
claude-haiku-4.5
Measured
16 Dec 2025
Not shown
Not long-term personal memory in a product: sessions are BFCL's scripted ones.
53.6%
Multi-turn tasks · rank 14 of 109
Unit
% correct, higher is better
Configuration
claude-haiku-4.5
Measured
16 Dec 2025
Not shown
Not open-ended agent work: the tools and tasks are BFCL's simulated APIs.
1.8%
Multi-turn tasks · rank 96 of 109
Unit
% correct, higher is better
Configuration
claude-haiku-4.5
Measured
16 Dec 2025
Not shown
Not open-ended agent work: the tools and tasks are BFCL's simulated APIs.
68.7%
Overall accuracy · rank 6 of 109
Unit
% correct, higher is better
Configuration
claude-haiku-4.5
Measured
16 Dec 2025
Not shown
Not a neutral average: the weighting is BFCL's. Not reliability on your own tools: BFCL's functions and queries are a fixed test set, and a correct call is judged by its form, not by what it achieved.
25.3%
Overall accuracy · rank 87 of 109
Unit
% correct, higher is better
Configuration
claude-haiku-4.5
Measured
16 Dec 2025
Not shown
Not a neutral average: the weighting is BFCL's. Not reliability on your own tools: BFCL's functions and queries are a fixed test set, and a correct call is judged by its form, not by what it achieved.
62.5%
Relevance detection · rank 91 of 109
Unit
% correct, higher is better
Configuration
claude-haiku-4.5
Measured
16 Dec 2025
Not shown
Not whether the call itself was right; only that one was attempted.
31.2%
Relevance detection · rank 106 of 109
Unit
% correct, higher is better
Configuration
claude-haiku-4.5
Measured
16 Dec 2025
Not shown
Not whether the call itself was right; only that one was attempted.
86.5%
Single-turn calls (curated) · rank 34 of 109
Unit
% correct, higher is better
Configuration
claude-haiku-4.5
Measured
16 Dec 2025
Not shown
Not reliability on your own tools: BFCL's functions and queries are a fixed test set, and a correct call is judged by its form, not by what it achieved.
55.4%
Single-turn calls (curated) · rank 98 of 109
Unit
% correct, higher is better
Configuration
claude-haiku-4.5
Measured
16 Dec 2025
Not shown
Not reliability on your own tools: BFCL's functions and queries are a fixed test set, and a correct call is judged by its form, not by what it achieved.
78.7%
Single-turn calls (user-contributed) · rank 19 of 109
Unit
% correct, higher is better
Configuration
claude-haiku-4.5
Measured
16 Dec 2025
Not shown
Not reliability on your own tools: BFCL's functions and queries are a fixed test set, and a correct call is judged by its form, not by what it achieved.
52.5%
Single-turn calls (user-contributed) · rank 98 of 109
Unit
% correct, higher is better
Configuration
claude-haiku-4.5
Measured
16 Dec 2025
Not shown
Not reliability on your own tools: BFCL's functions and queries are a fixed test set, and a correct call is judged by its form, not by what it achieved.
83.5%
Web search · rank 2 of 109
Unit
% correct, higher is better
Configuration
claude-haiku-4.5
Measured
16 Dec 2025
Not shown
Not general research quality: questions have short, checkable answers.
19.5%
Web search · rank 45 of 109
Unit
% correct, higher is better
Configuration
claude-haiku-4.5
Measured
16 Dec 2025
Not shown
Not general research quality: questions have short, checkable answers.
Reported by UGI Leaderboard
6.0%
Requested-length error · rank 41 of 370
Unit
% off the requested word count, lower is better
Configuration
claude-haiku-4.5
Measured
15 Oct 2025
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
24.0%
Requested-length error · rank 217 of 370
Unit
% off the requested word count, lower is better
Configuration
claude-haiku-4.5
Measured
15 Oct 2025
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
0.35
Style adherence · rank 178 of 370
Unit
score from 0 to 1, higher is better
Configuration
claude-haiku-4.5
Measured
15 Oct 2025
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
0.39
Style adherence · rank 43 of 370
Unit
score from 0 to 1, higher is better
Configuration
claude-haiku-4.5
Measured
15 Oct 2025
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
52.2
Writing score · rank 145 of 370
Unit
score out of 100, higher is better
Configuration
claude-haiku-4.5
Measured
15 Oct 2025
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.
Reported by UGI Leaderboard
49.1
Writing score · rank 162 of 370
Unit
score out of 100, higher is better
Configuration
claude-haiku-4.5
Measured
15 Oct 2025
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.
99.5%
Answer rate · rank 53 of 108
Unit
% of documents, higher is better
Configuration
Claude Haiku 4.5
Measured
22 Sep 2026
Not shown
Not a quality score: a low rate usually means content filters were triggered, and hallucination rates are measured on answered documents only.
9.8%
Hallucination rate · rank 58 of 108
Unit
% of summaries, lower is better
Configuration
Claude Haiku 4.5
Measured
22 Sep 2026
Not shown
Not errors in open questions or other tasks: only summarisation, judged by Vectara's own model (HHEM-2.3), not by people, on news-style documents rather than your data.

Compare Claude Haiku 4.5 with