Models / Claude Fable 5

Anthropic

Claude Fable 5

36 published results from 4 sources. Each card shows where the number comes from and what it does not measure. Results are never combined into one score.

Provider
Anthropic
Sources
4
Our benchmarks
0
Price
Not yet published

Reported by others

0.03
Confirmed task success · rank 12 of 46
Unit
IPS effect estimate, higher is better
Range
0.01 to 0.06
Sample
31532 observations
Configuration
Claude Fable 5
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
0.17
Praise over complaint · rank 3 of 46
Unit
IPS effect estimate, higher is better
Range
0.12 to 0.21
Sample
10471 observations
Configuration
Claude Fable 5
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
0.11
Steerability · rank 1 of 46
Unit
IPS effect estimate, higher is better
Range
0.09 to 0.14
Sample
32795 observations
Configuration
Claude Fable 5
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
0.00
Tool grounding · rank 1 of 46
Unit
IPS effect estimate, higher is better
Range
0.00 to 0.00
Sample
2120388 observations
Configuration
Claude Fable 5
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
1,499
Overall · rank 1 of 44
Unit
Arena rating, higher is better
Range
1,491 to 1,507
Sample
8497 votes
Configuration
Claude Fable 5
Measured
13 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,229
Overall · rank 2 of 34
Unit
Arena rating, higher is better
Range
1,221 to 1,237
Sample
41795 votes
Configuration
Claude Fable 5
Measured
24 Aug 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,483
Overall · rank 3 of 177
Unit
Arena rating, higher is better
Range
1,480 to 1,487
Sample
36148 votes
Configuration
Claude Fable 5 (high)
Measured
25 Sep 2026
Not shown
Not an error rate: claims that cannot be checked on the web are skipped, and preference still carries most of the weight.
1,503
Business, management and finance · rank 1 of 402
Unit
Arena rating, higher is better
Range
1,496 to 1,511
Sample
7467 votes
Configuration
Claude Fable 5 (high)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,501
Creative writing · rank 1 of 407
Unit
Arena rating, higher is better
Range
1,493 to 1,509
Sample
7516 votes
Configuration
Claude Fable 5 (high)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,550
Expert prompts · rank 1 of 359
Unit
Arena rating, higher is better
Range
1,540 to 1,560
Sample
4088 votes
Configuration
Claude Fable 5 (high)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,509
Instruction following · rank 1 of 409
Unit
Arena rating, higher is better
Range
1,503 to 1,515
Sample
13049 votes
Configuration
Claude Fable 5 (high)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,504
Overall · rank 1 of 409
Unit
Arena rating, higher is better
Range
1,499 to 1,508
Sample
36462 votes
Configuration
Claude Fable 5 (high)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,508
Writing, literature and language · rank 1 of 408
Unit
Arena rating, higher is better
Range
1,501 to 1,515
Sample
9956 votes
Configuration
Claude Fable 5 (high)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
Reported by OpenHands Index
81.0%
Average · rank 1 of 34
Unit
% resolved, higher is better
Configuration
claude-fable-5
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by OpenHands Index
70.6%
Front end (SWE-Bench Multimodal) · rank 1 of 34
Unit
% resolved, higher is better
Configuration
claude-fable-5
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by OpenHands Index
62.5%
Greenfield (Commit0) · rank 1 of 34
Unit
% resolved, higher is better
Configuration
claude-fable-5
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by OpenHands Index
84.2%
Information gathering (GAIA) · rank 2 of 34
Unit
% resolved, higher is better
Configuration
claude-fable-5
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by OpenHands Index
95.8%
Issue resolution (SWE-Bench) · rank 1 of 34
Unit
% resolved, higher is better
Configuration
claude-fable-5
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by OpenHands Index
91.9%
Testing (SWT-Bench) · rank 1 of 34
Unit
% resolved, higher is better
Configuration
claude-fable-5
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by tau2-bench
28.9%
Consistency (pass^4) · rank 5 of 21
Unit
% of tasks, higher is better
Configuration
claude-fable-5
Measured
23 Jul 2026
Not shown
Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.
Reported by tau2-bench
39.7%
Task success (pass^1) · rank 8 of 21
Unit
% of tasks, higher is better
Configuration
claude-fable-5
Measured
23 Jul 2026
Not shown
Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.
Reported by UGI Leaderboard
2.0%
Requested-length error · rank 7 of 370
Unit
% off the requested word count, lower is better
Configuration
Claude Fable 5 (high)
Measured
10 Jun 2026
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
2.0%
Requested-length error · rank 7 of 370
Unit
% off the requested word count, lower is better
Configuration
claude-fable-5
Measured
10 Jun 2026
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
2.0%
Requested-length error · rank 7 of 370
Unit
% off the requested word count, lower is better
Configuration
claude-fable-5
Measured
10 Jun 2026
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
2.0%
Requested-length error · rank 7 of 370
Unit
% off the requested word count, lower is better
Configuration
claude-fable-5
Measured
10 Jun 2026
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
2.0%
Requested-length error · rank 7 of 370
Unit
% off the requested word count, lower is better
Configuration
claude-fable-5
Measured
10 Jun 2026
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
0.40
Style adherence · rank 30 of 370
Unit
score from 0 to 1, higher is better
Configuration
Claude Fable 5 (high)
Measured
10 Jun 2026
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
0.39
Style adherence · rank 40 of 370
Unit
score from 0 to 1, higher is better
Configuration
claude-fable-5
Measured
10 Jun 2026
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
0.38
Style adherence · rank 71 of 370
Unit
score from 0 to 1, higher is better
Configuration
claude-fable-5
Measured
10 Jun 2026
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
0.38
Style adherence · rank 54 of 370
Unit
score from 0 to 1, higher is better
Configuration
claude-fable-5
Measured
10 Jun 2026
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
0.41
Style adherence · rank 17 of 370
Unit
score from 0 to 1, higher is better
Configuration
claude-fable-5
Measured
10 Jun 2026
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
74.2
Writing score · rank 5 of 370
Unit
score out of 100, higher is better
Configuration
Claude Fable 5 (high)
Measured
10 Jun 2026
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.
Reported by UGI Leaderboard
71.9
Writing score · rank 14 of 370
Unit
score out of 100, higher is better
Configuration
claude-fable-5
Measured
10 Jun 2026
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.
Reported by UGI Leaderboard
73.7
Writing score · rank 6 of 370
Unit
score out of 100, higher is better
Configuration
claude-fable-5
Measured
10 Jun 2026
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.
Reported by UGI Leaderboard
72.6
Writing score · rank 8 of 370
Unit
score out of 100, higher is better
Configuration
claude-fable-5
Measured
10 Jun 2026
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.
Reported by UGI Leaderboard
71.7
Writing score · rank 15 of 370
Unit
score out of 100, higher is better
Configuration
claude-fable-5
Measured
10 Jun 2026
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.

Compare Claude Fable 5 with