Benchmarks
What each benchmark measures, who measured it, and how today's models do. Sources are never blended into one score.
Measured by Spring Prompt
- BulletBenchWhen thinking time comes off the clock, which models are quick enough to still make good decisions?
- CatalogBenchCan a model turn a product feed and photos into a listing that is ready to publish, without making things up?
- ROASBenchGiven a year of ad spend decisions, which models grow revenue and which ones burn the budget?
Reported by others
- Arena (formerly LMArena)Head-to-head human preference across all text prompts.
- Berkeley Function Calling Leaderboard (BFCL) V4BFCL's own weighted average across its test categories.
- Microsoft STATE-BenchShare of customer support tasks completed correctly, averaged over repeated runs.
- OpenHands IndexThe equally weighted average of the five category scores below.
- tau2-benchShare of airline bookings and changes tasks completed within policy, averaged over trials.
- UGI LeaderboardUGI's blend of intelligence, style, repetition and length adherence in writing, tuned to average human preference.
- Vectara Hallucination LeaderboardShare of document summaries that contain something the document does not support.