Make the AI inside your product measurably better.
We benchmark AI models on real business work and publish everything: results, method and limits. Then we apply the same checks to the AI in your product, so you know what to fix before it ships.
See the benchmarksTalk to us- Our own benchmarks
- Ranges, not just ranks
- Method on every page
The model sees the photos and the supplier's copy.
Porthkerris Cornish Coast London Dry Gin 70cl 41.3% ABV
- ✓41.3% ABV, 70clMatches the back label
- ✓Distilled in ScotlandRead from the label, not the supplier's "Cornish coast"
- ✗"Small-batch distilled"Repeated from the supplier's copy. Nothing supports it.
- ✗"Hand-foraged samphire"Repeated from the supplier's copy. Nothing supports it.
Benchmarks built from real business work.
CatalogBench
Can a model turn a product feed and photos into a listing that is ready to publish, without making things up?
ROASBench
Given a year of ad spend decisions, which models grow revenue and which ones burn the budget?
BulletBench
When thinking time comes off the clock, which models are quick enough to still make good decisions?
Ask for copy that sells, and most models stop being safe to publish.
On the factual brief, the best models wrote listings that were publish-ready in all three runs for about seven products in ten. Given the same facts and asked for copy that sells, OpenAI's GPT-6 models kept most of that. Most other models fell close to zero, largely by adding claims the product data does not support.
Read the findings →Start with the decision, not the leaderboard.
- 01
Define
Agree what good looks like, and which failures must never ship.
- 02
Test
Run models and prompts on real work. Hard checks first, judgement only where needed.
- 03
Improve
Fix what the failures point to, then check the fix on cases held back.
- 04
Prove
A clear call to ship or not, with the evidence and its limits in plain view.
Choosing a model?
Results from our benchmarks and 7 licensed sources, never blended into one score.
Test the AI before your customers do.
Shipping AI in your product? We run the same kind of checks on it, and tell you what to fix first.
Findings and field notes.
- 24 Jul 2026 · Evals
Gemini 3.6 Flash Wins the Minute. Gemini 3.5 Flash-Lite Wins the Second.
Google's two new fast models split BulletBench's hardest latency tests: Gemini 3.6 Flash tops the 60-second board, while Gemini 3.5 Flash-Lite nearly wins Lightning without losing a single game on time.
- 21 Jul 2026 · Evals
AI Is Now a Better Football Pundit Than the Pundit. It Still Lost to One Line of Code.
We had six frontier AI models predict all 32 knockout matches of the 2026 World Cup, blind. They out-called a top BBC pundit. Then a one-line rule, always back the higher-ranked team, beat every single one of them on…
- 21 Jul 2026 · Evals
Six Models, One Mind: What the World Cup Revealed About How AIs Actually Think
Across 34 World Cup fixtures, six frontier AI models from five different labs picked the same winner 31 times. Then we gave them memory of the tournament and they changed zero of 96 picks. The herd is real, and it has…