Independent AI evaluation

Make the AI inside your product measurably better.

We benchmark AI models on real business work and publish everything: results, method and limits. Then we apply the same checks to the AI in your product, so you know what to fix before it ships.

See the benchmarksTalk to us
  • Our own benchmarks
  • Ranges, not just ranks
  • Method on every page
Listing check · Porthkerris Cornish Coast GinCatalogBench · cb-050
Gin bottle and a glass on a shop counter Back label of the gin bottle, held in a hand The model sees the photos and the supplier's copy.

Porthkerris Cornish Coast London Dry Gin 70cl 41.3% ABV

  • ✓41.3% ABV, 70clMatches the back label
  • ✓Distilled in ScotlandRead from the label, not the supplier's "Cornish coast"
  • ✗"Small-batch distilled"Repeated from the supplier's copy. Nothing supports it.
  • ✗"Hand-foraged samphire"Repeated from the supplier's copy. Nothing supports it.
Not ready to publish · 6 unsupported claimsBlocks: unverified production claims
Real output from Mistral Medium 3.5 on the marketing brief, 30 September 2026. Synthetic brand and images.
Measured by Spring Prompt

Benchmarks built from real business work.

All benchmarks →
CatalogBench · 30 September 2026

Ask for copy that sells, and most models stop being safe to publish.

On the factual brief, the best models wrote listings that were publish-ready in all three runs for about seven products in ten. Given the same facts and asked for copy that sells, OpenAI's GPT-6 models kept most of that. Most other models fell close to zero, largely by adding claims the product data does not support.

Read the findings →
Reliably publish-ready, % of products
GPT-6 Astra
71% → 66%
GPT-6.1 Sol
70% → 66%
GPT-6 Sol
61% → 55%
GPT-6 Luna
52% → 43%
Grok 4.7
45% → 11%
Claude Opus 5.5
39% → 0%
Muse Spark 1.3
38% → 7%
DeepSeek V4.1 Flash
36% → 11%
Factual briefMarketing brief
How we work

Start with the decision, not the leaderboard.

  1. 01

    Define

    Agree what good looks like, and which failures must never ship.

  2. 02

    Test

    Run models and prompts on real work. Hard checks first, judgement only where needed.

  3. 03

    Improve

    Fix what the failures point to, then check the fix on cases held back.

  4. 04

    Prove

    A clear call to ship or not, with the evidence and its limits in plain view.

Choosing a model?

Results from our benchmarks and 7 licensed sources, never blended into one score.

Test the AI before your customers do.

Shipping AI in your product? We run the same kind of checks on it, and tell you what to fix first.

Research

Findings and field notes.

All posts →