Benchmarks / ROASBench

Measured by Spring Prompt

ROASBench

Given twelve months of paid-marketing decisions for a simulated brand, which models grow the business and which overspend?

Results dated
29 Sep 2026
Models
17
Unit
score out of 100
Licence
Spring Prompt original
ROASBench: overall score, score out of 100, higher is better
#ModelOverall score
score out of 100, higher is better
Return on ad spend
profit per £1 spent
Months over budget
months of 12
Cost of a run
US dollars
1 GPT-6 AstraOpenAI
56.6
£0.670$1.23
2 GPT-6 SolOpenAI
54.1
£0.650$0.24
3 GPT-6 LunaOpenAI
53.4
£0.660$0.0142
4 Claude Sonnet 5.5Anthropic
52.0
£0.700$0.42
5 Claude Opus 5.5Anthropic
51.6
£0.690$1.06
6 Claude Fable 5.1Anthropic
51.5
£0.640$2.49
7 Gemini 3.8 FlashGoogle
51.4
£0.650$0.16
8 Kimi K3Moonshot AI
44.1
£0.500$1.01
9 Gemini 3.1 Pro PreviewGoogle
43.5
£0.420$0.52
10 Grok 4.7xAI
36.6
£0.230$0.42
11 Muse Spark 1.3Meta
34.1
£0.290$0.18
12 DeepSeek V4 Pro 0423DeepSeek
28.2
£0.260$0.20
13 Qwen3.8 Max (0902)Alibaba
25.4
£0.092$0.73
14 Gemini 3.5 Flash LiteGoogle
25.4
£0.010$0.0390
15 GLM 5.3Z.ai
24.8
£0.010$0.13
16 Claude Haiku 4.5Anthropic
17.5
−£0.212$0.0298
17 Mistral Medium 3.5Mistral
15.2
−£0.173$0.15

Each model runs at its provider's default reasoning setting. Some providers think at length by default and others barely at all, so this is what you get without tuning.

What it measures

  • Budget allocation across channels, month by month
  • Reacting to last month's results
  • Staying within budget

What it does not measure

  • Real ad performance: the market is a deterministic simulation
  • Creative quality of the ad copy
  • Any brand other than one invented skincare company

Method

  • A deterministic simulator scores every plan, so runs are reproducible
  • Critical failures (overspend, broken plans) are reported per model
  • Scores are not comparable with the 2026 v1 results