Findings and field notes.
What our benchmarks found, how we run them, and notes on building AI products you can trust.
- 24 Jul 2026 · Evals
Gemini 3.6 Flash Wins the Minute. Gemini 3.5 Flash-Lite Wins the Second.
Google's two new fast models split BulletBench's hardest latency tests: Gemini 3.6 Flash tops the 60-second board, while Gemini 3.5 Flash-Lite nearly wins Lightning without losing a single game on time.
- 21 Jul 2026 · Evals
AI Is Now a Better Football Pundit Than the Pundit. It Still Lost to One Line of Code.
We had six frontier AI models predict all 32 knockout matches of the 2026 World Cup, blind. They out-called a top BBC pundit. Then a one-line rule, always back the higher-ranked team, beat every single one of them on…
- 21 Jul 2026 · Evals
Six Models, One Mind: What the World Cup Revealed About How AIs Actually Think
Across 34 World Cup fixtures, six frontier AI models from five different labs picked the same winner 31 times. Then we gave them memory of the tournament and they changed zero of 96 picks. The herd is real, and it has…
- 21 Jul 2026 · Evals
The AI That Won Our World Cup Benchmark Didn't Predict Best. It Read the Rules.
GLM-5.2 won World Cup Bench by two points while spending the fewest reasoning tokens and posting an ordinary hit rate. It found the one free lunch in our scoring rules and ate it 34 times. A story about Goodhart's law…
- 20 Jul 2026 · AI
How to Continue a Claude Code Conversation in Codex After Hitting Your Usage Limit
Ran out of Claude Code allowance halfway through a task? Here is how to move the conversation into Codex and continue without losing the decisions behind the code.
- 13 Jul 2026 · Evals
GPT-5.6 Luna Was Meant to Be Fast. It Still Lost to the Clock.
Luna arrived as the fastest GPT-5.6. Across 96 BulletBench games, it lost 71 on time. The result exposes the model the market still has not built: genuinely smart, sub-second and predictable.
- 3 Jul 2026 · Evals
BulletBench: we made 23 AI models play speed chess, and the clock was the judge
Every leaderboard measures how smart a model is. None measure how smart it is per second. So we built one: AI models play speed chess on a real clock, and slow genius loses on time.
- 1 Jul 2026 · Evals
Sonnet 5 gave the sharpest marketing analysis we've ever tested. It also lost money every single time.
A ROASBench case study in why fluent reasoning isn't the same as good decisions, and why for this model more thinking made it worse.
- 25 Mar 2026 · AI
Claude vs Gemini vs GPT in a 12-Month Marketing Simulation
ROAS Bench is one of the clearest examples of where frontier models diverge in practice: not on prose quality, but on economic compounding.
- 25 Mar 2026 · AI
How We Built ROASBench to Feel Like Real Growth Work
ROASBench was built from operator experience, not benchmark theater. We designed it to feel like real performance marketing: stateful, constrained, path-dependent, and economically unforgiving.
- 25 Mar 2026
LiteLLM alternatives for 2026
If you’re looking for LiteLLM alternatives, you’re usually trying to solve one of two problems: * you need a Python library that makes it easy to switch between LLM providers * you need an AI gateway / routing layer…
- 23 Mar 2026 · AI
Why Most LLMs Still Can't Run Growth
ROAS Bench shows that growth is not a copywriting problem. It is a compounding systems problem, and most models still break the system faster than they improve it.
- 20 Mar 2026 · AI
The AI Marketing Benchmark That Punishes Plausible-Sounding Strategy
Most LLMs can sound like a competent growth lead for one turn. ROAS Bench is interesting because it makes them live with the consequences for twelve months.
- 11 Mar 2026
Gemini Embedding 2 Just Launched - So We Benchmarked It
Google launched gemini-embedding-2-preview on March 10, 2026 as its first multimodal embedding model, with one shared embedding space for text, images, video, audio, and PDFs. Google specifically positions it for…
- 17 Dec 2025
The Great AI Gifting Showdown: Which Model Should You Trust for Christmas Shopping?
It’s that time of year again. You’re out and about, the clock is ticking, and you still haven't found the perfect gift for your partner, your roommate, or that difficult-to-shop-for in-law. Naturally, many of us are…
- 20 Nov 2025
Google Gemini 3 Review: The Benchmarks Actually Match the Hype 🤯
So, on Tuesday Google launched Gemini 3. The hype was massive leading up to this, and honestly? It is justified. It is really, really good. Trying to explain how good is difficult without getting bogged down in…
- 13 Nov 2025
GPT-5.1 First Look: Smarter, Warmer… But Not a Breakthrough
2025’s flagship model season kicked off yesterday with the unexpected arrival of GPT-5.1, with OpenAI getting their release out before Gemini 3. While we’re still waiting for API access (and therefore can’t run proper…
- 12 Nov 2025
How to prepare for Gemini 3 + GPT 5.1
Here we go again: new-flagship season. Google’s Gemini 3 has been peeking through A/B tests in AI Studio and docs watchers have noticed model lifecycle shuffles, while OpenAI is lining up a GPT-5.1 family (base…