Lightning: strength against thinking time
Each configuration's median time per move at 10 seconds plus 1 a move. Right of the line, a model thinks longer than the increment it earns and the clock drains.
Benchmarks / BulletBench
When every second of thinking time comes off the clock, which models are fast enough to still make good decisions?
Each configuration's median time per move at 10 seconds plus 1 a move. Right of the line, a model thinks longer than the increment it earns and the clock drains.
| # | Model | Lightning 10+1 · 95% range ladder Elo, higher is better | Lightning 10+1 ladder Elo | Blitz 3+2 ladder Elo | Move time milliseconds | Lost on time % of games | Cost per game US dollars |
|---|---|---|---|---|---|---|---|
| 1 | Gemini 3.8 Flash (low reasoning)Google |
1,068 888–1,241 |
197 | – | 1.8 s | 83.3% | $0.0041 |
| 1 | Gemini 3.1 Flash Lite (minimal reasoning)Google |
912 697–1,087 |
780 | – | 0.8 s | 0.0% | $0.0058 |
| 1 | Gemini 3.5 Flash (minimal reasoning)Google |
810 582–1,006 |
727 | – | 1.2 s | 33.3% | $0.0213 |
| 1 | Gemini 3.5 Flash LiteGoogle |
780 586–956 |
892 | 707 | 0.9 s | 0.0% | $0.0070 |
| 2 | Gemini 3.6 Flash (minimal reasoning)Google |
628 429–849 |
33 | – | 1.5 s | 91.7% | $0.0050 |
| 2 | Mistral Medium 3.5 (no reasoning)Mistral |
572 428–708 |
529 | – | 0.5 s | 0.0% | $0.0491 |
| 2 | Gemini 3.1 Pro Preview (low reasoning)Google |
529 329–698 |
– | – | – | – | – |
| 3 | Gemini 3.8 FlashGoogle |
529 263–688 |
– | 1,126 | – | – | – |
| 3 | Inkling Small (no reasoning)Thinkingmachines |
529 393–652 |
571 | – | 0.6 s | 0.0% | $0.0077 |
| 3 | Claude Haiku 4.5Anthropic |
445 135–626 |
33 | 491 | 1.3 s | 83.3% | $0.0095 |
| 3 | GPT-6 Sol (no reasoning)OpenAI |
445 215–639 |
– | – | – | – | – |
| 5 | Seed-2.0-Mini (minimal reasoning)Bytedance |
453 355–538 |
571 | – | 0.3 s | 0.0% | $0.0019 |
| 5 | Mistral Medium 3.5Mistral |
385 276–464 |
529 | 555 | 0.5 s | 0.0% | $0.0297 |
| 5 | Mistral Small 4 (no reasoning)Mistral |
367 126–564 |
571 | – | 0.6 s | 0.0% | $0.0041 |
| 5 | GPT-6 Astra (low reasoning)OpenAI |
323 0–528 |
– | – | – | – | – |
| 5 | GPT-6 Luna (no reasoning)OpenAI |
249 0–453 |
– | – | – | – | – |
| 5 | GPT-5.4 Mini (no reasoning)OpenAI |
197 0–445 |
571 | – | 0.9 s | 0.0% | $0.0136 |
| 5 | GPT-6 AstraOpenAI |
197 0–439 |
– | 839 | – | – | – |
| 11 | Muse Spark 1.3 (minimal reasoning)Meta |
33 0–270 |
– | – | – | – | – |
| 11 | Kimi K3 (low reasoning)Moonshot AI |
33 0–270 |
– | – | – | – | – |
| 11 | GPT-5.4 Nano (no reasoning)OpenAI |
33 0–273 |
197 | – | 1.0 s | 75.0% | $0.0025 |
| 11 | GLM 5.3 (low reasoning)Z.ai |
33 0–270 |
– | – | – | – | – |
| 11 | GLM 5.3Z.ai |
33 0–270 |
– | – | – | – | – |
Models share a rank when their ranges overlap. Each model runs at its provider's default reasoning setting and, where the provider offers a faster one, again at its fastest setting, listed separately. Blitz 3+2 was run at default settings only.
These configurations lost their first four games on time, so they were stopped there and are not rated for that clock. Being too slow is the result, not a fault.
Each model's clock after every move of one game. A fast model earns back more than it spends; a slow one runs out within a few moves, whatever the position.
The same at 60 seconds for the whole game, no increment. A 40-move game leaves 1.5 seconds a move.
Share of games lost on time, and moves with no legal reply, by clock.
| Model | On time, Lightning | On time, Bullet | Invalid moves, Lightning | Invalid moves, Bullet |
|---|---|---|---|---|
| Gemini 3.5 Flash LiteGoogle | 0.0% | 0.0% | 0.0% | 0.0% |
| Gemini 3.1 Flash Lite (minimal reasoning)Google | 0.0% | 8.3% | 0.2% | 0.0% |
| Gemini 3.5 Flash (minimal reasoning)Google | 33.3% | 25.0% | 0.6% | 0.5% |
| Seed-2.0-Mini (minimal reasoning)Bytedance | 0.0% | 0.0% | 0.8% | 0.4% |
| Mistral Small 4 (no reasoning)Mistral | 0.0% | 25.0% | 1.0% | 0.6% |
| GPT-5.4 Mini (no reasoning)OpenAI | 0.0% | 50.0% | 0.0% | 0.0% |
| Inkling Small (no reasoning)Thinkingmachines | 0.0% | 0.0% | 5.3% | 7.2% |
| Mistral Medium 3.5 (no reasoning)Mistral | 0.0% | 0.0% | 0.0% | 0.4% |
| Mistral Medium 3.5Mistral | 0.0% | 8.3% | 0.3% | 0.8% |
| Gemini 3.8 Flash (low reasoning)Google | 83.3% | 50.0% | 0.0% | 0.0% |
| GPT-5.4 Nano (no reasoning)OpenAI | 75.0% | 66.7% | 4.0% | 3.1% |
| Claude Haiku 4.5Anthropic | 83.3% | 25.0% | 0.0% | 0.3% |
| Claude Sonnet 5.5 (low reasoning)Anthropic | 91.7% | – | 0.0% | – |
| Gemini 3.6 Flash (minimal reasoning)Google | 91.7% | 41.7% | 0.0% | 0.3% |
The model gets the position, the moves so far, its remaining clock and the legal moves.
It answers with one move. The whole API round trip, including any hidden reasoning, comes off its clock.
Stockfish replies instantly at the model's current ladder level.
Win and the next game is a level up; lose and it is a level down; draw and it stays. The results give a rating with a 95% interval.
Many products need an answer in about a second: routing, triage, autocomplete, live agents. BulletBench asks which models can still think usefully at that speed. Chess gives a hard, objective score, and the clock makes slow thinking a real cost instead of a free extra.
Twelve games per clock give wide intervals: models whose ranges overlap share a rank. Latency depends on the provider's servers on the day. The ladder's Elo anchors come from v1's Stockfish; this edition runs Stockfish 19, so ratings are not comparable with v1's.