Lightning: strength against thinking time
Each configuration's median time per move at 10 seconds plus 1 a move. Right of the line, a model thinks longer than the increment it earns and the clock drains.
Benchmarks / BulletBench
When every second of thinking time comes off the clock, which models are fast enough to still make good decisions?
Each configuration's median time per move at 10 seconds plus 1 a move. Right of the line, a model thinks longer than the increment it earns and the clock drains.
| # | Model | Lightning 10+1 · 95% range ladder Elo, higher is better | Bullet 60s ladder Elo | Blitz 3+2 ladder Elo | Move time milliseconds | Lost on time % of games | Cost per game US dollars |
|---|---|---|---|---|---|---|---|
| 1 | Gemini 3.5 Flash LiteGoogle |
892 641–1,081 |
780 | 707 | 0.9 s | 0.0% | $0.0070 |
| 1 | Gemini 3.1 Flash Lite (minimal reasoning)Google |
780 589–940 |
912 | – | 0.8 s | 0.0% | $0.0058 |
| 1 | Gemini 3.5 Flash (minimal reasoning)Google |
727 567–890 |
810 | – | 1.2 s | 33.3% | $0.0213 |
| 1 | Seed-2.0-Mini (minimal reasoning)Bytedance |
571 420–710 |
453 | – | 0.3 s | 0.0% | $0.0019 |
| 1 | Mistral Small 4 (no reasoning)Mistral |
571 420–699 |
367 | – | 0.6 s | 0.0% | $0.0041 |
| 1 | GPT-5.4 Mini (no reasoning)OpenAI |
571 412–721 |
197 | – | 0.9 s | 0.0% | $0.0136 |
| 1 | Inkling Small (no reasoning)Thinkingmachines |
571 428–698 |
529 | – | 0.6 s | 0.0% | $0.0077 |
| 1 | Mistral Medium 3.5 (no reasoning)Mistral |
529 361–669 |
572 | – | 0.5 s | 0.0% | $0.0491 |
| 1 | Mistral Medium 3.5Mistral |
529 308–680 |
385 | 555 | 0.5 s | 0.0% | $0.0297 |
| 4 | Gemini 3.8 Flash (low reasoning)Google |
197 0–439 |
1,068 | – | 1.8 s | 83.3% | $0.0041 |
| 5 | GPT-5.4 Nano (no reasoning)OpenAI |
197 0–425 |
33 | – | 1.0 s | 75.0% | $0.0025 |
| 10 | Claude Haiku 4.5Anthropic |
33 0–248 |
445 | 491 | 1.3 s | 83.3% | $0.0095 |
| 10 | Claude Sonnet 5.5 (low reasoning)Anthropic |
33 0–270 |
– | – | 2.1 s | 91.7% | $0.0080 |
| 10 | Gemini 3.6 Flash (minimal reasoning)Google |
33 0–248 |
628 | – | 1.5 s | 91.7% | $0.0050 |
Models share a rank when their ranges overlap. Each model runs at its provider's default reasoning setting and, where the provider offers a faster one, again at its fastest setting, listed separately. Blitz 3+2 was run at default settings only.
These configurations lost their first four games on time, so they were stopped there and are not rated for that clock. Being too slow is the result, not a fault.
Each model's clock after every move of one game. A fast model earns back more than it spends; a slow one runs out within a few moves, whatever the position.
The same at 60 seconds for the whole game, no increment. A 40-move game leaves 1.5 seconds a move.
Share of games lost on time, and moves with no legal reply, by clock.
| Model | On time, Lightning | On time, Bullet | Invalid moves, Lightning | Invalid moves, Bullet |
|---|---|---|---|---|
| Gemini 3.5 Flash LiteGoogle | 0.0% | 0.0% | 0.0% | 0.0% |
| Gemini 3.1 Flash Lite (minimal reasoning)Google | 0.0% | 8.3% | 0.2% | 0.0% |
| Gemini 3.5 Flash (minimal reasoning)Google | 33.3% | 25.0% | 0.6% | 0.5% |
| Seed-2.0-Mini (minimal reasoning)Bytedance | 0.0% | 0.0% | 0.8% | 0.4% |
| Mistral Small 4 (no reasoning)Mistral | 0.0% | 25.0% | 1.0% | 0.6% |
| GPT-5.4 Mini (no reasoning)OpenAI | 0.0% | 50.0% | 0.0% | 0.0% |
| Inkling Small (no reasoning)Thinkingmachines | 0.0% | 0.0% | 5.3% | 7.2% |
| Mistral Medium 3.5 (no reasoning)Mistral | 0.0% | 0.0% | 0.0% | 0.4% |
| Mistral Medium 3.5Mistral | 0.0% | 8.3% | 0.3% | 0.8% |
| Gemini 3.8 Flash (low reasoning)Google | 83.3% | 50.0% | 0.0% | 0.0% |
| GPT-5.4 Nano (no reasoning)OpenAI | 75.0% | 66.7% | 4.0% | 3.1% |
| Claude Haiku 4.5Anthropic | 83.3% | 25.0% | 0.0% | 0.3% |
| Claude Sonnet 5.5 (low reasoning)Anthropic | 91.7% | – | 0.0% | – |
| Gemini 3.6 Flash (minimal reasoning)Google | 91.7% | 41.7% | 0.0% | 0.3% |
The model gets the position, the moves so far, its remaining clock and the legal moves.
It answers with one move. The whole API round trip, including any hidden reasoning, comes off its clock.
Stockfish replies instantly at the model's current ladder level.
Win and the next game is a level up; lose and it is a level down; draw and it stays. The results give a rating with a 95% interval.
Many products need an answer in about a second: routing, triage, autocomplete, live agents. BulletBench asks which models can still think usefully at that speed. Chess gives a hard, objective score, and the clock makes slow thinking a real cost instead of a free extra.
Twelve games per clock give wide intervals: models whose ranges overlap share a rank. Latency depends on the provider's servers on the day. The ladder's Elo anchors come from v1's Stockfish; this edition runs Stockfish 19, so ratings are not comparable with v1's.