Benchmarks / Microsoft STATE-Bench

Reported by Microsoft STATE-Bench

Microsoft STATE-Bench

Share of tasks completed correctly in all five of five runs.

Results dated
25 May 2026 to 29 May 2026
Models
5
Unit
% of tasks
Licence
MIT License

Full results

STATE-Bench v0.7: consistency (pass^5), % of tasks, higher is better
#ModelSTATE-Bench v0.7: consistency (pass^5)
% of tasks, higher is better
1 GPT-5.4 (high reasoning, statebench)OpenAI
38.0%
2 Claude Opus 4.7 (high reasoning, statebench)Anthropic
33.9%
3 Kimi K2.6 (statebench)Moonshot AI
29.3%
4 GPT-5.4 (default reasoning, statebench)OpenAI
26.2%
5 DeepSeek V4 Pro 0423 (statebench)DeepSeek
25.3%

Results as published by Microsoft STATE-Bench; we do not re-run them.

What it measures

Share of tasks completed correctly in all five of five runs.

What it does not measure

Not the model alone: it runs inside STATE-Bench's agent loop against simulated users, and versions of the benchmark are not comparable.

Failures

Failures count against a model: a product with no usable output is a failed listing. They are listed here so you can see why.

Not ranked

  • GPT 5.1: memory-track entries are agent-learning systems, not models
  • GPT 5.1 + Foundry Memory: memory-track entries are agent-learning systems, not models
  • GPT 5.4: memory-track entries are agent-learning systems, not models
  • GPT 5.4 + Foundry Memory: memory-track entries are agent-learning systems, not models

STATE-Bench by Microsoft and STATE-Bench contributors, MIT License.