Benchmarks / tau2-bench
Reported by tau2-bench
tau2-bench
Share of telecom account and technical support tasks completed within policy, averaged over trials.
- Results dated
- 24 Feb 2026 to 26 May 2026
- Models
- 8
- Unit
- % of tasks
- Licence
- MIT License
Full results
| # | Model | tau2 v1.0.1, telecom: task success (pass^1) % of tasks, higher is better |
|---|---|---|
| 1 | Qwen3.5 397B A17B (thinking reasoning, tau2)Alibaba |
97.8%
|
| 2 | Claude Opus 4.5 (high reasoning, tau2)Anthropic |
92.3%
|
| 3 | gemini-3-flash (high reasoning, tau2)Google |
91.2%
|
| 4 | gemini-3-pro (high reasoning, tau2)Google |
91.0%
|
| 5 | GPT-5.2 (high reasoning, tau2)OpenAI |
89.7%
|
| 6 | GLM 5 (thinking reasoning, tau2)Z.ai |
86.8%
|
| 7 | Claude Sonnet 4.5 (thinking reasoning, tau2)Anthropic |
84.9%
|
| 8 | GPT-5.2 (no reasoning, tau2)OpenAI |
57.2%
|
Results as published by tau2-bench; we do not re-run them.
What it measures
Share of telecom account and technical support tasks completed within policy, averaged over trials.
What it does not measure
Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.
Source
Failures
Failures count against a model: a product with no usable output is a failed listing. They are listed here so you can see why.
Not ranked
- Distyl ButtonAgent: custom agent scaffold, not the standard harness
- RAFT-30B-A3B: custom agent scaffold, not the standard harness
tau2-bench by Sierra Research and contributors, MIT License.