Benchmarks / tau2-bench
Reported by tau2-bench
tau2-bench
Share of banking questions answered from a policy knowledge base tasks completed within policy, averaged over trials.
- Results dated
- 5 May 2026 to 3 Aug 2026
- Models
- 21
- Unit
- % of tasks
- Licence
- MIT License
Full results
| # | Model | tau2 v1.0.1, banking knowledge, alltools retrieval: task success (pass^1) % of tasks, higher is better |
|---|---|---|
| 1 | Qwen3.8 Max (0902) (xhigh reasoning, tau2)Alibaba |
55.2%
|
| 2 | Claude Opus 5 (max reasoning, tau2)Anthropic |
48.7%
|
| 3 | Grok 4.5 (high reasoning, tau2)xAI |
47.9%
|
| 4 | GPT-5.6 Sol (xhigh reasoning, tau2)OpenAI |
46.9%
|
| 5 | GPT-5.5 (xhigh reasoning, tau2)OpenAI |
44.6%
|
| 6 | Muse Spark 1.1 (xhigh reasoning, tau2)Meta |
40.5%
|
| 7 | Claude Opus 4.7 (max reasoning, tau2)Anthropic |
40.2%
|
| 8 | Claude Fable 5 (max reasoning, tau2)Anthropic |
39.7%
|
| 8 | Claude Opus 4.8 (max reasoning, tau2)Anthropic |
39.7%
|
| 10 | GPT-5.4 (xhigh reasoning, tau2)OpenAI |
39.4%
|
| 11 | Kimi K3 (max reasoning, tau2)Moonshot AI |
37.1%
|
| 11 | GLM 5.2 (xhigh reasoning, tau2)Z.ai |
37.1%
|
| 13 | GPT-5.2 (high reasoning, tau2)OpenAI |
32.2%
|
| 14 | Claude Opus 4.6 (max reasoning, tau2)Anthropic |
27.3%
|
| 15 | Gemini 3.1 Pro Preview (high reasoning, tau2)Google |
26.0%
|
| 16 | Inkling (max reasoning, tau2)Thinkingmachines |
25.0%
|
| 17 | Claude Opus 4.5 (high reasoning, tau2)Anthropic |
24.7%
|
| 18 | Grok 4.20 (high reasoning, tau2)xAI |
18.0%
|
| 19 | grok-4-fast (high reasoning, tau2)xAI |
15.7%
|
| 20 | Gemini 2.5 Pro (high reasoning, tau2)Google |
13.7%
|
| 21 | grok-4-1-fast (high reasoning, tau2)xAI |
13.1%
|
Results as published by tau2-bench; we do not re-run them.
What it measures
Share of banking questions answered from a policy knowledge base tasks completed within policy, averaged over trials.
What it does not measure
Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.
Source
Failures
Failures count against a model: a product with no usable output is a failed listing. They are listed here so you can see why.
Not ranked
- Distyl ButtonAgent: custom agent scaffold, not the standard harness
- RAFT-30B-A3B: custom agent scaffold, not the standard harness
tau2-bench by Sierra Research and contributors, MIT License.