Benchmarks / tau2-bench

Reported by tau2-bench

tau2-bench

Share of banking questions answered from a policy knowledge base tasks completed within policy, averaged over trials.

Results dated
5 May 2026 to 3 Aug 2026
Models
21
Unit
% of tasks
Licence
MIT License

Full results

tau2 v1.0.1, banking knowledge, alltools retrieval: task success (pass^1), % of tasks, higher is better
#Modeltau2 v1.0.1, banking knowledge, alltools retrieval: task success (pass^1)
% of tasks, higher is better
1 Qwen3.8 Max (0902) (xhigh reasoning, tau2)Alibaba
55.2%
2 Claude Opus 5 (max reasoning, tau2)Anthropic
48.7%
3 Grok 4.5 (high reasoning, tau2)xAI
47.9%
4 GPT-5.6 Sol (xhigh reasoning, tau2)OpenAI
46.9%
5 GPT-5.5 (xhigh reasoning, tau2)OpenAI
44.6%
6 Muse Spark 1.1 (xhigh reasoning, tau2)Meta
40.5%
7 Claude Opus 4.7 (max reasoning, tau2)Anthropic
40.2%
8 Claude Fable 5 (max reasoning, tau2)Anthropic
39.7%
8 Claude Opus 4.8 (max reasoning, tau2)Anthropic
39.7%
10 GPT-5.4 (xhigh reasoning, tau2)OpenAI
39.4%
11 Kimi K3 (max reasoning, tau2)Moonshot AI
37.1%
11 GLM 5.2 (xhigh reasoning, tau2)Z.ai
37.1%
13 GPT-5.2 (high reasoning, tau2)OpenAI
32.2%
14 Claude Opus 4.6 (max reasoning, tau2)Anthropic
27.3%
15 Gemini 3.1 Pro Preview (high reasoning, tau2)Google
26.0%
16 Inkling (max reasoning, tau2)Thinkingmachines
25.0%
17 Claude Opus 4.5 (high reasoning, tau2)Anthropic
24.7%
18 Grok 4.20 (high reasoning, tau2)xAI
18.0%
19 grok-4-fast (high reasoning, tau2)xAI
15.7%
20 Gemini 2.5 Pro (high reasoning, tau2)Google
13.7%
21 grok-4-1-fast (high reasoning, tau2)xAI
13.1%

Results as published by tau2-bench; we do not re-run them.

What it measures

Share of banking questions answered from a policy knowledge base tasks completed within policy, averaged over trials.

What it does not measure

Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.

Failures

Failures count against a model: a product with no usable output is a failed listing. They are listed here so you can see why.

Not ranked

  • Distyl ButtonAgent: custom agent scaffold, not the standard harness
  • RAFT-30B-A3B: custom agent scaffold, not the standard harness

tau2-bench by Sierra Research and contributors, MIT License.