Benchmarks / tau2-bench

Reported by tau2-bench

tau2-bench

Share of banking questions answered from a policy knowledge base tasks completed within policy in all four trials.

Results dated
27 Feb 2026
Models
2
Unit
% of tasks
Licence
MIT License

Full results

tau2 v1.0.1, banking knowledge, text-emb-3-large retrieval: consistency (pass^4), % of tasks, higher is better
#Modeltau2 v1.0.1, banking knowledge, text-emb-3-large retrieval: consistency (pass^4)
% of tasks, higher is better
1 Qwen3.5 397B A17B (thinking reasoning, tau2)Alibaba
5.2%
2 GLM 5 (thinking reasoning, tau2)Z.ai
3.1%

Results as published by tau2-bench; we do not re-run them.

What it measures

Share of banking questions answered from a policy knowledge base tasks completed within policy in all four trials.

What it does not measure

Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.

Failures

Failures count against a model: a product with no usable output is a failed listing. They are listed here so you can see why.

Not ranked

  • Distyl ButtonAgent: custom agent scaffold, not the standard harness
  • RAFT-30B-A3B: custom agent scaffold, not the standard harness

tau2-bench by Sierra Research and contributors, MIT License.