Models / Grok 4.20

xAI

Grok 4.20

2 published results from 1 source. Each card shows where the number comes from and what it does not measure. Results are never combined into one score.

Provider
xAI
Sources
1
Our benchmarks
0
Price
Not yet published

Reported by others

Reported by tau2-bench
8.2%
Consistency (pass^4) · rank 18 of 21
Unit
% of tasks, higher is better
Configuration
Grok 4.20 (high reasoning, tau2)
Measured
5 May 2026
Not shown
Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.
Reported by tau2-bench
18.0%
Task success (pass^1) · rank 18 of 21
Unit
% of tasks, higher is better
Configuration
Grok 4.20 (high reasoning, tau2)
Measured
5 May 2026
Not shown
Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.