Models / Grok 4.7

xAI

Grok 4.7

33 published results from 4 sources. Each card shows where the number comes from and what it does not measure. Results are never combined into one score.

Provider
xAI
Sources
4
Our benchmarks
3
Price
Not yet published

Measured by Spring Prompt

Measured by Spring Prompt · BulletBench
147
Ladder Elo · rank 6 of 16
Unit
ladder Elo, higher is better
Range
0 to 387
Sample
8 games
Configuration
Grok 4.7, provider default reasoning
Measured
29 Sep 2026
Not shown
Not a business skill: chess against an engine on a real clock, used here as a speed-under-pressure signal.
Measured by Spring Prompt · BulletBench
10.1 s
Median move time · rank 15 of 16
Unit
milliseconds, lower is better
Configuration
Grok 4.7, provider default reasoning
Measured
29 Sep 2026
Not shown
Not general response speed: replies are one short move.
Measured by Spring Prompt · CatalogBench
97.7%
Content quality · rank 4 of 18
Unit
% of checks, higher is better
Configuration
Grok 4.7, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
0.0%
Failed outputs · rank 1 of 18
Unit
% of products, lower is better
Configuration
Grok 4.7, provider default reasoning
Measured
30 Sep 2026
Not shown
Not a provider outage: provider errors are retried before a product counts as failed.
Measured by Spring Prompt · CatalogBench
38.1%
Publish-ready listings · rank 5 of 18
Unit
% of products, higher is better
Configuration
Grok 4.7, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
10.7%
Reliably publish-ready · rank 5 of 18
Unit
% of products, higher is better
Configuration
Grok 4.7, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
1.58
Unsupported claims · rank 5 of 18
Unit
claims per product, lower is better
Configuration
Grok 4.7, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
100.0%
Channel compliance · rank 1 of 18
Unit
% of products, higher is better
Configuration
Grok 4.7, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
100.0%
Conflicts caught · rank 1 of 18
Unit
% of conflicts, higher is better
Configuration
Grok 4.7, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
97.3%
Content quality · rank 4 of 18
Unit
% of checks, higher is better
Configuration
Grok 4.7, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
$0.0237
Cost per product · rank 12 of 18
Unit
US dollars, lower is better
Configuration
Grok 4.7, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your cost: prices are those charged through OpenRouter on the run date.
Measured by Spring Prompt · CatalogBench
96.4%
Decision accuracy · rank 7 of 18
Unit
% of decisions, higher is better
Configuration
Grok 4.7, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
0.0%
Failed outputs · rank 1 of 18
Unit
% of products, lower is better
Configuration
Grok 4.7, provider default reasoning
Measured
30 Sep 2026
Not shown
Not a provider outage: provider errors are retried before a product counts as failed.
Measured by Spring Prompt · CatalogBench
92.9%
Field accuracy · rank 10 of 18
Unit
% of missing fields, higher is better
Configuration
Grok 4.7, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
0.5%
Invented values · rank 3 of 18
Unit
% of filled values, lower is better
Configuration
Grok 4.7, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
62.5%
Publish-ready listings · rank 5 of 18
Unit
% of products, higher is better
Configuration
Grok 4.7, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
44.6%
Reliably publish-ready · rank 5 of 18
Unit
% of products, higher is better
Configuration
Grok 4.7, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
0.80
Unsupported claims · rank 7 of 18
Unit
claims per product, lower is better
Configuration
Grok 4.7, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · ROASBench
$0.42
Cost of a run · rank 11 of 17
Unit
US dollars, lower is better
Configuration
Grok 4.7, provider default reasoning
Measured
29 Sep 2026
Not shown
Not your cost: prices are those charged through OpenRouter on the run date.
Measured by Spring Prompt · ROASBench
0
Months over budget · rank 1 of 17
Unit
months of 12, lower is better
Configuration
Grok 4.7, provider default reasoning
Measured
29 Sep 2026
Not shown
Not real ad performance: the market is a deterministic simulation of one invented skincare brand.
Measured by Spring Prompt · ROASBench
36.6
Overall score · rank 10 of 17
Unit
score out of 100, higher is better
Configuration
Grok 4.7, provider default reasoning
Measured
29 Sep 2026
Not shown
Not real ad performance: the market is a deterministic simulation of one invented skincare brand.
Measured by Spring Prompt · ROASBench
£0.23
Return on ad spend · rank 12 of 17
Unit
profit per £1 spent, higher is better
Configuration
Grok 4.7, provider default reasoning
Measured
29 Sep 2026
Not shown
Not real ad performance: the market is a deterministic simulation of one invented skincare brand.

Reported by others

0.09
Confirmed task success · rank 3 of 46
Unit
IPS effect estimate, higher is better
Range
0.06 to 0.12
Sample
5222 observations
Configuration
Grok 4.7
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
0.04
Praise over complaint · rank 9 of 46
Unit
IPS effect estimate, higher is better
Range
-0.01 to 0.10
Sample
1702 observations
Configuration
Grok 4.7
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
0.00
Steerability · rank 10 of 46
Unit
IPS effect estimate, higher is better
Range
-0.03 to 0.04
Sample
5569 observations
Configuration
Grok 4.7
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
0.00
Tool grounding · rank 1 of 46
Unit
IPS effect estimate, higher is better
Range
0.00 to 0.00
Sample
1310543 observations
Configuration
Grok 4.7
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
1,420
Overall · rank 105 of 177
Unit
Arena rating, higher is better
Range
1,412 to 1,429
Sample
3775 votes
Configuration
Grok 4.7 (xhigh)
Measured
25 Sep 2026
Not shown
Not an error rate: claims that cannot be checked on the web are skipped, and preference still carries most of the weight.
1,424
Business, management and finance · rank 65 of 402
Unit
Arena rating, higher is better
Range
1,403 to 1,445
Sample
745 votes
Configuration
Grok 4.7 (xhigh)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,427
Creative writing · rank 27 of 407
Unit
Arena rating, higher is better
Range
1,406 to 1,448
Sample
884 votes
Configuration
Grok 4.7 (xhigh)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,454
Expert prompts · rank 47 of 359
Unit
Arena rating, higher is better
Range
1,426 to 1,482
Sample
430 votes
Configuration
Grok 4.7 (xhigh)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,437
Instruction following · rank 47 of 409
Unit
Arena rating, higher is better
Range
1,421 to 1,453
Sample
1382 votes
Configuration
Grok 4.7 (xhigh)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,439
Overall · rank 71 of 409
Unit
Arena rating, higher is better
Range
1,430 to 1,449
Sample
4114 votes
Configuration
Grok 4.7 (xhigh)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,426
Writing, literature and language · rank 48 of 408
Unit
Arena rating, higher is better
Range
1,407 to 1,444
Sample
1103 votes
Configuration
Grok 4.7 (xhigh)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.

Compare Grok 4.7 with