Models / Inkling

Thinkingmachines

Inkling

13 published results from 2 sources. Each card shows where the number comes from and what it does not measure. Results are never combined into one score.

Provider
Thinkingmachines
Sources
2
Our benchmarks
0
Price
Not yet published

Reported by others

-0.22
Confirmed task success · rank 43 of 46
Unit
IPS effect estimate, higher is better
Range
-0.25 to -0.19
Sample
24482 observations
Configuration
Inkling
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
-0.22
Praise over complaint · rank 44 of 46
Unit
IPS effect estimate, higher is better
Range
-0.25 to -0.20
Sample
9597 observations
Configuration
Inkling
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
-0.12
Steerability · rank 42 of 46
Unit
IPS effect estimate, higher is better
Range
-0.14 to -0.10
Sample
31566 observations
Configuration
Inkling
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
-0.00
Tool grounding · rank 29 of 46
Unit
IPS effect estimate, higher is better
Range
-0.00 to 0.00
Sample
1538933 observations
Configuration
Inkling
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
1,452
Overall · rank 46 of 177
Unit
Arena rating, higher is better
Range
1,448 to 1,455
Sample
32469 votes
Configuration
Inkling
Measured
25 Sep 2026
Not shown
Not an error rate: claims that cannot be checked on the web are skipped, and preference still carries most of the weight.
1,444
Business, management and finance · rank 59 of 402
Unit
Arena rating, higher is better
Range
1,436 to 1,452
Sample
6290 votes
Configuration
Inkling
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,384
Creative writing · rank 99 of 407
Unit
Arena rating, higher is better
Range
1,376 to 1,392
Sample
6718 votes
Configuration
Inkling
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,474
Expert prompts · rank 42 of 359
Unit
Arena rating, higher is better
Range
1,464 to 1,484
Sample
3981 votes
Configuration
Inkling
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,427
Instruction following · rank 76 of 409
Unit
Arena rating, higher is better
Range
1,420 to 1,433
Sample
11891 votes
Configuration
Inkling
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,442
Overall · rank 73 of 409
Unit
Arena rating, higher is better
Range
1,438 to 1,447
Sample
32778 votes
Configuration
Inkling
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,403
Writing, literature and language · rank 91 of 408
Unit
Arena rating, higher is better
Range
1,396 to 1,411
Sample
8708 votes
Configuration
Inkling
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
Reported by tau2-bench
11.3%
Consistency (pass^4) · rank 14 of 21
Unit
% of tasks, higher is better
Configuration
inkling
Measured
24 Jul 2026
Not shown
Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.
Reported by tau2-bench
25.0%
Task success (pass^1) · rank 16 of 21
Unit
% of tasks, higher is better
Configuration
inkling
Measured
24 Jul 2026
Not shown
Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.

Compare Inkling with