Models / Kimi K2 0711

Moonshot AI

Kimi K2 0711

17 published results from 3 sources. Each card shows where the number comes from and what it does not measure. Results are never combined into one score.

Provider
Moonshot AI
Sources
3
Our benchmarks
0
Price
Not yet published

Reported by others

1,409
Business, management and finance · rank 108 of 402
Unit
Arena rating, higher is better
Range
1,400 to 1,418
Sample
4529 votes
Configuration
Kimi K2 0711
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,371
Creative writing · rank 119 of 407
Unit
Arena rating, higher is better
Range
1,361 to 1,381
Sample
3371 votes
Configuration
Kimi K2 0711
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,417
Expert prompts · rank 115 of 359
Unit
Arena rating, higher is better
Range
1,401 to 1,432
Sample
1472 votes
Configuration
Kimi K2 0711
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,381
Instruction following · rank 146 of 409
Unit
Arena rating, higher is better
Range
1,373 to 1,389
Sample
6708 votes
Configuration
Kimi K2 0711
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,418
Overall · rank 111 of 409
Unit
Arena rating, higher is better
Range
1,413 to 1,423
Sample
26960 votes
Configuration
Kimi K2 0711
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,375
Writing, literature and language · rank 138 of 408
Unit
Arena rating, higher is better
Range
1,367 to 1,383
Sample
5928 votes
Configuration
Kimi K2 0711
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
29.0%
Memory · rank 23 of 109
Unit
% correct, higher is better
Configuration
kimi-k2
Measured
16 Dec 2025
Not shown
Not long-term personal memory in a product: sessions are BFCL's scripted ones.
50.6%
Multi-turn tasks · rank 16 of 109
Unit
% correct, higher is better
Configuration
kimi-k2
Measured
16 Dec 2025
Not shown
Not open-ended agent work: the tools and tasks are BFCL's simulated APIs.
59.1%
Overall accuracy · rank 11 of 109
Unit
% correct, higher is better
Configuration
kimi-k2
Measured
16 Dec 2025
Not shown
Not a neutral average: the weighting is BFCL's. Not reliability on your own tools: BFCL's functions and queries are a fixed test set, and a correct call is judged by its form, not by what it achieved.
75.0%
Relevance detection · rank 62 of 109
Unit
% correct, higher is better
Configuration
kimi-k2
Measured
16 Dec 2025
Not shown
Not whether the call itself was right; only that one was attempted.
81.6%
Single-turn calls (curated) · rank 69 of 109
Unit
% correct, higher is better
Configuration
kimi-k2
Measured
16 Dec 2025
Not shown
Not reliability on your own tools: BFCL's functions and queries are a fixed test set, and a correct call is judged by its form, not by what it achieved.
78.7%
Single-turn calls (user-contributed) · rank 19 of 109
Unit
% correct, higher is better
Configuration
kimi-k2
Measured
16 Dec 2025
Not shown
Not reliability on your own tools: BFCL's functions and queries are a fixed test set, and a correct call is judged by its form, not by what it achieved.
66.5%
Web search · rank 19 of 109
Unit
% correct, higher is better
Configuration
kimi-k2
Measured
16 Dec 2025
Not shown
Not general research quality: questions have short, checkable answers.
Reported by UGI Leaderboard
22.0%
Requested-length error · rank 202 of 370
Unit
% off the requested word count, lower is better
Configuration
Kimi K2 0711
Measured
7 Sep 2025
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
0.29
Style adherence · rank 351 of 370
Unit
score from 0 to 1, higher is better
Configuration
Kimi K2 0711
Measured
7 Sep 2025
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
52.5
Writing score · rank 143 of 370
Unit
score out of 100, higher is better
Configuration
Kimi K2 0711
Measured
7 Sep 2025
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.

Compare Kimi K2 0711 with