Models / DeepSeek V4.1 Flash

DeepSeek

DeepSeek V4.1 Flash

27 published results from 2 sources. Each card shows where the number comes from and what it does not measure. Results are never combined into one score.

Provider
DeepSeek
Sources
2
Our benchmarks
1
Price
Not yet published

Measured by Spring Prompt

Measured by Spring Prompt · CatalogBench
95.4%
Content quality · rank 13 of 18
Unit
% of checks, higher is better
Configuration
DeepSeek V4.1 Flash, provider default reasoning
Measured
30 Sep 2026
Failures
invalid JSON: 1 of 168 attempts
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
0.6%
Failed outputs · rank 15 of 18
Unit
% of products, lower is better
Configuration
DeepSeek V4.1 Flash, provider default reasoning
Measured
30 Sep 2026
Failures
invalid JSON: 1 of 168 attempts
Not shown
Not a provider outage: provider errors are retried before a product counts as failed.
Measured by Spring Prompt · CatalogBench
24.4%
Publish-ready listings · rank 6 of 18
Unit
% of products, higher is better
Configuration
DeepSeek V4.1 Flash, provider default reasoning
Measured
30 Sep 2026
Failures
invalid JSON: 1 of 168 attempts
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
10.7%
Reliably publish-ready · rank 5 of 18
Unit
% of products, higher is better
Configuration
DeepSeek V4.1 Flash, provider default reasoning
Measured
30 Sep 2026
Failures
invalid JSON: 1 of 168 attempts
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
2.79
Unsupported claims · rank 7 of 18
Unit
claims per product, lower is better
Configuration
DeepSeek V4.1 Flash, provider default reasoning
Measured
30 Sep 2026
Failures
invalid JSON: 1 of 168 attempts
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
99.4%
Channel compliance · rank 9 of 18
Unit
% of products, higher is better
Configuration
DeepSeek V4.1 Flash, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
100.0%
Conflicts caught · rank 1 of 18
Unit
% of conflicts, higher is better
Configuration
DeepSeek V4.1 Flash, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
96.3%
Content quality · rank 11 of 18
Unit
% of checks, higher is better
Configuration
DeepSeek V4.1 Flash, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
$0.0061
Cost per product · rank 3 of 18
Unit
US dollars, lower is better
Configuration
DeepSeek V4.1 Flash, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your cost: prices are those charged through OpenRouter on the run date.
Measured by Spring Prompt · CatalogBench
96.2%
Decision accuracy · rank 8 of 18
Unit
% of decisions, higher is better
Configuration
DeepSeek V4.1 Flash, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
0.0%
Failed outputs · rank 1 of 18
Unit
% of products, lower is better
Configuration
DeepSeek V4.1 Flash, provider default reasoning
Measured
30 Sep 2026
Not shown
Not a provider outage: provider errors are retried before a product counts as failed.
Measured by Spring Prompt · CatalogBench
94.4%
Field accuracy · rank 3 of 18
Unit
% of missing fields, higher is better
Configuration
DeepSeek V4.1 Flash, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
3.1%
Invented values · rank 15 of 18
Unit
% of filled values, lower is better
Configuration
DeepSeek V4.1 Flash, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
58.3%
Publish-ready listings · rank 6 of 18
Unit
% of products, higher is better
Configuration
DeepSeek V4.1 Flash, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
35.7%
Reliably publish-ready · rank 8 of 18
Unit
% of products, higher is better
Configuration
DeepSeek V4.1 Flash, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
0.72
Unsupported claims · rank 5 of 18
Unit
claims per product, lower is better
Configuration
DeepSeek V4.1 Flash, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.

Reported by others

0.12
Confirmed task success · rank 2 of 46
Unit
IPS effect estimate, higher is better
Range
0.11 to 0.13
Sample
48064 observations
Configuration
DeepSeek V4.1 Flash
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
0.03
Praise over complaint · rank 11 of 46
Unit
IPS effect estimate, higher is better
Range
0.01 to 0.05
Sample
17188 observations
Configuration
DeepSeek V4.1 Flash
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
-0.03
Steerability · rank 24 of 46
Unit
IPS effect estimate, higher is better
Range
-0.05 to -0.02
Sample
50671 observations
Configuration
DeepSeek V4.1 Flash
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
0.00
Tool grounding · rank 25 of 46
Unit
IPS effect estimate, higher is better
Range
0.00 to 0.00
Sample
12155899 observations
Configuration
DeepSeek V4.1 Flash
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
1,471
Overall · rank 10 of 177
Unit
Arena rating, higher is better
Range
1,464 to 1,478
Sample
6378 votes
Configuration
DeepSeek V4.1 Flash (max)
Measured
25 Sep 2026
Not shown
Not an error rate: claims that cannot be checked on the web are skipped, and preference still carries most of the weight.
1,481
Business, management and finance · rank 2 of 402
Unit
Arena rating, higher is better
Range
1,464 to 1,498
Sample
1245 votes
Configuration
DeepSeek V4.1 Flash (max)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,440
Creative writing · rank 20 of 407
Unit
Arena rating, higher is better
Range
1,423 to 1,457
Sample
1297 votes
Configuration
DeepSeek V4.1 Flash (max)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,518
Expert prompts · rank 3 of 359
Unit
Arena rating, higher is better
Range
1,497 to 1,538
Sample
775 votes
Configuration
DeepSeek V4.1 Flash (max)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,478
Instruction following · rank 7 of 409
Unit
Arena rating, higher is better
Range
1,466 to 1,490
Sample
2369 votes
Configuration
DeepSeek V4.1 Flash (max)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,477
Overall · rank 13 of 409
Unit
Arena rating, higher is better
Range
1,469 to 1,484
Sample
6630 votes
Configuration
DeepSeek V4.1 Flash (max)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,455
Writing, literature and language · rank 15 of 408
Unit
Arena rating, higher is better
Range
1,440 to 1,470
Sample
1679 votes
Configuration
DeepSeek V4.1 Flash (max)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.

Compare DeepSeek V4.1 Flash with