Models / Muse Spark 1.3

Meta

Muse Spark 1.3

34 published results from 4 sources. Each card shows where the number comes from and what it does not measure. Results are never combined into one score.

Provider
Meta
Sources
4
Our benchmarks
3
Price
Not yet published

Measured by Spring Prompt

Measured by Spring Prompt · BulletBench
0
Ladder Elo · rank 12 of 16
Unit
ladder Elo, higher is better
Configuration
Muse Spark 1.3, provider default reasoning
Measured
30 Sep 2026
Not shown
Not a business skill: chess against an engine on a real clock, used here as a speed-under-pressure signal.
Measured by Spring Prompt · BulletBench
20.7 s
Median move time · rank 16 of 16
Unit
milliseconds, lower is better
Configuration
Muse Spark 1.3, provider default reasoning
Measured
30 Sep 2026
Not shown
Not general response speed: replies are one short move.
Measured by Spring Prompt · CatalogBench
97.7%
Content quality · rank 5 of 18
Unit
% of checks, higher is better
Configuration
Muse Spark 1.3, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
0.0%
Failed outputs · rank 1 of 18
Unit
% of products, lower is better
Configuration
Muse Spark 1.3, provider default reasoning
Measured
30 Sep 2026
Not shown
Not a provider outage: provider errors are retried before a product counts as failed.
Measured by Spring Prompt · CatalogBench
14.3%
Publish-ready listings · rank 7 of 18
Unit
% of products, higher is better
Configuration
Muse Spark 1.3, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
7.1%
Reliably publish-ready · rank 7 of 18
Unit
% of products, higher is better
Configuration
Muse Spark 1.3, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
3.44
Unsupported claims · rank 8 of 18
Unit
claims per product, lower is better
Configuration
Muse Spark 1.3, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
100.0%
Channel compliance · rank 1 of 18
Unit
% of products, higher is better
Configuration
Muse Spark 1.3, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
100.0%
Conflicts caught · rank 1 of 18
Unit
% of conflicts, higher is better
Configuration
Muse Spark 1.3, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
97.2%
Content quality · rank 5 of 18
Unit
% of checks, higher is better
Configuration
Muse Spark 1.3, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
$0.0164
Cost per product · rank 10 of 18
Unit
US dollars, lower is better
Configuration
Muse Spark 1.3, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your cost: prices are those charged through OpenRouter on the run date.
Measured by Spring Prompt · CatalogBench
96.7%
Decision accuracy · rank 5 of 18
Unit
% of decisions, higher is better
Configuration
Muse Spark 1.3, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
0.0%
Failed outputs · rank 1 of 18
Unit
% of products, lower is better
Configuration
Muse Spark 1.3, provider default reasoning
Measured
30 Sep 2026
Not shown
Not a provider outage: provider errors are retried before a product counts as failed.
Measured by Spring Prompt · CatalogBench
92.9%
Field accuracy · rank 10 of 18
Unit
% of missing fields, higher is better
Configuration
Muse Spark 1.3, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
0.7%
Invented values · rank 5 of 18
Unit
% of filled values, lower is better
Configuration
Muse Spark 1.3, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
57.7%
Publish-ready listings · rank 8 of 18
Unit
% of products, higher is better
Configuration
Muse Spark 1.3, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
37.5%
Reliably publish-ready · rank 7 of 18
Unit
% of products, higher is better
Configuration
Muse Spark 1.3, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
0.73
Unsupported claims · rank 6 of 18
Unit
claims per product, lower is better
Configuration
Muse Spark 1.3, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · ROASBench
$0.18
Cost of a run · rank 7 of 17
Unit
US dollars, lower is better
Configuration
Muse Spark 1.3, provider default reasoning
Measured
29 Sep 2026
Not shown
Not your cost: prices are those charged through OpenRouter on the run date.
Measured by Spring Prompt · ROASBench
0
Months over budget · rank 1 of 17
Unit
months of 12, lower is better
Configuration
Muse Spark 1.3, provider default reasoning
Measured
29 Sep 2026
Not shown
Not real ad performance: the market is a deterministic simulation of one invented skincare brand.
Measured by Spring Prompt · ROASBench
34.1
Overall score · rank 11 of 17
Unit
score out of 100, higher is better
Configuration
Muse Spark 1.3, provider default reasoning
Measured
29 Sep 2026
Not shown
Not real ad performance: the market is a deterministic simulation of one invented skincare brand.
Measured by Spring Prompt · ROASBench
£0.29
Return on ad spend · rank 10 of 17
Unit
profit per £1 spent, higher is better
Configuration
Muse Spark 1.3, provider default reasoning
Measured
29 Sep 2026
Not shown
Not real ad performance: the market is a deterministic simulation of one invented skincare brand.

Reported by others

0.08
Confirmed task success · rank 4 of 46
Unit
IPS effect estimate, higher is better
Range
0.07 to 0.10
Sample
43358 observations
Configuration
Muse Spark 1.3
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
0.04
Praise over complaint · rank 10 of 46
Unit
IPS effect estimate, higher is better
Range
0.02 to 0.06
Sample
15832 observations
Configuration
Muse Spark 1.3
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
0.02
Steerability · rank 11 of 46
Unit
IPS effect estimate, higher is better
Range
0.00 to 0.03
Sample
45532 observations
Configuration
Muse Spark 1.3
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
0.00
Tool grounding · rank 1 of 46
Unit
IPS effect estimate, higher is better
Range
0.00 to 0.00
Sample
5211195 observations
Configuration
Muse Spark 1.3
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
1,468
Overall · rank 6 of 44
Unit
Arena rating, higher is better
Range
1,450 to 1,486
Sample
1006 votes
Configuration
Muse Spark 1.3 (max)
Measured
13 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,479
Overall · rank 4 of 177
Unit
Arena rating, higher is better
Range
1,473 to 1,485
Sample
9672 votes
Configuration
Muse Spark 1.3 (max)
Measured
25 Sep 2026
Not shown
Not an error rate: claims that cannot be checked on the web are skipped, and preference still carries most of the weight.
1,506
Business, management and finance · rank 1 of 402
Unit
Arena rating, higher is better
Range
1,492 to 1,519
Sample
1912 votes
Configuration
Muse Spark 1.3 (max)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,450
Creative writing · rank 13 of 407
Unit
Arena rating, higher is better
Range
1,436 to 1,464
Sample
2035 votes
Configuration
Muse Spark 1.3 (max)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,527
Expert prompts · rank 1 of 359
Unit
Arena rating, higher is better
Range
1,510 to 1,544
Sample
1149 votes
Configuration
Muse Spark 1.3 (max)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,486
Instruction following · rank 4 of 409
Unit
Arena rating, higher is better
Range
1,476 to 1,496
Sample
3726 votes
Configuration
Muse Spark 1.3 (max)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,494
Overall · rank 2 of 409
Unit
Arena rating, higher is better
Range
1,488 to 1,501
Sample
10036 votes
Configuration
Muse Spark 1.3 (max)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,466
Writing, literature and language · rank 9 of 408
Unit
Arena rating, higher is better
Range
1,454 to 1,478
Sample
2655 votes
Configuration
Muse Spark 1.3 (max)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.

Compare Muse Spark 1.3 with