Meta’s Muse Spark 1.3 and Google’s Gemini 3.8 Flash launched within hours of each other on September 2–3, 2026, both pitched as cheap, high-capability reasoning and coding models. Three independent testers ran their own benchmarks on the exact same two models that same week β€” and got three different winners. There’s no single “which one wins” answer here, because the disagreement isn’t between vendor marketing and independent testing, it’s between independent testers themselves.

“Which wins” assumes one number should settle it β€” it can’t, because the testers don’t agree with each other

Normally a benchmark honesty-beat is about a vendor’s launch-day number versus independent verification. This launch is a step further: three separate independent evaluators, using three different test sets, reached three different verdicts on the same pair of models in the same week.

Independent testerMethodResult
MindStudio8-task coding, 3D, and agentic benchmarkGemini 3.8 Flash wins decisively: 65/80 (81.25%) vs Muse Spark 1.3’s 57/80 (71.25%)
Kingy.ai232-call direct test across coding and repository-repair tasksMuse Spark 1.3 wins decisively in the completed available-case analysis: 112/112 scored calls vs Gemini’s 91/112 β€” though Kingy’s own write-up flags this isn’t the full preregistered 240-call set, since some Gemini long-context calls hit an API quota
Artificial AnalysisIntelligence Index compositeRoughly tied: Muse Spark 1.3 scores 61 at $0.55 per task, Gemini 3.8 Flash scores 59 at $0.58 per task

Read that table as three genuinely different pictures of the same launch, not as a 2-1 vote for either model. Different task mixes, different harnesses, and different completion rates on the underlying test runs produce different rankings β€” which is the same lesson our piece on coding-agent harnesses makes in more general form, just now demonstrated by three outside evaluators disagreeing with each other rather than with a vendor.

“Biggest jump we’ve made so far” hides a real regression

Mark Zuckerberg introduced Muse Spark 1.3 on social media as delivering “frontier performance almost too cheap to meter” and “the biggest jump we’ve made so far on coding and agentic work.” MindStudio’s own testing complicates that framing directly: Muse Spark 1.3 actually regressed from its predecessor, Muse Spark 1.2, on the same 8-task benchmark β€” dropping from 76.25% to 71.25% overall. The regression isn’t uniform: 1.3 set a new record on a 3D wristwatch task and improved on an elevator simulation, but fell hard on SVG generation and a 3D folding table β€” the exact tasks where 1.2 had excelled. The net effect is a trade, not a clean decline or a clean improvement, and it’s a genuine regression sitting underneath a launch framed as pure uplift β€” the same pattern GPT-6 Astra’s own GDPval-AA v2 regression showed the same week, on a different lab’s model.

A single benchmark number can hide an entirely different underlying system

Meta’s own materials cite a 75.4% score for Muse Spark 1.3 on the DeepSWE benchmark, ahead of Gemini 3.8 Flash’s reported 74%. Kingy.ai’s own analysis makes a point worth repeating plainly: Meta’s evaluation used a specific internal agent framework or benchmark-provider harness for some tests and named agent products or a mini-swe setup for others, while the comparison numbers for other models came from a mix of vendor self-reports and independent leaderboards. Subtracting Meta’s 75.4% from a third party’s 74% and declaring a 1.4-point win treats two numbers produced by different underlying systems as if they were the same measurement. This is the exact confound covered at length in our harness piece and again in our Qwen vs Claude Opus 5 comparison, and it’s showing up in essentially every major model launch this year β€” treat any cross-lab benchmark subtraction with the same skepticism regardless of which two models are being compared.

The best numbers aren’t necessarily the numbers you can access

Meta’s most favorable comparisons for Muse Spark 1.3 reportedly come from a configuration not yet broadly available to developers β€” the same access-gating pattern seen with GPT-6 Astra’s cyber-capable configuration staying inside OpenAI’s Daybreak program at launch. Before choosing a model based on a headline comparison, check which access tier that specific number was measured on, and confirm your own account tier actually gets it β€” a pattern worth watching for on every major release from here forward, not just this one.

Price comparisons need the same caution as the benchmarks

Gemini 3.8 Flash’s list price β€” $0.75 per million input tokens, $3.75 per million output tokens β€” looks meaningfully cheaper than Muse Spark on a per-token basis. But Artificial Analysis’s own cost-per-task figures, which account for how many tokens each model actually uses to complete a task, put the two within three cents of each other ($0.55 for Muse Spark 1.3, $0.58 for Gemini 3.8 Flash). A model that uses more tokens per task can erase a lower list price entirely β€” the same caution our guide to free vs. paid AI tools applies to sticker-price comparisons generally, and the same one our Qwen vs Claude Opus 5 piece raised about reasoning tokens quietly inflating a real bill past the advertised rate.

What this actually means if you’re choosing a model to build on

Given three independent testers disagreeing with each other this sharply, no public benchmark table β€” including the one in this article β€” should be the deciding factor for your specific workload. Test both models directly against a representative sample of your own tasks before committing either into a production pipeline, the same recommendation our harness piece makes for choosing between coding agents generally. If your work runs primarily through a specific tool rather than a raw API, check how that tool exposes these models first β€” our Cursor vs Claude Code vs Copilot comparison covers how much the surrounding tool changes the practical outcome independent of the underlying model. And if part of your evaluation includes other recent budget-tier launches, our GLM-5.3 review and our Grok 4.6 review cover two more models chasing the same cheap-frontier-performance positioning this week.

Who should sit this one out

If you’re already running a workflow on a different model β€” Claude, GPT, GLM, Qwen, or an earlier Gemini or Muse Spark version β€” and it’s working, this week’s launch pair isn’t a strong enough or stable enough signal to switch on. The three-way disagreement above means the ground truth here is still being worked out by the people whose job is measuring it; there’s no cost to waiting for that to settle before spending engineering time on a migration.

Does Muse Spark 1.3 actually beat Gemini 3.8 Flash?

It depends entirely on whose independent test you trust β€” MindStudio’s own benchmark found Gemini winning decisively, Kingy.ai’s own test found Muse Spark winning decisively, and Artificial Analysis’s Intelligence Index found them roughly tied. There is no single settled answer this week.

Why do different benchmark tests give completely different winners for the same two models?

Each independent tester used a different task mix, a different underlying harness or agent framework, and in at least one case an incomplete test run due to an API quota limit. Benchmarks measure the combination of a model and its test setup, not the model in isolation, which is why three genuinely independent tests can disagree this sharply.

Did Muse Spark 1.3 actually get worse at anything compared to version 1.2?

Yes, according to MindStudio’s testing. It regressed on SVG generation and a 3D folding table task, dropping its overall score on their 8-task benchmark from 76.25% to 71.25%, even while improving on other tasks like a 3D wristwatch build and an elevator simulation.

Is Gemini 3.8 Flash or Muse Spark 1.3 cheaper to run?

Gemini’s list price per token is lower, but Artificial Analysis’s cost-per-completed-task figures β€” which account for actual token usage β€” put the two within three cents of each other. Compare cost per task on your own workload rather than relying on list price alone.

Can I access Muse Spark 1.3’s best-performing configuration right now?

Not necessarily. Reporting on the launch indicates Meta’s most favorable comparisons come from a configuration not yet broadly available to developers, similar to how GPT-6 Astra’s top-scoring configuration stayed limited to OpenAI’s gated Daybreak program at launch. Confirm what your specific account tier actually provides before assuming a headline number applies to you.

Which one should I actually use for my own coding or automation work?

Test both on a representative sample of your own tasks rather than choosing from a public benchmark table. Given how much the three independent tests disagree with each other this week, no generic comparison β€” including this one β€” reliably predicts how either model will perform on your specific workload.

Shurah is the founder of AI Tools Daily, tracking pricing, licensing and policy changes across AI tools so readers can make decisions without wading through marketing claims themselves.