This week has marked a significant milestone for Meta's artificial intelligence initiatives, with the company transitioning its Muse Code coding agent from beta status and introducing three new subscription tiers. On Wednesday evening, Meta unveiled Muse Spark 1.3, the newest iteration of its economical reasoning model, asserting that it has achieved unprecedented performance gains in coding and agentic capabilities.

Accessible via Muse Code and the Meta Model API, Muse Spark 1.3 represents the latest evolution of the reasoning model Meta debuted in April, arriving within weeks of Muse Spark 1.2's introduction. Meta CEO Mark Zuckerberg promoted the release on social media, characterizing it as delivering "frontier performance almost too cheap to meter."

This is the biggest jump we've made so far on coding and agentic work.

Mark Zuckerberg, Meta CEO, on X

Open-weight release 'coming soon'

Zuckerberg announced that open-weight versions of Muse Spark will arrive "coming soon," potentially enabling developers to download and execute the model on their own systems rather than relying on Meta's cloud-hosted API.

The scope of this openness remains uncertain, as Meta has not yet disclosed the licensing framework that will govern the release. Meta's prior open-weight offerings have included different usage constraints, and these forthcoming terms will dictate the extent to which Spark can be altered, shared or deployed independently.

Zuckerberg also previewed Meta's anticipated next-generation model, referred to internally as Watermelon, accompanied by a watermelon emoji. According to a Business Insider report from July, this model represents a substantially larger architecture than Spark and remained under development at that time. No official timeline has been provided for Watermelon's release, though Zuckerberg's messaging suggests it may arrive relatively soon.

https://x.com/finkd/status/2095232032896946311?ref_src=twsrc%5Etfw

Regarding Spark 1.3 itself, Meta has made substantial performance assertions. Zuckerberg shared a benchmark comparison positioning the model against OpenAI's GPT-5.6 Sol and Anthropic's Claude Opus 5 across coding, agentic, computer-use and extended-context evaluations.

Key performance metrics include a 75.4% score on the DeepSWE coding benchmark and 98.1% on the 512K-1M variant of the MRCR long-context test.

These figures originate from Meta's own testing methodology. The company generated Spark 1.3 results through the Meta Model API, while comparison data comes from a combination of Meta's evaluations, public leaderboards and figures disclosed by competing model developers. Meta characterizes its assessment of third-party models as "best-effort," indicating the comparison should not be interpreted as a standardized independent evaluation conducted under uniform conditions.

Muse Spark 1.3: Benchmarked
Muse Spark 1.3: Benchmarked

An important qualifier: Meta evaluated Spark 1.3 using its new "max" reasoning configuration for the primary comparisons, whereas the highest reasoning tier currently available to most developers is "xhigh." The max setting remains in restricted preview while Meta conducts additional safety assessments.

"Gloves are off": How Spark 1.3 stacks up

Independent evaluation by Artificial Analysis offers third-party perspective on the models. The San Francisco-based benchmarking specialist assigned Muse Spark 1.3 xhigh a score of 61 on its Intelligence Index, representing a four-point improvement over Spark 1.2 and matching GPT-5.6 Sol max, Grok 4.6 high and Claude Opus 5 high. Testing of the restricted-preview max variant yielded a score of 62, positioning it behind only Claude Fable 5.1 and Claude Opus 5 among comparable models at launch.

Alex Volkov, AI evangelist at infrastructure provider CoreWeave, cited the benchmark results as evidence of Meta's rapid convergence with leading competitors.

Artificial Analysis Intelligence Index (credit: Artificial Analysis)
Artificial Analysis Intelligence Index (credit: Artificial Analysis)

Damn, gloves are off!

Alex Volkov, CoreWeave, on X

Volkov characterized the performance as "quite the statement" from Meta and predicted "busy weeks ahead of us."

Pricing represents another component of Meta's competitive strategy. Artificial Analysis identifies Spark 1.3 xhigh as the "most cost-efficient model" at its measured intelligence level, with its 61 Intelligence Index score combined with relatively modest per-task expenses positioning it on the company's Pareto efficiency frontier.

Intelligence Index vs. Cost per Intelligence Index Task (Credit: Artificial Analysis)
Intelligence Index vs. Cost per Intelligence Index Task (Credit: Artificial Analysis)

According to Artificial Analysis calculations, Spark 1.3 xhigh costs approximately $0.55 per Intelligence Index task, the lowest among models scoring 59 or higher. GPT-5.6 Sol max and Grok 4.6 high, both matching Spark at 61, cost $0.95 and $0.94 per task respectively.

Spark 1.3 carries higher per-task expenses than its predecessor Spark 1.2, which cost $0.40. Artificial Analysis attributes this increase primarily to the newer model requiring roughly 57% additional input tokens during agentic evaluations.

Cost per Intelligence Index Task (Credit: Artificial Analysis)
Cost per Intelligence Index Task (Credit: Artificial Analysis)

Muse meets Gemini: A 'playground slapfight'

Shortly before Meta's Spark 1.3 announcement, Google introduced Gemini 3.8 Flash, its third Flash iteration within six weeks. Google similarly promoted its model as its most capable Flash variant for reasoning and coding, while maintaining introductory pricing at $0.75 per million input tokens and $3.75 per million output tokens.

Artificial Analysis initially positioned Gemini 3.8 Flash (high) on its Intelligence-versus-Cost Pareto frontier—the group of models for which no alternative exists that is simultaneously more capable and less expensive. Gemini achieved a 59 Intelligence Index score at $0.58 per task.

Within hours, however, Muse Spark 1.3 xhigh arrived with a 61 score and $0.55 per task, surpassing Gemini 3.8 Flash on both dimensions and displacing it from the frontier.

Google was ahead only a few hours.

Benjamin Marie, AI researcher, on X

The rapid turnover did not escape notice in the AI community. AI researcher Benjamin Marie observed the brief window of Gemini's advantage on X.

https://x.com/bnjmn_marie/status/2095251751737971190?ref_src=twsrc%5Etfw

Florian Brand, research engineer at Prime Intellect, summarized the shift with brevity: "Gemini 3.8 held a spot at the pareto frontier for *checks notes* 3.5 hours."

Meta's chief AI officer Alexandr Wang also commented, leveraging the Artificial Analysis findings to critique Google's position.

Corey Quinn, co-founder and chief cloud economist at Duckbill, offered a skeptical perspective on the exchange, suggesting that neither Meta nor Google are meaningfully driving the broader AI competition. "Meta casting shade at Google in AI is a playground slapfight outside a MMA championship," Quinn wrote on X.

https://x.com/alexandr_wang/status/2095249704888197175?ref_src=twsrc%5Etfw

Muse Code enters the scene

Muse Spark 1.3 marks the fourth iteration Meta has released since April, following Muse Spark 1.1 in July and 1.2 in August. However, the more strategically significant element of Meta's initiative involves the agent framework being constructed around the model.

This framework takes shape through Muse Code, Meta's terminal-based coding agent, which entered beta alongside Spark 1.2 in early August. The agent officially launched on Tuesday with subscription plans beginning at $5 monthly and a new SDK entering developer preview.

While Meta is evidently pursuing competition with leading model developers on raw capability—as demonstrated by Muse Spark's progression—the company is also competing at a higher level, where Anthropic's Claude Code and OpenAI's Codex have established benchmarks for how developers interact with agents in practice.

This context explains why the 1.3 release messaging emphasizes coding and agentic capabilities specifically. The model functions as a central component of a broader developer platform where Meta is competing on pricing as vigorously as on technical performance.