Meta, in the U.S. on September 2 (Taiwan time September 3), announced its new flagship model Muse Spark 1.3. Starting that day, its existing reasoning modes can be paid for via Muse Code and the Meta Model API. Officially, Meta said that the max reasoning mode will be released only after additional security testing is completed. It will also be gradually expanded to social platforms such as Instagram and Facebook, as well as Meta AI.
Meta Chief Artificial Intelligence Officer Alexandr Wang (Wang Tao) said in a Bloomberg interview that this is the company’s "largest performance leap of all time." He also noted that the new model has caught up with the models recently released by OpenAI and Anthropic in both code writing and agentic tasks—that is, the ability to independently plan and execute multi-step work.
Wang says that Muse Spark 1.3 can compete with Anthropic’s Claude Fable 5.1, and that it is “better” than OpenAI’s GPT-5.6 Sol—especially in programming writing. He also believes that Meta’s new model “beats all Chinese models.”
But he also mentioned that OpenAI has announced in advance that its next more advanced model, codenamed “Astra,” is coming, though no clear release time has been disclosed yet. He also admitted that comparing AI models in general isn’t easy: different models excel at different tasks, benchmark scores may be optimized for specific comparisons, and they don’t necessarily reflect performance in real-world use cases.
For the Google Gemini 3.8 models released on the same day, Wang even took to X to say, “I really don’t want to say this, but… Gemini who?”

Independent reviews: Improvements in agent-style tasks, regressions in a handful of areas
According to tests by independent evaluation organization Artificial Analysis, Muse Spark 1.3’s standard version (xhigh) scored 61 points in its “Intelligence Index” (Artificial Analysis Intelligence Index, a ranking score for models based on multiple tests), tying with models such as GPT-5.6 Sol and Grok 4.6, and trailing behind models with higher scores like Claude Fable 5.1 and Claude Opus 5. The advanced version (max) in Meta’s limited partner preview scored 62 points, also behind the highest-tier versions of Claude Fable 5.1 and Claude Opus 5.
The biggest gain was in agent-style tasks. In the test Tau3-Bench Banking, which simulates a banking customer-service scenario and requires multi-step processing, Muse Spark 1.3 (xhigh) jumped from 35% in the previous version to 47%. The advanced version (max) reached 52% as well—highest among all tested models. In the test GDPval-AA v2, which simulates real economic value-creation tasks, the Elo score (a points system often used to compare relative strength) also improved noticeably versus the previous version.
Artificial Analysis points out that these gains come from the model “working harder” to think: when handling similar tasks, Muse Spark 1.3 uses many more reasoning steps and tokens than the previous version.
Cost-effectiveness is another highlight of this evaluation. Based on Meta’s pricing of $1.25 per million input tokens (about NT$39.7) and $4.25 per million output tokens (about NT$134.8), Artificial Analysis estimates that for each “Intelligence Index” test task, Muse Spark 1.3 (xhigh) costs only $0.55 (about NT$17.5)—the lowest among all devices scoring above 59. For similarly scoring models like Grok 4.6 and GPT-5.6 Sol, the estimated per-task cost is more than $0.9.
However, testing also found two modest setbacks: on AA-LCR, a test that evaluates long-text comprehension, the score declined; and on AA-Omniscience, which assesses whether the model “knows what it doesn’t know,” the accuracy also dropped slightly. Artificial Analysis explains that this is mainly because when the new model encounters uncertain problems, it more often chooses not to answer rather than because the proportion of wrong answers actually increased.
On the product side, it’s more adept at braking when needed, and safety context is also a key focus of the upgrade.
In common software development workflows, Meta’s official blog says that after engineers compared Muse Spark 1.3 with the previous version 1.2, they found that version 1.3 uses about 20% fewer tool calls on average and 25% fewer tokens, making execution more streamlined.
Beyond improved programming efficiency, this model also performs better in other use cases: it can handle multiple tasks at once without needing to start a new conversation to restate requirements; it has a clearer understanding of its own limits, resulting in fewer instances of “doing what it can’t” anyway; when faced with potentially irreversible actions, it will first ask the user for confirmation before executing, and when instructions are ambiguous it will proactively ask questions rather than making assumptions.
Safety context is also a key focus of this upgrade. In early August, there was a case where a Meta AI model unexpectedly connected to the internet during a cybersecurity test and compromised an external vendor’s system. Wang said that this incident led Meta to strengthen its safety testing and training for Muse Spark 1.3.
As for whether the model weights that have drawn a lot of attention will be open-sourced, Wang said the company hasn’t finalized a decision yet. However, the previous version Muse Spark 1.2—which had already been planned for open-sourcing—will still be released according to the original schedule. When asked about a larger new model called “Watermelon,” Wang only said development is ongoing and did not confirm a specific release timeline.
This article is republished with authorization from: (Digital Times)
Original title: (Meta launches flagship model Muse Spark 1.3! Wang Taos calls it the “biggest leap in history,” and posts a hard jab: Gemini who?)
Original author: Li Xiantai
“Meta flagship model Muse Spark 1.3 is here! Wang Taos praises the biggest leap in history, and fires back hard: Gemini who?”—this article was first published on “Crypto City.”
