Meta's foundational model, Muse Spark 1.2, achieved a high score in a third-party performance analysis, demonstrating rapid growth in the Muse series just four months after its launch.

It has been revealed that
Muse Spark 1.2: Improved Agentic Performance at Higher Cost per Task
https://artificialanalysis.ai/articles/muse-spark-1-2
Muse Spark 1.2 (xhigh) - Intelligence, Performance & Price Analysis
https://artificialanalysis.ai/models/muse-spark-1-2
Muse Spark 1.2 (xhigh) has a score of '54' in the Artificial Analysis Intelligence Index. Muse Spark 1.0, released in April 2026, scored '43,' and Muse Spark 1.1, released in July 2026, scored '51,' meaning Meta has dramatically increased its benchmark score in just four months.
Furthermore, the score for Muse Spark 1.2 (xhigh) was almost the same as GPT-5.5 (xhigh)'s '55' and Grok 4.5 (high)'s '54,' and slightly lower than the most advanced models at the time of writing, such as Claude Opus 5 (max)'s '61,' Claude Fable 5 (fallback)'s '60,' GPT-5.6 Sol (max)'s '59,' and Kimi K3 (max)'s '57.' Artificial Analysis, which conducts the Artificial Analysis Intelligence Index, points out that 'this is a record tied for third place among American companies' AI, alongside SpaceX AI.'

The graph below summarizes the Artificial Analysis Intelligence Index score (vertical axis) and cost (horizontal axis). Muse Spark 1.2 is a high-performance model compared to models with equivalent cost.

In GDPval-AA v2, a benchmark test for measuring the practical performance capabilities of AI agents, the score for Muse Spark 1.2 rose from '1371' for Muse Spark 1.1 to '1631'. Only Claude Opus 5 (max) achieved a score higher than Muse Spark 1.2 ('1852'), Claude Fable 5 (fallback) achieved a score of '1743', GPT-5.6 Sol (max) achieved a score of '1730', and Kimi K3 (max) achieved a score of '1685'.

In Terminal-Bench v2.1, a benchmark test that evaluates the AI agent's ability to perform complex tasks in a real terminal environment, the score improved from '78%' to '80%'.

In the τ³-Banking benchmark test, which measures how accurately AI agents can perform financial customer support tasks, the score improved from '25%' to '27%'.

In SciCode, a benchmark test that evaluates how accurately AI models can generate practical numerical computation code used in scientific research, the score dropped from 58.2% to 56.4%. However, only Claude Fable 5 (fallback), Kimi K3 (max), and Muse Spark 1.1 recorded better scores than Muse Spark 1.2.

In Humanity's Last Exam, a benchmark test that measures whether AI models possess the highest level of human expertise and genuine reasoning capabilities, the score dropped from 45.1% to 43.9%.

In the AA-Omniscience Index, a benchmark test that rigorously measures the accuracy of the AI model's expertise, its ability to suppress hallucinations, and its understanding of its own knowledge boundaries, the score increased from '18' to '22'. The incidence of hallucinations decreased from '38%' to '28%', and the accuracy rate decreased from '41%' to '38%'.

Furthermore, in the Vals Index , Vals AI 's benchmark, Muse Spark 1.2 achieved the fifth highest score of '71.88%'. The cost per test was the cheapest among the top five AI models at $0.69 (approximately 110 yen).
Muse Spark 1.2 just cracked the top 5 on the Vals Index, at just $0.69 per test. This is 3x cheaper than Kimi and 10x or more cheaper than Fable, Opus, and 5.6 Sol. pic.twitter.com/E2pFFqQiBI
— Vals AI (@ValsAI) August 6, 2026
Cline, an open-source AI coding agent that runs on VS Code, reported: 'We tried Meta's new Muse Code agent, but it has a bug that prevents signing in from Docker containers. So we did a fun experiment. Meta claims that Muse Spark 1.2 was co-trained with their Muse agent harness. So we extracted instructions from their system prompts and added them to the Cline harness.' They then asked this harness to fix the bug from our repository and compared the results to the original Cline agent harness. The results were 2.7 times less token usage (19.7 million → 7.2 million), 2 times faster completion (49 minutes → 24 minutes), and 2.4 times lower cost ($7.69 → $3.25),' praising the improvements in Muse Spark 1.2.
We tried using Meta's new Muse Code agent, but it has a bug that doesn't let it sign in from a docker container.
— Cline (@cline) August 6, 2026
So we did a fun experiment: Meta claims Muse Spark 1.2 was co-trained with their Muse agent harness. So we extracted instructions from their system prompt and added… https://t.co/kOkilroGEK pic.twitter.com/zhxsezv6gs
AI researcher Rihard Jarc said, 'It's quite remarkable, considering the timeframe, that Meta appears to have surpassed Google's AI models in quality in many use cases with Muse Spark 1.2. Given that Meta is already planning to release a more powerful model than Muse Spark, the high-performance model (Watermelon) should deliver Claude Fable-level performance. After a year of major overhauls, Meta finally seems to have a scaling foundation for its AI lab. The delivery speed is also very good. Given the rumors that DeepSeek is planning to raise prices, it's clear that having sufficient computing resources to serve customers is important. Having a great model is meaningless if most people can't use it. Meta is one of the few companies with computing power on par with or greater than Anthropic and OpenAI.'
A few thoughts on the $META AI model's progress, because I think it is significant.
— Rihard Jarc (@RihardJarc) August 6, 2026
1. It does seem that $META has now leapfrogged $GOOGL in model quality when it comes to Muse 1.2 for many use cases, which is very surprising given the timeframe.
2. This is still the “Muse… https://t.co/1tqMtgjGKG
Related Posts:
in AI, Posted by logu_ii







