Benchmark Gains and Autonomous Workflows
Independent evaluations of frontier artificial intelligence models increasingly emphasize execution in practical workflows over pure academic recall. Meta has launched Muse Spark 1.3 through Muse Code and the Meta Model API, marking another rapid update in its enterprise tooling lineup. The release structure makes the xhigh tier available immediately, while the more compute-heavy max tier operates as a limited partner preview.
Evaluation data from Artificial Analysis shows that the model concentrates its performance gains in simulated autonomous environments. On the τ³-Banking benchmark, which tests agents operating tools inside a simulated banking environment, the max variant reaches 52 percent, taking the top spot on the leaderboard. The xhigh tier achieves 47 percent on the same test, tying Claude Fable 5.1 (max) and GLM-5.3-Flash.
Coding and professional workflow benchmarks show similar gains, although the model trails top competitors in overall scores. On Terminal-Bench 2.1, Muse Spark 1.3 climbs to 85 percent on xhigh and 86 percent on max, while Claude Fable 5.1 max maintains the lead at 91.4 percent. On GDPval-AA v2, which measures performance across real-world professional tasks, Meta reaches 1,709 points on xhigh and 1,754 on max, trailing Claude Fable 5.1 (max) at 1,853.
The Economics of Inference
Performance in specialized domains outside practical tool use remains mixed across standard evaluations. On GPQA Diamond, Muse Spark 1.3 reaches 94 percent, placing it below Gemini 3.8 Flash (high) at 95.3 percent and Grok 4.6 (high) at 94.9 percent.
Despite trailing top competitors on certain academic benchmarks, the model introduces substantial pricing pressure across enterprise deployments. Muse Spark 1.3 costs $1.25 per million input tokens and $4.25 per million output tokens, which calculates to $0.55 per index task.
No model scoring 59 points or higher on the Intelligence Index is cheaper than Meta's $0.55 per task.
Direct rivals operating at the same index level run between $0.94 and $1.23 per task. On the Intelligence Index, Muse Spark 1.3 reaches 61 points on xhigh and 62 points on max, while max consumes 62 percent more reasoning tokens than the xhigh tier.
Meta has confirmed that an open-weights version of Muse Spark is forthcoming, creating an additional long-term consideration for corporate infrastructure budgets. By combining aggressive pricing on API endpoints with plans for downloadable weights, the rollout forces closed-ecosystem model providers to justify a significant price premium on comparable agentic workloads.