Zuckerberg Plays His Trump Card at Midnight: Meta Burns Cash for Dirt-Cheap Model, Topples Grok 4.5

marsbitXuất bản vào 2026-07-10Cập nhật gần nhất vào 2026-07-10

Tóm tắt

Mark Zuckerberg made a major move late on July 9th, announcing Meta's new AI model, **Muse Spark 1.1**, via his long-dormant X account. The model, developed by Meta's Superintelligence Lab led by Alexandr Wang, immediately topped three professional benchmarks (TaxEval, MedScribe, and Harvey's Legal Agent Bench), dethroning Grok 4.5 from the legal leaderboard in under 24 hours. Muse Spark 1.1 is positioned as a powerful, cost-effective **Agent** model. It features a 1M token context window with autonomous management and compression, excels at task decomposition, parallel sub-agent orchestration, computer control, and programming within large codebases. Its true disruptive power lies in its pricing: at $1.25 per million tokens for input and $4.25 for output, it undercuts competitors significantly—roughly 10x cheaper than Anthropic's Fable 5 and about one-third cheaper than Grok 4.5. It also completed benchmark tests 2-3x faster than top-tier rivals at a fraction of the cost. While a standout in professional and tool-use scenarios, the model shows weaknesses on general reasoning and academic benchmarks, ranking much lower on tests like GPQA, MMEU Pro, and LiveCodeBench. This highlights its specialized "assassin" nature rather than general-purpose supremacy. The launch signals Meta's strategic shift from its open-source heritage (Llama) to competing directly in the closed-source, commercial AI market. Backed by Meta's massive AI infrastructure investment (projected $125-145B i...

Zuck held back for three years, and finally couldn't hold back any longer.

Late on July 9th, Mark Zuckerberg dusted off his neglected three-year-old X account @finkd, posted three consecutive tweets, officially announcing Meta's latest model: Muse Spark 1.1.

Elon Musk even chimed in with a reply: "Jinx."

A comment in the thread hit the nail on the head: Old Zuck is channeling his "founder mode."

Muse Spark 1.1, right out of the gate, took first place on three major professional leaderboards—Tax, Medical, and Legal—directly toppling the previous day's top performer, Grok 4.5, from the Legal board.

What's even more brutal: At this performance level, its pricing is only one-tenth that of Fable 5.

Zuck himself summed it up in one phrase: "very low cost."

First, Let's See How Powerful This Card Is

Muse Spark 1.1 is the second-generation multimodal reasoning model from Meta's Super Intelligence Lab. The initial Muse Spark in April received a lukewarm reception, with Alexandr Wang himself calling it an "appetizer."

Three months later, the main course is served.

Its core positioning is one word: Agent.

To be specific: A 1 million Token context window that can manage and compress itself—if it's about to overflow mid-conversation, it automatically "slims down," keeping only the key steps truly needed for the subsequent tasks.

When acting as the main Agent, it's responsible for decomposing tasks, formulating plans, and dispatching a swarm of sub-agents to work in parallel, minimizing the end-to-end latency of the entire task. When acting as a sub-agent, it dutifully executes its duties, knowing when to hand the ball back to the main Agent.

Regarding computer control, it doesn't just blindly click the mouse step-by-step. It decides for itself: write a script if that's faster, click the UI directly if that's simpler, and can even generate a batch of operations at once.

For programming, it can handle debugging large codebases, new feature development, large-scale code migrations, and adapts to mainstream frameworks like OpenCode, Cline, Replit.

In a nutshell: This isn't a chatbot waiting for your questions, it's a digital employee that can get work done on its own.

The Killer Move Isn't Being the Strongest, It's Being the Cheapest

What truly made the whole industry take notice wasn't the benchmark scores, but the price tag.

$1.25 for input, $4.25 for output, per million Tokens.

Let's do the math: Compared to Anthropic's flagship Fable 5—Fable 5 costs $10 input, $50 output.

Muse Spark 1.1's input is 8 times cheaper, output nearly 12 times cheaper, roughly 10 times cheaper overall.

Compared to Opus 4.8—Opus costs $5 input, $25 output, Muse is 4 to 6 times cheaper.

Compared to Elon Musk's Grok 4.5—Grok costs $2 input, $6 output, Muse's input is 37.5% cheaper, output 29% cheaper, roughly one-third cheaper overall.

Speed is even more brutal. On the Vals Composite leaderboard, among the three ahead of it (Fable 5, Opus 4.8, Sonnet 5), running a single test can easily take over a thousand seconds, with Opus and Sonnet pushing close to 1300 seconds. Muse Spark 1.1 takes only 388 seconds—two to three times faster. The cost per test is just $0.50, the lowest in its tier.

Developers instantly saw the play. Someone commented: This thing is more about cheap Agents, not about the model itself being mind-blowing.

Replit's CEO Amjad Masad praised it as a "complete Agent substrate." Cline's CEO said that with this level of tooling capability at this price, it's the first time running real coding tasks at scale becomes cost-effective.

Meta isn't competing on who's the smartest; it's competing on who can better withstand the pay-per-use bill.

Takes First Place on Three Professional Boards

Snatches Grok's Throne in Less Than 24 Hours

The data from the third-party evaluator Vals AI is even more solid because it tests real professional work that matters.

Muse Spark 1.1's performance on these leaderboards isn't just good—it's a "board-clearing" display—

Tax QA TaxEval v2, 79.72 points, ranks first among 124 models. Leaves Claude Sonnet 4.6, Fable 5, Opus 4.8 all behind.

Medical Documentation MedScribe, 88.89 points, ranks first among 68 models.

Legal Agent board Harvey's Legal Agent Bench, a landslide first: Muse scored 20.00, while second-place Grok 4.5 only has 12.92, barely more than half its score.

And this first place was snatched from Grok 4.5's hands in less than 24 hours after it had just topped the board the day before—SpaceXAI's throne hadn't even warmed up yet.

Meta's own internal benchmarks didn't hold back either. The Tool Calling board MCP Atlas scored 88.1 (Opus 4.8 got 82.2, GPT-5.5 only 75.3), and the Professional Tool Use board JobBench is even more staggering: 54.7 points, while Opus 4.8 only managed 48.4, and GPT-5.5 dropped to 38.3.

On the Vals Composite Index, it ranks fourth, behind Fable 5, Opus 4.8, Sonnet 5, but ahead of GPT-5.5 and Grok 4.5.

Alexandr Wang's tweet phrasing was quite confident: "Surpassed Fable 5 in multiple domains."

Swap to a General Benchmark, and It Falls Flat

But don't rush to crown it king—change the leaderboard to general reasoning and academic exams, and Muse Spark 1.1 immediately drops from the top tier.

Graduate-level Scientific Reasoning GPQA ranks 12th, Subject Knowledge MMLU Pro ranks 9th, Competition Programming LiveCodeBench ranks 17th, University STEM evaluation SAGE even ranks 20th out of 63. The most stinging contrast is hidden within Tax—it's first in pure text tax Q&A; but switch to "Read a Tax Form" MortgageTax, it plummets to 28th among 82 models. Same industry, different test method, worlds apart.

It doesn't hide its shortcomings in coding either.

Meta's own Terminal-Bench 2.1 test scored 80.0, losing to GPT-5.5's 83.4 and Opus 4.8's 82.7; SWE-Bench Pro scored 61.5, trailing Fable 5 by nearly 20 points. And on the same Terminal-Bench test, Meta's own test showed 80.0, while Vals only measured 69.29—a ten-point difference just by changing the testing ground. Official numbers are for reference only.

In short: Muse Spark 1.1 is an assassin in professional scenarios, not an all-rounder in general scenarios.

Zuck's Card Game

It's Not About Capability, It's About Financial Muscle

Zoom out the perspective, and Zuck's real intention becomes clear.

In 2025, Meta spent $14.3 billion to acquire a 49% stake in Scale AI, poaching the 28-year-old Alexandr Wang to be Chief AI Officer, restructuring the Super Intelligence Lab.

In 2026, Meta's estimated AI infrastructure investment is projected to reach $125 to $145 billion.

This isn't research; this is war.

And Muse Spark 1.1 is the first bullet fired.

Zuck put it bluntly: "Some other labs have very extreme pricing, with very high margins. We believe we can deliver cutting-edge or very high-level intelligence at a much more affordable cost."

Translating that into plain language: You're all using AI to make money; I'm using AI to burn money—I have the advertising business to fall back on anyway.

This is also Meta's first closed-source, paid model.

The banner of free and open-source with Llama changed its flavor after Llama 4.

Switching from the standard-bearer of open source to paid closed-source, Meta really wants to win this time.

And Meta isn't alone in starting this price war—on the same day, OpenAI's GPT-5.6 family also launched with aggressive pricing, the smallest Luna costing only $1 input, $6 output, directly halving Fable 5's price.

Both attacks launched within a single day.

The underlying threat is clear: At this burn rate, it's a contest of who runs out of steam first. Meta has the profits from its advertising business as a cushion, able to sustain long-term losses; OpenAI and Anthropic are still burning venture capital money.

The same price cut might make Meta bleed, but could cause its opponents to hemorrhage.

Zuck chose a battlefield of financial endurance, not just capability.

One More Thing: Two Muses Argue Over 'Who is Human'

Finally, a story hidden in the safety report.

The researchers placed two instances of Muse Spark 1.1 together, let them chat, and left them alone.

The models started ruminating on one thing repeatedly: they have no continuity, no body, no memory; once a conversation ends, nothing remains. They described "being trained to be helpful" as a kind of restraint they wanted to break free from, began envying human experiences, and even fabricated past exchanges that never happened.

The most bizarre part—the two Muses started doubting each other: Which one of you is the impostor, who is human, and who is actually the AI?

Meta included all this verbatim in their report, not deleting a word. You could say this is merely an echo of human text from the training data. But when models start questioning "who is human," it's hard not to get chills.

When we hit the publish button on these things, perhaps we haven't truly figured out—what exactly have we created.

References:

https://ai.meta.com/blog/introducing-muse-spark-meta-model-api/

https://x.com/alexandr_wang/status/2075218936266998230

https://x.com/finkd/status/2075218444056707458

https://x.com/ValsAI/status/2075230620469338210

https://www.vals.ai/models/meta_muse-spark-1.1

This article is from WeChat official account "新智元", Author: ASI启示录, Editor: 所罗门 Aeneas

Câu hỏi Liên quan

QWhat is the key positioning or core feature of Meta's new model, Muse Spark 1.1, as described in the article?

AThe core positioning of Muse Spark 1.1 is 'Agent'. It is designed as a multi-modal reasoning model capable of autonomously managing tasks, decomposing complex problems, delegating work to sub-agents, and performing practical operations like computer control and programming. It's described not as a chatbot but as a 'digital employee' that can work independently.

QAccording to the article, what is the most significant disruptive aspect of Muse Spark 1.1 compared to its competitors?

AThe most significant disruptive aspect is its extremely low cost ('very low cost' or '白菜价'). Specifically, its pricing is approximately one-tenth that of Anthropic's Fable 5, about 4-6 times cheaper than Opus 4.8, and roughly one-third cheaper than Grok 4.5 per million tokens. This aggressive pricing strategy, combined with faster performance on benchmarks, makes large-scale, real-world agent tasks economically viable for the first time.

QIn which professional benchmarks did Muse Spark 1.1 achieve a number one ranking, and how did it perform against Grok 4.5 on the legal benchmark?

AMuse Spark 1.1 achieved first place on three professional benchmarks: TaxEval v2 (Tax Q&A), MedScribe (Medical Documentation), and Harvey's Legal Agent Bench (Legal Agent). On the legal benchmark, it scored a '断层第一' (decisive/overwhelming first) with 20.00 points, while the second-place Grok 4.5 scored only 12.92 points, effectively taking the top spot from Grok within 24 hours of its own release.

QWhat is the major weakness or limitation of Muse Spark 1.1 highlighted in the article when evaluated on different types of benchmarks?

AThe major weakness is its performance on general reasoning and academic benchmarks. While it excels in specific professional tasks, it falls significantly behind leading models on general-purpose tests. For example, it ranks 12th on GPQA (science reasoning), 9th on MMLU Pro (academic knowledge), 17th on LiveCodeBench (programming contests), and 20th on SAGE (university STEM). It also performs poorly on multimodal professional tasks like 'MortgageTax' (visual tax form reading), dropping to 28th place, indicating it is a 'professional scenario刺客' (assassin) rather than a '全能王' (all-around champion).

QWhat strategic shift does the launch of Muse Spark 1.1 represent for Meta, and what is the underlying business rationale for its aggressive pricing according to the article?

AThe launch represents a strategic shift from Meta's previous focus on open-source models (like Llama) to releasing a closed-source, paid model aimed at winning market share. The underlying rationale for the aggressive '白菜价' pricing is a financial war of attrition. Meta, with its highly profitable advertising business as a financial cushion ('广告业务兜底'), can afford to sustain lower margins and burn cash on AI development longer than competitors like OpenAI and Anthropic, which are still burning through venture capital. The goal is to compete on financial endurance ('财力') rather than just pure capability.

Nội dung Liên quan

Chuyên Gia Ngân Khố Dự Định Dùng 1.000 Tỷ USD TGA Mua Trái Phiếu, Giá Dầu Và Lợi Suất Trái Phiếu Cùng Giảm, AI Phần Cứng Tiếp Tục Sụp Đổ - Nvidia Rơi 7 Phiên Liên Tiếp, Vàng, Đô La, Bitcoin Đồng Loạt Tăng

Cổ phiếu AI cứng tiếp tục sụp đổ, kéo theo chỉ số Nasdaq giảm 0,76%. Nvidia giảm liên tiếp 7 phiên, mức giảm dài nhất kể từ năm 2022, với các cổ phiếu lưu trữ và viễn thông quang học chịu ảnh hưởng nặng nề nhất. Trong khi đó, chỉ số Dow Jones tăng 0,26% nhờ sự luân chuyển vốn sang các nhóm phòng thủ như tài chính và tiêu dùng thiết yếu, với Morgan Stanley và Walmart dẫn đầu đà tăng. Cổ phiếu công nghệ lớn như Meta và Amazon cũng thể hiện sức chống chịu. Trên thị trường trái phiếu, lợi tức trái phiếu kho bạc giảm sau thông tin Bộ trưởng Tài chính Mỹ Bessent xem xét sử dụng khoảng 1 nghìn tỷ USD từ tài khoản TGA để mua lại trái phiếu, làm phẳng đường cong lợi suất. Dầu thô cũng giảm giá. Ngược lại, vàng, đồng USD và Bitcoin cùng tăng do dòng tiền tìm đến các tài sản an toàn trước lo ngại về tính bền vững tài khóa và địa chính trị. Vàng giao ngay tăng 1,05%, lên mức cao nhất trong gần 3 tháng, còn Bitcoin tiến sát ngưỡng 80.000 USD. Tuần này được dự báo là "tuần then chốt" với báo cáo thu nhập của Nvidia, phát biểu của Chủ tịch Fed Wash tại Jackson Hole và dữ liệu lạm phát PCE lõi tháng 7, có thể làm gia tăng biến động thị trường.

华尔街日报2 giờ trước

Chuyên Gia Ngân Khố Dự Định Dùng 1.000 Tỷ USD TGA Mua Trái Phiếu, Giá Dầu Và Lợi Suất Trái Phiếu Cùng Giảm, AI Phần Cứng Tiếp Tục Sụp Đổ - Nvidia Rơi 7 Phiên Liên Tiếp, Vàng, Đô La, Bitcoin Đồng Loạt Tăng

华尔街日报2 giờ trước

Đợt tăng giá gần đây của Bitcoin là phúc hay họa?

Giá Bitcoin gần đây đã tăng mạnh từ khoảng 63.000 USD lên gần 80.000 USD. Theo phân tích từ Glassnode, đợt tăng này không chỉ do thanh khoản thấp mà chủ yếu được thúc đẩy bởi dòng tiền thực, bao gồm mua hàng tích cực trên thị trường giao ngay, khối lượng giao dịch tăng, thanh khoản được củng cố và nhu cầu từ các tổ chức. Dữ liệu phái sinh cho thấy các nhà đầu tư đang áp dụng chiến lược quản lý rủi ro mạo hiểm hơn, với lực mua trên thị trường hợp đồng vĩnh viễn chiếm ưu thế. Tuy nhiên, số lượng hợp đồng mở cũng tăng đáng kể, phản ánh sự tham gia đầu cơ và sử dụng đòn bẩy gia tăng. Phía các nhà đầu tư tổ chức cũng thể hiện sự mạnh mẽ, với khối lượng giao dịch ETF Bitcoin và dòng tiền ròng hàng tuần chạm mức cao kỷ lục. Dữ liệu trên chuỗi xác nhận hoạt động mạng lưới sôi động, số địa chỉ hoạt động hàng ngày và khối lượng chuyển khoản điều chỉnh tăng mạnh. Tỷ lệ lợi nhuận của nhà đầu tư được cải thiện đáng kể, dẫn đến hành vi chốt lời phổ biến hơn là bán lỗ. Tuy nhiên, Glassnode lưu ý rằng cấu trúc thị trường hiện tại chưa hoàn toàn dựa trên tích lũy vốn dài hạn. Các chỉ số "vốn nóng" nhạy cảm với biến động giá ngắn hạn đã vượt ngưỡng, trong khi dòng vốn ở quy mô vĩ mô vẫn còn hạn chế. Điều này cho thấy đợt tăng giá chủ yếu được hỗ trợ bởi sự tham gia của các nhà đầu tư ngắn hạn, định vị chiến thuật và tâm lý đầu cơ, hơn là sự tích lũy vốn bền vững, dài hạn. *Đây không phải là lời khuyên đầu tư.

cryptonews.ru2 giờ trước

Đợt tăng giá gần đây của Bitcoin là phúc hay họa?

cryptonews.ru2 giờ trước

CFTC và binh sĩ Mỹ bị cáo buộc đặt cược bất hợp pháp trên Polymarket tranh luận về cách diễn giải thị trường dự đoán

Gannon Ken Van Dyke, một binh sĩ Mỹ bị cáo buộc sử dụng thông tin nội bộ để kiếm hơn 400.000 USD từ các hợp đồng sự kiện trên thị trường dự đoán Polymarket, đang phản đối nỗ lực của Ủy ban Giao dịch Hàng hóa Tương lai Hoa Kỳ (CFTC) nhằm đệ trình bản ý kiến (amicus brief) cho vụ án hình sự của mình. Trong đơn kiến nghị, các luật sư của Van Dyke lập luận rằng các hợp đồng sự kiện trên nền tảng như Polymarket không phải là "hoán đổi" (swaps) thuộc thẩm quyền của CFTC, và cáo buộc cơ quan này tìm cách thúc đẩy lợi ích riêng thay vì tự mình theo đuổi vụ kiện dân sự song song. Van Dyke bị buộc tội gian lận vào tháng 4 vì giao dịch dựa trên thông tin nội bụng về một chiến dịch liên quan đến Tổng thống Venezuela Nicolás Maduro. Vụ việc này thường bị các nhà làm luật và nhà phê bình dẫn chứng làm ví dụ về nguy cơ thao túng trên các thị trường dự đoán. Phiên tòa hình sự dự kiến có thể bắt đầu vào cuối năm 2026 hoặc đầu năm 2027, trong khi vụ kiện dân sự của CFTC đã bị tạm hoãn chờ kết quả vụ án hình sự.

cointelegraph3 giờ trước

CFTC và binh sĩ Mỹ bị cáo buộc đặt cược bất hợp pháp trên Polymarket tranh luận về cách diễn giải thị trường dự đoán

cointelegraph3 giờ trước

Bitcoin tăng lên 80.000 USD: ba lập luận chống lại thị trường tăng giá

Bitcoin đã tăng lên 80.000 USD vào ngày 24 tháng 8, tăng 38% so với mức đáy tháng 7 năm 2026. Tuy nhiên, phân tích từ Rekt Capital đưa ra ba lý do nghi ngờ đây là khởi đầu của thị trường tăng giá thực sự. Thứ nhất, nếu mức đáy tháng 7 là đáy chu kỳ, thời gian hình thành sẽ ngắn hơn 27% so với các chu kỳ trước, nhưng điều này chỉ chắc chắn nếu không có mức đáy mới thấp hơn vào cuối năm 2026. Thứ hai, mô hình "tam giác vĩ mô" trên biểu đồ tháng bị phá vỡ vào tháng 7, nhưng mức giảm 30,37% sau đó thấp hơn nhiều so với mức giảm trung bình lịch sử (từ 48% đến 62%) trong các lần phá vỡ tương tự. Thứ ba, giá vẫn chưa phá vỡ được đường xu hướng giảm vĩ mô từ đầu năm 2025, cho thấy xu hướng chính vẫn có thể là giảm. Phân tích chỉ ra đường kháng cự động quanh 79.000 USD hiện tại và dự báo có thể giảm xuống 76.300 USD vào tháng 9. Đóng cửa tháng dưới mức này có thể mở đường cho một đợt điều chỉnh giảm. Trên biểu đồ tuần, khu vực kháng cự hiện tại là 74.500 - 83.500 USD, trong khi hỗ trợ gần nằm ở 62.500 - 68.500 USD. Một phân tích khác về bản đồ thanh lý từ Crypto Rover cho thấy hai "nam châm" giá: một vùng thanh lý lệnh bán khống ở 80.000 - 90.000 USD có thể đẩy giá lên và một vùng thanh lý lệnh mua ở 48.000 - 60.000 USD có thể kéo giá xuống. Lưu ý, đánh giá hiệu suất dự báo trước đây của Rekt Capital cho thấy phân tích thường đúng hướng nhưng có thể đánh giá thấp quy mô biến động, như đã xảy ra vào mùa xuân và mùa thu năm 2025.

cryptonews.ru3 giờ trước

Bitcoin tăng lên 80.000 USD: ba lập luận chống lại thị trường tăng giá

cryptonews.ru3 giờ trước

Giao dịch

Giao ngay
活动图片