Claude's Two New Models Leaked in Real-World Tests, Spatial Reasoning is Divine, Computing Power Directly Dried Up

marsbitXuất bản vào 2026-08-25Cập nhật gần nhất vào 2026-08-25

Tóm tắt

Summary: Claude's two new leaked models, "Marshmallow" and "Melon," demonstrate breakthrough capabilities in 3D spatial reasoning and architectural layout generation, performing complex tasks in a single attempt (one-shot). However, they consume massive amounts of computational power through extensive "thinking tokens" for deep internal reasoning, often nearing system limits. Their sudden emergence follows the reportedly disappointing release of Opus 5, leading to speculation they are either enhanced Opus 5.1 versions or next-generation Sonnet/Haiku models. These developments signal Anthropic's intense focus on advancing AI's deep reasoning and 3D understanding, with significant potential implications for gaming, architecture, and digital twin simulations.

The two new Claude models—Marshmallow and Melon—that were just exposed yesterday have real-world test results appearing today!

Some have found through testing that the two new models are terrifyingly powerful in 3D reinforcement learning and architectural spatial layout.

Furthermore, their consumption of computing power is also incredibly intense.

Before giving an answer, they consume massive amounts of "Thinking Tokens" for deep logical reasoning, repeatedly even pushing against the system's usage limits!

Judging from this model, Claude is evolving into the next-generation AI giant possessing "slow thinking" and deep spatial physical reasoning capabilities.

One-Shot Mastery of 3D Reinforcement Learning and Architectural Layout

Some testers gave this evaluation after testing: Anthropic is really pushing 3D reinforcement learning like crazy.

The spatial layout of buildings is just ridiculously good, and all of this is generated in one step (One-Shot)!

You must know, for LLMs, spatial imagination has always been a fatal weakness.

Getting an AI to understand the physical coordinate relationships and gravity laws of tables and chairs in a 3D room, or planning the layout of a complex building cluster with intricate traffic flows and load-bearing structures, is as difficult as climbing to the sky.

But this time, Claude's "Marshmallow" and "Melon" seem to have crossed this chasm.

They can directly understand and generate complex 3D coordinates and spatial relationships.

Moreover, testing found their architectural layouts are exceptionally outstanding, indicating the new models also perform powerfully when dealing with topological, geometric, and physical constraints.

And, they output in One-Shot, without needing humans to repeatedly adjust prompts to correct errors; after deep internal deliberation, the model can output perfect 3D scene data in one go.

This hides immense commercial and industrial value.

Imagine, future game developers only need to describe in natural language: "Help me generate a medieval-style fortress, containing three defense towers and a hidden underground passage," and Claude can directly output perfect 3D model code or layout parameters.

Architects can directly let the AI generate dozens of preliminary layout schemes that conform to mechanical structures based on terrain and lighting conditions. The construction of the metaverse, industrial digital twins, autonomous driving simulation environments...

All of this will be completely reshaped by this powerful 3D generation capability!

The Double Whammy of Marshmallow and Melon

Yesterday, the whole internet was flooded with marshmallow and melon.

The internal codenames for the two models exposed this time are claude-marshmallow-eap (Marshmallow) and claude-melon-eap (Melon).

This naming convention of "food + EAP" continues their tradition from claude-horchata-eap (Horchata, a Mexican drink), which was briefly tested in July of this year.

Where did these two models suddenly come from?

According to data intercepted by developers from API traffic logs, during the period of 00:45-00:57Z on August 21, the ID claude-marshmallow-ht-eap was called a high frequency of 57 times in API traffic!

Subsequently, it prominently appeared in the Claude Code model list as a "Custom model" with an astonishing 1 million Token ultra-long context window.

Although these two models currently don't have first-party API endpoints and seem limited to red team testers or internal core personnel, the early sporadic test feedback is explosive enough.

"Marshmallow" is considered slightly stronger in overall capability than "Melon."

In daily conversation and logical interaction, the experience of "Marshmallow" is even better than the current Opus 5, with dialogue appearing more natural and pleasant.

Although their current performance still seems below Anthropic's top-tier "Fable" level, are they the legendary Opus 5.1? Sonnet 5.1? Or a new iteration of Haiku?

There is no conclusion yet.

Computing Power Devourer: The Crazy "Thinking Tokens"

Not only that, multiple testers have discovered an extremely anomalous phenomenon: when using these two new models, their consumption of computing power reaches insane levels!

A tester exclaimed on X: "One thing I noticed about both these models, they use massive amounts of thinking tokens to the point where I directly triggered the limit multiple times during testing!"

What are "Thinking Tokens"? Why are they so important?

In past large model interactions, AI often acted like a "fast-talking" respondent, generating answers word by word instantly based on probability distributions.

Claude's new generation of models has completely changed this logic. Before finally providing an answer, the model generates a massive amount of "invisible" tokens in the background for extremely deep self-deduction, chain-of-thought construction, and logical trial and error.

This "regardless of cost" pouring of computational power releases an extremely strong signal: Anthropic is secretly tackling the bottleneck of AI deep reasoning!

This also confirms previous industry speculation—the competition among large models is shifting from "computing power stacking in the training phase" to "computing power consumption in the inference phase."

When Claude starts thinking so frantically that it hits the limit, the quality of the answers it outputs will form a dimensional blow against traditional models.

Suspected Emergency Firefighting? Opus 5's "Revenge Battle"

The sudden leak of these two mysterious models is also timed quite subtly.

Just less than a month ago, on July 24, 2026, Anthropic released the highly anticipated flagship model Opus 5.

Before that, Sonnet 5, released on June 30, was widely praised.

However, Opus 5 suffered a disastrous defeat.

A developer mocked mercilessly on X: "The negative feedback on Opus 5 is so severe that Anthropic had to immediately roll out two new models to try and replace it... I'm dying of laughter."

Indeed, Opus 5 seems to have failed to meet user expectations for a "cross-generational leap." For Anthropic, which is eager to go public, this is undoubtedly a major blow.

Therefore, the exposure of "Marshmallow" and "Melon" at this critical juncture has sparked infinite speculation from the outside world.

Speculation One: Opus 5.1's Last Stand Counterattack.

Many believe these two models, closely associated with Opus 5.1, are precisely the "power-enhanced versions" that Anthropic urgently retooled and rebuilt to address Opus 5's flaws.

By introducing a powerful "Thinking Tokens" mechanism, the logical ceiling of the model is forcibly elevated.

Speculation Two: A Surprise Attack by the New Generation Sonnet / Haiku.

Some testers also think that, considering their astonishing speed and cost-effectiveness, they might be updated versions of Sonnet or Haiku, after all, they haven't been crowned with the top-tier "Fable" title.

But regardless of which scenario, Anthropic is clearly getting restless.

They are frantically accelerating the iteration speed, even to the point of directly releasing cutting-edge models that extremely consume computing power for testing. X

Developers can no longer contain their excitement: "The next few weeks to months are going to get extremely interesting!"

Can Marshmallow and Melon wash away the shame of Opus 5's predecessor and let Anthropic win another round?

References:

https://x.com/Lentils80/status/2091704307863142812?s=20

https://x.com/NFT_Chen/status/2091764673767198730?s=20

This article is from the WeChat public account "New Zhiyuan", author: ASI Revelation; Editor: Aeneas

Câu hỏi Liên quan

QWhat are the two new Claude models mentioned in the article, and what are their internal code names?

AThe two new Claude models mentioned are internally codenamed 'claude-marshmallow-eap' (Marshmallow) and 'claude-melon-eap' (Melon).

QAccording to the article, in which specific areas do the new Claude models demonstrate exceptional performance?

AThe new Claude models demonstrate exceptional performance in 3D reinforcement learning and building spatial layout tasks. They excel at generating complex 3D coordinates, understanding spatial relationships, and creating architecturally sound layouts in a single attempt (One-Shot).

QWhat is a 'thinking token' as described in the context of the new Claude models, and why is it significant?

AA 'thinking token' refers to the invisible, intermediate tokens the new Claude models generate internally before producing a final answer. They are used for deep self-reasoning, chain-of-thought construction, and logical trial-and-error. This is significant because it represents a shift towards AI with 'slow thinking' and deep reasoning capabilities, consuming significant computational power for higher quality outputs.

QWhat is one potential commercial or industrial application suggested for the new models' 3D spatial reasoning capabilities?

AOne potential application is for game developers to use natural language prompts to generate complete 3D model code or layout parameters for complex structures like a medieval fortress with defense towers and hidden tunnels, drastically speeding up development.

QWhat is one speculation in the article about why these two new models were leaked shortly after the release of Opus 5?

AOne speculation is that these models, potentially Opus 5.1 variants, are an emergency response or 'powered-up version' from Anthropic to address the negative reception and perceived shortcomings of the recently released Opus 5 model.

Nội dung Liên quan

TON Đặt Ngày 1 Tháng 9 Là Thời Hạn Cuối Cùng Để Đóng Cổng Cầu Kế Thừa

TON Foundation đã xác nhận sẽ ngừng hoạt động vĩnh viễn cầu nối (bridge) cũ vào ngày 1 tháng 9, đặt ra hạn chót cho người dùng nắm giữ TON được đóng gói (wrapped) và các tài sản bridge liên quan để chuyển chúng về dạng gốc. Việc đóng cửa ảnh hưởng đến bridge-v3.ton.org. Người dùng nắm giữ Wrapped TON dưới dạng token ERC-20 trên Ethereum hoặc BNB Chain, hoặc các j-token như jUSDT trên TON, cần phải chuyển tài sản về trước hạn chót để tránh mất quyền truy cập. Đây là một quá trình chuyển đổi cơ sở hạ tầng có kế hoạch, không phải là sự cố bảo mật. Các cầu nối là bộ phận nhạy cảm, tạo rủi ro vận hành. Nếu bị ngừng hoạt động, người dùng cần hướng dẫn rõ ràng và đủ thời gian để di chuyển tiền. Token phụ thuộc vào bridge có thể trở nên khó đổi hoặc di chuyển nếu người dùng không hành động kịp thời. Tài sản được đóng gói (wrapped) cần được đặc biệt chú ý. Chúng không giống với tài sản gốc và phụ thuộc vào cơ sở hạ tầng bridge. Người dùng nên làm theo hướng dẫn chính thức, sử dụng giao diện bridge đúng và tránh các liên kết lừa đảo. Một đợt ngừng hoạt động có kế hoạch vẫn tạo ra rủi ro, chủ yếu từ việc người dùng không phối hợp kịp thời: bỏ lỡ thông báo, chần chừ, dùng sai giao diện hoặc hiểu nhầm tài sản bị ảnh hưởng. Các mạng lưới có thể cho ngừng hoạt động cầu nối cũ vì nhiều lý do: được thay thế bởi cơ sở hạ tầng mới, chi phí bảo trì cao, hoặc không còn phù hợp với lộ trình. Điều quan trọng là người dùng có đủ thời gian và hướng dẫn đơn giản để di chuyển tài sản an toàn. Mốc quan trọng tiếp theo là hạn chót 1 tháng 9. Cho đến lúc đó, người nắm giữ wrapped TON và j-token nên xác nhận xem mình có bị ảnh hưởng không và sử dụng các kênh chính thức của TON để chuyển tài sản về. Sau hạn chót, quyền truy cập qua đường dẫn cũ có thể bị hạn chế hoặc không thể thực hiện được.

bitcoinist12 phút trước

TON Đặt Ngày 1 Tháng 9 Là Thời Hạn Cuối Cùng Để Đóng Cổng Cầu Kế Thừa

bitcoinist12 phút trước

Kiểm tra nhu cầu: Cá voi Bitcoin kiếm được lợi nhuận kỷ lục 1,2 tỷ USD, còn chủ sở hữu Ethereum quay trở lại vùng có lãi

Các cá voi Bitcoin mới đã ghi nhận lợi nhuận kỷ lục hơn 1,2 tỷ USD chỉ trong ba ngày sau khi giá phục hồi. Ngày 20/8 đạt đỉnh với khoảng 614 triệu USD lợi nhuận thực hiện, cũng là mức cao nhất trong ngày từ trước đến nay. Theo CryptoQuant, việc chốt lời bắt đầu khi Bitcoin vượt lên trên giá thực hiện trung bình (khoảng 68.900 USD) của nhóm cá voi ngắn hạn này. Vào ngày 23/8, Bitcoin giao dịch quanh 77.700 USD, cao hơn khoảng 12,8% so với giá gốc trung bình của họ. Tình trạng này được coi là một bài kiểm tra quan trọng đối với nhu cầu thị trường: nếu Bitcoin có thể giữ được trên mức giá gốc này và khối lượng chốt lời trở lại bình thường, điều đó cho thấy nhu cầu mới có khả năng hấp thụ nguồn cung bán ra. Về phía Ethereum, các nhà đầu tư lớn cũng đã trở lại vùng lợi nhuận chưa thực hiện sau đợt tăng giá. Tuy nhiên, theo nhà phân tích Darkfost từ CryptoQuant, mức độ sinh lời hiện tại vẫn tương đối thấp và khó có khả năng tạo ra áp lực bán đáng kể. Hệ số lợi nhuận/thua lỗ chưa thực hiện cho các nhóm nắm giữ khác nhau dao động từ 0,075 đến 0,38. Ông lưu ý rằng việc tăng khả năng sinh lời này có thể là tín hiệu tích cực, cải thiện tâm lý của các cá voi, vì họ đã chịu khoản lỗ đáng kể vào tháng 6 trước khi Ethereum tăng hơn 65% từ đó đến nay.

cryptonews.ru34 phút trước

Kiểm tra nhu cầu: Cá voi Bitcoin kiếm được lợi nhuận kỷ lục 1,2 tỷ USD, còn chủ sở hữu Ethereum quay trở lại vùng có lãi

cryptonews.ru34 phút trước

Đội ngũ Kinetiq công bố mạng L2 Elysium dành cho Hyperliquid

Vào ngày 24 tháng 8, giao thức staking thanh khoản Kinetiq đã công bố Elysium - một mạng L2 mới cho hệ sinh thái Hyperliquid. Mạng này nhằm tăng thông lượng cho HyperEVM và đơn giản hóa việc ra mắt thị trường giao ngay, token và các ứng dụng DeFi. Gas trên Elysium sẽ được thanh toán bằng $HYPE. Mạng L2 này được lên kế hoạch tích hợp trực tiếp với công cụ giao dịch HyperCore, cho phép các ứng dụng truy cập vào tính thanh khoản và dữ liệu sổ lệnh của nó. Một lý do cho sự phát triển của Elysium là những hạn chế của HyperEVM, như thông lượng thấp và phí cao khi tải mạng nặng. Elysium hứa hẹn tốc độ xử lý giao dịch và tạo khối cao hơn đáng kể so với HyperEVM, nhắm đến các ứng dụng cần cập nhật trạng thái thường xuyên như giao dịch giao ngay tần suất cao và AMM. Elysium cũng đơn giản hóa quy trình niêm yết tài sản mới trong hệ sinh thái Hyperliquid, hợp nhất các bước từ cung cấp thanh khoản ban đầu đến việc ra mắt thị trường giao ngay và sau đó là hợp đồng tương lai vĩnh viễn. Về mô hình doanh thu, 50% phí từ trình tự (sequencer) sẽ được dùng để mua lại và đốt token $KNTQ, 25% dành cho các nhà phát triển ứng dụng trên Elysium và 25% còn lại chuyển vào kho bạc của Kinetiq. Sản phẩm chính của Kinetiq vẫn là kHYPE, một token đại diện cho $HYPE được stake và có thể sử dụng trong DeFi. Hyperliquid gần đây đã trở thành một trong ba sàn giao dịch hợp đồng tương lai vĩnh viễn lớn nhất.

cryptonews.ru36 phút trước

Đội ngũ Kinetiq công bố mạng L2 Elysium dành cho Hyperliquid

cryptonews.ru36 phút trước

Hyperliquid Policy Center kêu gọi các cơ quan quản lý Hoa Kỳ tạo khung pháp lý cho hợp đồng tương lai vĩnh viễn

Trung tâm Chính sách Hyperliquid kêu gọi SEC và CFTC của Mỹ tạo khuôn khổ pháp lý cho các hợp đồng tương lai vĩnh viễn (perpetual futures). Tổ chức phi lợi nhuận này, thành lập vào tháng 2/2026 với mục tiêu thúc đẩy tài chính phi tập trung (DeFi), đề xuất thay đổi cách phân loại các hợp đồng này, đặc biệt là phái sinh trên chứng khoán, để đơn giản hóa việc niêm yết và giao dịch. Trọng tâm đề xuất là việc tài sản cơ sở nên xác định thẩm quyền giám sát, không làm thay đổi bản chất kinh tế của hợp đồng. Hyperliquid cho rằng các hợp đồng tương lai vĩnh viễn trên cổ phiếu nên được phân loại là hợp đồng tương lai thông thường, chịu sự giám sát chung của CFTC và SEC, thay vì là hoán đổi dựa trên chứng khoán (SBS). Các lập luận ủng hộ bao gồm tính chất tiêu chuẩn hóa, khả năng thanh khoản bằng giao dịch đối ứng, giá công khai và việc không làm phát sinh quyền sở hữu tài sản cơ sở. Họ kiến nghị SEC và CFTC ban hành hướng dẫn chung, áp dụng cách tiếp cận nhất quán cho các hợp đồng trên nhiều loại tài sản (như Bitcoin, dầu, vàng, cổ phiếu) và hiện đại hóa khung pháp lý. Tuy nhiên, đề xuất này vấp phải sự phản đối từ các sàn giao dịch truyền thống như CME Group và ICE. Các tập đoàn này cho rằng các nền tảng như Hyperliquid chịu các yêu cầu quản lý lỏng lẻo hơn và đã kêu gọi siết chặt quy định. Đáng chú ý, CME Group thậm chí đã kiện CFTC liên quan đến việc phân loại các hợp đồng tương lai vĩnh viễn, cho rằng chúng nên được xếp vào danh mục hoán đổi để áp dụng các quy định nghiêm ngặt hơn về ký quỹ, thanh toán bù trừ và vốn.

cryptonews.ru37 phút trước

Hyperliquid Policy Center kêu gọi các cơ quan quản lý Hoa Kỳ tạo khung pháp lý cho hợp đồng tương lai vĩnh viễn

cryptonews.ru37 phút trước

Giao dịch

Giao ngay
活动图片