Claude's Two New Models Leaked in Real-World Tests, Spatial Reasoning is Divine, Computing Power Directly Dried Up

marsbitОпубліковано о 2026-08-25Востаннє оновлено о 2026-08-25

Анотація

Summary: Claude's two new leaked models, "Marshmallow" and "Melon," demonstrate breakthrough capabilities in 3D spatial reasoning and architectural layout generation, performing complex tasks in a single attempt (one-shot). However, they consume massive amounts of computational power through extensive "thinking tokens" for deep internal reasoning, often nearing system limits. Their sudden emergence follows the reportedly disappointing release of Opus 5, leading to speculation they are either enhanced Opus 5.1 versions or next-generation Sonnet/Haiku models. These developments signal Anthropic's intense focus on advancing AI's deep reasoning and 3D understanding, with significant potential implications for gaming, architecture, and digital twin simulations.

The two new Claude models—Marshmallow and Melon—that were just exposed yesterday have real-world test results appearing today!

Some have found through testing that the two new models are terrifyingly powerful in 3D reinforcement learning and architectural spatial layout.

Furthermore, their consumption of computing power is also incredibly intense.

Before giving an answer, they consume massive amounts of "Thinking Tokens" for deep logical reasoning, repeatedly even pushing against the system's usage limits!

Judging from this model, Claude is evolving into the next-generation AI giant possessing "slow thinking" and deep spatial physical reasoning capabilities.

One-Shot Mastery of 3D Reinforcement Learning and Architectural Layout

Some testers gave this evaluation after testing: Anthropic is really pushing 3D reinforcement learning like crazy.

The spatial layout of buildings is just ridiculously good, and all of this is generated in one step (One-Shot)!

You must know, for LLMs, spatial imagination has always been a fatal weakness.

Getting an AI to understand the physical coordinate relationships and gravity laws of tables and chairs in a 3D room, or planning the layout of a complex building cluster with intricate traffic flows and load-bearing structures, is as difficult as climbing to the sky.

But this time, Claude's "Marshmallow" and "Melon" seem to have crossed this chasm.

They can directly understand and generate complex 3D coordinates and spatial relationships.

Moreover, testing found their architectural layouts are exceptionally outstanding, indicating the new models also perform powerfully when dealing with topological, geometric, and physical constraints.

And, they output in One-Shot, without needing humans to repeatedly adjust prompts to correct errors; after deep internal deliberation, the model can output perfect 3D scene data in one go.

This hides immense commercial and industrial value.

Imagine, future game developers only need to describe in natural language: "Help me generate a medieval-style fortress, containing three defense towers and a hidden underground passage," and Claude can directly output perfect 3D model code or layout parameters.

Architects can directly let the AI generate dozens of preliminary layout schemes that conform to mechanical structures based on terrain and lighting conditions. The construction of the metaverse, industrial digital twins, autonomous driving simulation environments...

All of this will be completely reshaped by this powerful 3D generation capability!

The Double Whammy of Marshmallow and Melon

Yesterday, the whole internet was flooded with marshmallow and melon.

The internal codenames for the two models exposed this time are claude-marshmallow-eap (Marshmallow) and claude-melon-eap (Melon).

This naming convention of "food + EAP" continues their tradition from claude-horchata-eap (Horchata, a Mexican drink), which was briefly tested in July of this year.

Where did these two models suddenly come from?

According to data intercepted by developers from API traffic logs, during the period of 00:45-00:57Z on August 21, the ID claude-marshmallow-ht-eap was called a high frequency of 57 times in API traffic!

Subsequently, it prominently appeared in the Claude Code model list as a "Custom model" with an astonishing 1 million Token ultra-long context window.

Although these two models currently don't have first-party API endpoints and seem limited to red team testers or internal core personnel, the early sporadic test feedback is explosive enough.

"Marshmallow" is considered slightly stronger in overall capability than "Melon."

In daily conversation and logical interaction, the experience of "Marshmallow" is even better than the current Opus 5, with dialogue appearing more natural and pleasant.

Although their current performance still seems below Anthropic's top-tier "Fable" level, are they the legendary Opus 5.1? Sonnet 5.1? Or a new iteration of Haiku?

There is no conclusion yet.

Computing Power Devourer: The Crazy "Thinking Tokens"

Not only that, multiple testers have discovered an extremely anomalous phenomenon: when using these two new models, their consumption of computing power reaches insane levels!

A tester exclaimed on X: "One thing I noticed about both these models, they use massive amounts of thinking tokens to the point where I directly triggered the limit multiple times during testing!"

What are "Thinking Tokens"? Why are they so important?

In past large model interactions, AI often acted like a "fast-talking" respondent, generating answers word by word instantly based on probability distributions.

Claude's new generation of models has completely changed this logic. Before finally providing an answer, the model generates a massive amount of "invisible" tokens in the background for extremely deep self-deduction, chain-of-thought construction, and logical trial and error.

This "regardless of cost" pouring of computational power releases an extremely strong signal: Anthropic is secretly tackling the bottleneck of AI deep reasoning!

This also confirms previous industry speculation—the competition among large models is shifting from "computing power stacking in the training phase" to "computing power consumption in the inference phase."

When Claude starts thinking so frantically that it hits the limit, the quality of the answers it outputs will form a dimensional blow against traditional models.

Suspected Emergency Firefighting? Opus 5's "Revenge Battle"

The sudden leak of these two mysterious models is also timed quite subtly.

Just less than a month ago, on July 24, 2026, Anthropic released the highly anticipated flagship model Opus 5.

Before that, Sonnet 5, released on June 30, was widely praised.

However, Opus 5 suffered a disastrous defeat.

A developer mocked mercilessly on X: "The negative feedback on Opus 5 is so severe that Anthropic had to immediately roll out two new models to try and replace it... I'm dying of laughter."

Indeed, Opus 5 seems to have failed to meet user expectations for a "cross-generational leap." For Anthropic, which is eager to go public, this is undoubtedly a major blow.

Therefore, the exposure of "Marshmallow" and "Melon" at this critical juncture has sparked infinite speculation from the outside world.

Speculation One: Opus 5.1's Last Stand Counterattack.

Many believe these two models, closely associated with Opus 5.1, are precisely the "power-enhanced versions" that Anthropic urgently retooled and rebuilt to address Opus 5's flaws.

By introducing a powerful "Thinking Tokens" mechanism, the logical ceiling of the model is forcibly elevated.

Speculation Two: A Surprise Attack by the New Generation Sonnet / Haiku.

Some testers also think that, considering their astonishing speed and cost-effectiveness, they might be updated versions of Sonnet or Haiku, after all, they haven't been crowned with the top-tier "Fable" title.

But regardless of which scenario, Anthropic is clearly getting restless.

They are frantically accelerating the iteration speed, even to the point of directly releasing cutting-edge models that extremely consume computing power for testing. X

Developers can no longer contain their excitement: "The next few weeks to months are going to get extremely interesting!"

Can Marshmallow and Melon wash away the shame of Opus 5's predecessor and let Anthropic win another round?

References:

https://x.com/Lentils80/status/2091704307863142812?s=20

https://x.com/NFT_Chen/status/2091764673767198730?s=20

This article is from the WeChat public account "New Zhiyuan", author: ASI Revelation; Editor: Aeneas

Пов'язані питання

QWhat are the two new Claude models mentioned in the article, and what are their internal code names?

AThe two new Claude models mentioned are internally codenamed 'claude-marshmallow-eap' (Marshmallow) and 'claude-melon-eap' (Melon).

QAccording to the article, in which specific areas do the new Claude models demonstrate exceptional performance?

AThe new Claude models demonstrate exceptional performance in 3D reinforcement learning and building spatial layout tasks. They excel at generating complex 3D coordinates, understanding spatial relationships, and creating architecturally sound layouts in a single attempt (One-Shot).

QWhat is a 'thinking token' as described in the context of the new Claude models, and why is it significant?

AA 'thinking token' refers to the invisible, intermediate tokens the new Claude models generate internally before producing a final answer. They are used for deep self-reasoning, chain-of-thought construction, and logical trial-and-error. This is significant because it represents a shift towards AI with 'slow thinking' and deep reasoning capabilities, consuming significant computational power for higher quality outputs.

QWhat is one potential commercial or industrial application suggested for the new models' 3D spatial reasoning capabilities?

AOne potential application is for game developers to use natural language prompts to generate complete 3D model code or layout parameters for complex structures like a medieval fortress with defense towers and hidden tunnels, drastically speeding up development.

QWhat is one speculation in the article about why these two new models were leaked shortly after the release of Opus 5?

AOne speculation is that these models, potentially Opus 5.1 variants, are an emergency response or 'powered-up version' from Anthropic to address the negative reception and perceived shortcomings of the recently released Opus 5 model.

Пов'язані матеріали

Demand Test: Bitcoin Whales Earn Record $1.2 Billion, Ethereum Holders Return to Profitability

In just three days after Bitcoin's price recovery, new Bitcoin whales have realized over $1.2 billion in profit, marking the largest profit-taking event for this cohort on record according to CryptoQuant. The peak occurred on August 20 with roughly $614 million, setting a daily record. Analysts note this selling pressure began after Bitcoin rose above the realized price of short-term whales, which was around $68,900. With Bitcoin trading near $77,700 on August 23, these whales were sitting on an average profit of about 12.8%. The market recovery allowed investors who were previously at breakeven or at a loss to lock in gains. CryptoQuant described the situation as a key test for Bitcoin demand; sustained prices above whale cost-basis with normalized profit-taking could signal strong new demand, while continued selling pressure could turn the rally into a mere break-even exit. Simultaneously, large Ethereum holders have also returned to an unrealized profit zone following its rally, as noted by CryptoQuant analyst Darkfost. Current profit levels, however, remain relatively low and are not seen as creating significant selling pressure. The Unrealized Profit/Loss Ratio for different whale cohorts stands at 0.075 (for 1k-10k ETH holders), 0.16 (10k-100k ETH), and 0.38 (over 100k ETH). This marks a significant improvement from June, when these whales were in substantial unrealized loss, with Ethereum having risen over 65% since then. The increased profitability is viewed as a positive sentiment signal for the Ethereum market.

cryptonews.ru4 хв тому

Demand Test: Bitcoin Whales Earn Record $1.2 Billion, Ethereum Holders Return to Profitability

cryptonews.ru4 хв тому

Kinetiq Team Announces Elysium L2 Network for Hyperliquid

On August 24, the liquid staking protocol Kinetiq announced Elysium, a new L2 network for the Hyperliquid ecosystem. It aims to increase HyperEVM's throughput and simplify the launch of spot markets, tokens, and DeFi applications. Elysium will use $HYPE for gas fees and plans direct integration with HyperCore's trading engine, giving apps access to its liquidity and orderbook data. Technical details and partners will be revealed later, with a launch date set for "soon." A primary reason for Elysium's development is to overcome HyperEVM's limitations, such as low throughput and rising fees during high load. Kinetiq claims the L2 will start with significantly higher block production and transaction speeds, later aiming to approach HyperCore's performance. This targets applications needing frequent state updates like high-frequency spot trading and AMMs, and will provide them deeper access to HyperCore orderbook data. The network also proposes to streamline the process of launching new assets within Hyperliquid, allowing a token to progress from AMM liquidity to HyperCore's spot orderbook and eventually to perp markets via HIP-3 in a unified flow. Regarding revenue, Kinetiq's model allocates 50% of sequencer fees to buy back and burn $KNTQ, 25% to developers using Elysium's block space, and 25% to the Kinetiq treasury. Elysium marks Kinetiq's expansion beyond its core liquid staking product, kHYPE.

cryptonews.ru5 хв тому

Kinetiq Team Announces Elysium L2 Network for Hyperliquid

cryptonews.ru5 хв тому

Торгівля

Спот
活动图片