The two new Claude models—Marshmallow and Melon—that were just exposed yesterday have real-world test results appearing today!
Some have found through testing that the two new models are terrifyingly powerful in 3D reinforcement learning and architectural spatial layout.

Furthermore, their consumption of computing power is also incredibly intense.
Before giving an answer, they consume massive amounts of "Thinking Tokens" for deep logical reasoning, repeatedly even pushing against the system's usage limits!
Judging from this model, Claude is evolving into the next-generation AI giant possessing "slow thinking" and deep spatial physical reasoning capabilities.

One-Shot Mastery of 3D Reinforcement Learning and Architectural Layout
Some testers gave this evaluation after testing: Anthropic is really pushing 3D reinforcement learning like crazy.
The spatial layout of buildings is just ridiculously good, and all of this is generated in one step (One-Shot)!


You must know, for LLMs, spatial imagination has always been a fatal weakness.
Getting an AI to understand the physical coordinate relationships and gravity laws of tables and chairs in a 3D room, or planning the layout of a complex building cluster with intricate traffic flows and load-bearing structures, is as difficult as climbing to the sky.
But this time, Claude's "Marshmallow" and "Melon" seem to have crossed this chasm.
They can directly understand and generate complex 3D coordinates and spatial relationships.
Moreover, testing found their architectural layouts are exceptionally outstanding, indicating the new models also perform powerfully when dealing with topological, geometric, and physical constraints.
And, they output in One-Shot, without needing humans to repeatedly adjust prompts to correct errors; after deep internal deliberation, the model can output perfect 3D scene data in one go.
This hides immense commercial and industrial value.
Imagine, future game developers only need to describe in natural language: "Help me generate a medieval-style fortress, containing three defense towers and a hidden underground passage," and Claude can directly output perfect 3D model code or layout parameters.
Architects can directly let the AI generate dozens of preliminary layout schemes that conform to mechanical structures based on terrain and lighting conditions. The construction of the metaverse, industrial digital twins, autonomous driving simulation environments...
All of this will be completely reshaped by this powerful 3D generation capability!
The Double Whammy of Marshmallow and Melon
Yesterday, the whole internet was flooded with marshmallow and melon.
The internal codenames for the two models exposed this time are claude-marshmallow-eap (Marshmallow) and claude-melon-eap (Melon).

This naming convention of "food + EAP" continues their tradition from claude-horchata-eap (Horchata, a Mexican drink), which was briefly tested in July of this year.
Where did these two models suddenly come from?
According to data intercepted by developers from API traffic logs, during the period of 00:45-00:57Z on August 21, the ID claude-marshmallow-ht-eap was called a high frequency of 57 times in API traffic!
Subsequently, it prominently appeared in the Claude Code model list as a "Custom model" with an astonishing 1 million Token ultra-long context window.
Although these two models currently don't have first-party API endpoints and seem limited to red team testers or internal core personnel, the early sporadic test feedback is explosive enough.



"Marshmallow" is considered slightly stronger in overall capability than "Melon."
In daily conversation and logical interaction, the experience of "Marshmallow" is even better than the current Opus 5, with dialogue appearing more natural and pleasant.
Although their current performance still seems below Anthropic's top-tier "Fable" level, are they the legendary Opus 5.1? Sonnet 5.1? Or a new iteration of Haiku?
There is no conclusion yet.

Computing Power Devourer: The Crazy "Thinking Tokens"
Not only that, multiple testers have discovered an extremely anomalous phenomenon: when using these two new models, their consumption of computing power reaches insane levels!
A tester exclaimed on X: "One thing I noticed about both these models, they use massive amounts of thinking tokens to the point where I directly triggered the limit multiple times during testing!"
What are "Thinking Tokens"? Why are they so important?
In past large model interactions, AI often acted like a "fast-talking" respondent, generating answers word by word instantly based on probability distributions.
Claude's new generation of models has completely changed this logic. Before finally providing an answer, the model generates a massive amount of "invisible" tokens in the background for extremely deep self-deduction, chain-of-thought construction, and logical trial and error.

This "regardless of cost" pouring of computational power releases an extremely strong signal: Anthropic is secretly tackling the bottleneck of AI deep reasoning!
This also confirms previous industry speculation—the competition among large models is shifting from "computing power stacking in the training phase" to "computing power consumption in the inference phase."
When Claude starts thinking so frantically that it hits the limit, the quality of the answers it outputs will form a dimensional blow against traditional models.
Suspected Emergency Firefighting? Opus 5's "Revenge Battle"
The sudden leak of these two mysterious models is also timed quite subtly.
Just less than a month ago, on July 24, 2026, Anthropic released the highly anticipated flagship model Opus 5.
Before that, Sonnet 5, released on June 30, was widely praised.
However, Opus 5 suffered a disastrous defeat.
A developer mocked mercilessly on X: "The negative feedback on Opus 5 is so severe that Anthropic had to immediately roll out two new models to try and replace it... I'm dying of laughter."
Indeed, Opus 5 seems to have failed to meet user expectations for a "cross-generational leap." For Anthropic, which is eager to go public, this is undoubtedly a major blow.
Therefore, the exposure of "Marshmallow" and "Melon" at this critical juncture has sparked infinite speculation from the outside world.
Speculation One: Opus 5.1's Last Stand Counterattack.
Many believe these two models, closely associated with Opus 5.1, are precisely the "power-enhanced versions" that Anthropic urgently retooled and rebuilt to address Opus 5's flaws.
By introducing a powerful "Thinking Tokens" mechanism, the logical ceiling of the model is forcibly elevated.
Speculation Two: A Surprise Attack by the New Generation Sonnet / Haiku.
Some testers also think that, considering their astonishing speed and cost-effectiveness, they might be updated versions of Sonnet or Haiku, after all, they haven't been crowned with the top-tier "Fable" title.

But regardless of which scenario, Anthropic is clearly getting restless.
They are frantically accelerating the iteration speed, even to the point of directly releasing cutting-edge models that extremely consume computing power for testing. X
Developers can no longer contain their excitement: "The next few weeks to months are going to get extremely interesting!"
Can Marshmallow and Melon wash away the shame of Opus 5's predecessor and let Anthropic win another round?
References:
https://x.com/Lentils80/status/2091704307863142812?s=20
https://x.com/NFT_Chen/status/2091764673767198730?s=20
This article is from the WeChat public account "New Zhiyuan", author: ASI Revelation; Editor: Aeneas





