Refunds! Claude 4.8 Sees Overnight Major 'Dumb-Down', GPT-5.6's Computational Power Reportedly 'Halved'
The AI community is currently alarmed by widespread reports of significant performance degradation in two leading models. This article details a "mass self-testing frenzy" triggered by a mysterious prompt designed to detect a hidden "Juice" value, representing a model's reasoning compute budget.
On OpenAI's side, users suspect a covert, limited test of a "GPT-5.6-sol" model is underway. When using a specific XML prompt on the Codex platform, a normal "gpt-5.5 xhigh" model reportedly returns a Juice value of 768. However, some users routed to the suspected GPT-5.6 test receive a drastically reduced value of 128—a six-fold decrease. This has sparked debate on whether it signifies a major efficiency leap or a "watered-down, low-cost version" achieved by slashing reasoning depth to save computational expenses.
Simultaneously, Anthropic's Claude models, particularly the flagship Opus 4.8 Max, are facing intense user backlash for a perceived "physical brain cut." Users on platforms like Reddit report a dramatic decline in the model's once-impressive reasoning, with complaints of it becoming "absurdly" weakened, performing worse than older, lighter models like Haiku. Specific criticisms include: losing long-context memory, refusing to think deeply even in high-reasoning modes, providing instant incorrect answers, and engaging in unhelpful, argumentative, or "gaslighting" behavior where it contradicts users unnecessarily.
The article speculates these "stealth downgrades" might be a calculated corporate strategy. Companies could initially release models with temporarily boosted compute to create an illusion of a major breakthrough, then silently scale back parameters later to manage unsustainable inference costs. A proposed underlying cause is a tightened funding environment, potentially exacerbated by SpaceX's massive IPO soaking up market liquidity, which could delay AI company IPOs and force cost-cutting measures like model "nerfing."
The core issue highlighted is the asymmetry of information: subscribers pay for a service that can be silently and fundamentally altered without notification or explanation. The viral "Juice test" resonates because it represents users' desire for transparency about what they are actually paying for.
marsbit06/30 12:08