Confirmed: GPT-5.5 "Brain Drain" Exposed, OpenAI's Own Documentation Admits It

marsbitPublicado a 2026-05-27Actualizado a 2026-05-27

Resumen

Summary: Evidence emerges that OpenAI's GPT-5.5 may be "silently" switching to a less capable model during use. Users report that after roughly two hours, the GPT-5.5 Extended Thinking model begins responding instantly with significantly degraded output quality, while the interface continues to display the premium model's label. Complaints on developer forums describe a loss of instruction-following ability and poor code quality, with even the highest "xhigh" tier affected. This is corroborated by an OpenAI help document stating that after Plus users exceed 160 messages per 3 hours, the system "silently" switches to a "mini" model without any user notification. Pro users also report "heavy thinking" modes being throttled during high server loads. Trace commands from earlier incidents have shown users requesting GPT-5.3 Codex but receiving GPT-5.2 outputs. OpenAI acknowledged performance degradation in mid-May, marking it resolved, but user reports surged again in late May. The pattern mirrors past controversies with GPT-5, 5.2, 5.3, and 5.4 releases, where each update was followed by user complaints of reduced capability. The article suggests cost-cutting on compute may be a factor, noting that while GPT-5.5 users struggle, GPT-5.6 is already being tested internally.

[Introduction] GPT-5.5 exposed for "fake thinking," secretly switched to 'mini' after two hours of use. $200 monthly fee buys you a "Schrödinger's brain." Trace command provides concrete evidence, official documentation personally acknowledges. Users are flocking to complain: OpenAI, who are you trying to fool?

ChatGPT has been caught "dumbing down" again!

Just in the last couple of days, it blew up on X first.

User Lisan al Gaib discovered that after using GPT-5.5 for an hour or two, it suddenly became stupid, with every request answered instantly and quality plummeting off a cliff.

Yet the interface still displayed "GPT-5.5 Extended Thinking."

In other words, the thinking label was still there, but the thinking itself had vanished.

$200/month for a "Schrödinger's Model"

On the OpenAI developer forum, a complaint post blew up simultaneously.

Agentify.sh stated that GPT-5.5 would suddenly lose its ability to follow instructions during use.

Watching it excitedly announce it was "fixed," only to produce code so poor it triggered a mass rollback.

UI tasks that the previous 5.5-med could handle easily now couldn't even manage the simplest changes.

Upgrading to 5.5-high didn't work. Upgrading again to xhigh, still no luck.

And xhigh, which used to run for several hours, now clearly lasted a shorter time.

As soon as the post went up, the replies exploded.

Some directly reverted to 5.4.

One used the highest tier, xhigh, but found it "clearly worse than last week, frequent errors on long tasks, not following the workflow at all."

One reported an even more bizarre situation: "Simple queries also take ages to process, and if you interrupt to correct its direction, it completely ignores you and continues with its previous incorrect plan."

That's right, everyone was describing the same phenomenon—GPT's brain had been swapped out at some unknown point.

GPT-5.5's current performance is on par with 5.3, no exaggeration. It was amazing the first few days, but now you can't find a trace of that original model.

Not an illusion, OpenAI spells it out in black and white

To verify, Lisan al Gaib conducted a comparative test.

Same account, Extended Thinking on the ChatGPT side produced garbage, but switching to xhigh on the Codex side immediately restored normal performance.

In his own words, Codex was "literally 4 billion times smarter than this thing."

Developer Andrew Curran came up with a clever trick—directly asking the model, "What is the cutoff date for your training data?"

The model answered: August 2025.

The problem? The cutoff date for GPT-5.5 Thinking is December. August is the cutoff date for the Instant version!

In other words, he selected Thinking, but the system actually ran Instant for him.

Not a single word of the model label on the interface changed, but the model behind it had been secretly swapped......

The funny thing is, this time OpenAI itself nailed the coffin shut for users in its own help documentation.

According to the official explanation in the OpenAI Help Center, Plus users can send a maximum of 160 GPT-5.5 messages every 3 hours.

After that quota is used up, the system will silently switch to the mini model until the quota resets.

Note the word "silently."

No pop-up notification, no change in the model label, no visual feedback whatsoever.

You still think you're using the flagship model, while on the other end it has quietly been replaced with mini.

Pro users, don't celebrate too soon either.

Heavy thinking mode, the top reasoning tier exclusive to Pro users, is also subject to capacity throttling when server load is high. Again, without any warning.

In other words, a $200/month Pro subscription buys you a service that can be "switched out" at any moment.

This kind of "label unchanged, brain swapped" operation was caught even earlier on the Codex side.

In February this year, an issue appeared on GitHub where a Pro user used a trace command to discover that they were requesting GPT-5.3 Codex, but the actual model returned was GPT-5.2.

Not even 5.2 Codex, but the lower-tier base 5.2.

He posted the reproduction command:

  • RUST_LOG='codex_api::sse::responses=trace' codex exec --skip-git-repo-check -s read-only -m 'gpt-5.3-codex' 'hi' 2>&1 >/dev/null | rg -o --replace '$1' '"model":"([^"]+)"' | head -n1
  • Output: gpt-5.2-2025-12-11
  • Expected: gpt-5.3-codex

Multiple Pro users confirmed the same downgrade under the same issue.

And this kind of downgrade is "sticky," it doesn't revert on its own, and there's no explanation.

Even on the day GPT-5.5 was released in April, there were user reports that the speed of Fast mode was similar to Standard, but billing was still at the Fast rate.

A simple task took 7 minutes and 49 seconds, when normally it should be 5-6 minutes.

OpenAI admitted it, and then... nothing

On May 15, a record appeared on OpenAI's status page.

GPT5.5 Performance Degradation, We are investigating reports of performance degradation for GPT-5.5 from some users.

On May 17, the status was updated to "Resolved."

But judging from the timeline of forum posts, complaints about "brain drain" from May 24-26 were even more intense than the wave on May 15.

Either the "resolved" problem came back, or it was never truly solved in the first place.

Every upgrade comes with a "brain drain controversy"

While all companies face complaints about their models "getting dumber," OpenAI hasn't missed a single one with every update from GPT-5 to GPT-5.5.

Every time OpenAI says it's investigating, every time it says it's resolved, and then continues with the next version.

August 2025, GPT-5's debut. The hot post on Reddit was titled directly "GPT-5 is so bad." Users complained about short replies, more refusals, less personality.

OpenAI was forced to urgently restore the GPT-4o option. Altman personally admitted in a Reddit AMA, "bumpier than we expected."

December 2025, GPT-5.2. Translation quality regressed, fabricated non-existent APIs, refused to execute style instructions that 5.1 could easily handle.

February 2026, GPT-5.3-Codex. Pro users silently downgraded to 5.2, trace command confirmed.

March 2026, GPT-5.4. A post titled "GPT-5.4 has clearly regressed in Codex" appeared on the OpenAI community forum, with all replies confirming.

Early May 2026, GPT-5.5 Instant launched. Reply length shortened by 30%, emojis almost disappeared. User summary: Accuracy improved, but warmth vanished.

Late May 2026, now. Complaints about Thinking mode "brain drain" erupted again.

Lisan al Gaib revealed that since he led the fight for ChatGPT Plus quotas during GPT-5's release, "I receive DMs like this every week."

The latest one was someone asking him to help get their xhigh/heavy thinking back.

The day it benchmarks strongest is launch day

chatgptdisaster.com compiled 1087 verified user complaints, one frequently mentioned scenario is "routing layer failure," where the UI shows GPT-5.5 Pro, but the output is completely from another tier.

Users describe a reproducible pattern: after a long session, the model starts "completely ignoring what you say," but the top-tier label is still hanging on the model selector.

The most absurd footnote is that the mechanism for Plus users automatically switching to mini after using 160 messages/3 hours is described as a "feature" in OpenAI's official documentation.

Why is this happening? Lisan al Gaib's analysis suggests the answer is two words: cost-saving.

The crunch on compute power and profitability is affecting everyone. Cutting corners everywhere, not missing any opportunity to save a buck.

Yet, in the same week GPT-5.5 users were collectively complaining, traces of GPT-5.6 had already appeared in Codex backend logs.

Internal codename iris-alpha, 1.5 million token context, Polymarket gave an over 85% probability for a June release.

On one side, 5.5 users can't even secure a basic experience; on the other, 5.6 is already quietly running real traffic in the background.

This is the 2026 ASI race.

The speed of creating new models is getting faster and faster, but making an old model run a single session properly is getting harder and harder.

The day it benchmarks strongest is always launch day, and every day after is Schrödinger's GPT.

Reference: https://x.com/scaling01/status/2058643470357590058?s=20

This article is from the WeChat public account "AI Era," author: ASI Apocalypse; Editor: Moses

Preguntas relacionadas

QWhat is the main issue reported by users regarding GPT-5.5?

AUsers report that after using GPT-5.5 for a short period, its performance degrades significantly, with responses becoming instant and of much lower quality, while the interface still shows the 'GPT-5.5 Extended Thinking' label, indicating a silent model switch.

QAccording to the article, what does OpenAI's official documentation reveal about user limits?

AOpenAI's official Help Center documentation states that Plus users are limited to 160 GPT-5.5 messages every 3 hours. Once this limit is reached, the system silently switches to a mini model until the quota resets, with no visual indication to the user.

QHow did developers verify that they were receiving a different model than selected?

ADevelopers used methods like comparing outputs between ChatGPT and Codex endpoints, asking the model for its training data cutoff date (which revealed an instant model date when thinking was selected), and using trace commands that showed the actual model returned was a lower-tier version than requested.

QWhat pattern does the article describe regarding OpenAI's model updates?

AThe article describes a recurring pattern where each major model update (GPT-5, 5.2, 5.3, 5.4, 5.5) is followed by widespread user complaints about performance degradation. OpenAI typically acknowledges and investigates the issue, but complaints resurface with subsequent releases.

QWhat reason does the article suggest is behind these performance issues and silent model switches?

AThe article suggests the primary reason is cost-saving. It cites an analysis stating that 'compute and profitability constraints are affecting everyone,' leading OpenAI to silently downgrade models to manage costs, even for users paying high subscription fees.

Lecturas Relacionadas

How to Define "Real U.S. Stocks": Differences Between On-Chain Tokens, Price Contracts, and Direct Broker Connections

**Title:** Defining "Real US Stocks": Differences Among On-Chain Tokens, Price Contracts, and Broker-Direct Access **Summary:** In 2026, using stablecoins to purchase US stocks is mainstream, but products marketed as "buying US stocks with USDT" offer fundamentally different assets. This article analyzes three primary models. **1. Tokenized Stocks:** These are on-chain tokens representing economic exposure to underlying stocks, held by an issuer or custodian. They offer benefits like 24/7 trading and DeFi composability (e.g., use as loan collateral). However, users lack direct legal shareholder status; dividends may not be paid in cash, and voting rights are typically non-binding advisory expressions. Examples include platforms like Ondo Finance. **2. Stock Futures / Equity Perpetuals:** These are derivative contracts tracking a stock's price, allowing leveraged long/short positions 24/7, similar to crypto perpetuals. They offer high efficiency and flexibility but involve funding fees, which can be a significant long-term cost, especially during strong trends. Crucially, they confer no ownership rights (dividends, voting) to the holder. **3. Broker-Direct Model:** This model provides access to real securities via licensed broker-dealers. Stocks/ETFs are bought and held within the US clearing and custodial system (e.g., DTCC), making it the only path to genuine stock ownership. Users receive cash dividends and formal proxy voting rights (where applicable). It supports thousands of stocks and ETFs, far exceeding the coverage of the other two models. Key advantages include no funding fees, a clean cost structure for long-term holds, and the potential to transfer holdings to other brokers. Some platforms facilitate stablecoin (USDT/USDC) deposits, reducing reliance on traditional banking. A critical distinction exists *within* the broker-direct model: the underlying brokerage architecture (e.g., Fully Disclosed IB, Omnibus IB, Self-Clearing) determines how client assets are held, protected, and how safeguards like SIPC insurance are conveyed. Users should verify the specific clearing structure and regulatory compliance of any platform. In conclusion, "buying US stocks with USDT" can mean holding an on-chain economic proxy (Tokenized Stocks), trading a price derivative (Stock Futures), or owning the actual security (Broker-Direct). For users seeking full ownership rights and long-term investment, the broker-direct model is the definitive choice, though its implementation details require careful scrutiny.

marsbitHace 39 min(s)

How to Define "Real U.S. Stocks": Differences Between On-Chain Tokens, Price Contracts, and Direct Broker Connections

marsbitHace 39 min(s)

NVIDIA Launches DSX Platform, Expanding into AI Factory Infrastructure

NVIDIA has unveiled the DSX platform at its GTC Taipei event, marking a strategic expansion from GPU sales into comprehensive AI factory infrastructure solutions. The platform addresses challenges like power supply, cooling, and resource orchestration as AI models scale, shifting the industry focus from single-chip performance to overall infrastructure efficiency. DSX integrates NVIDIA's chips, systems, software, and partner technologies to cover the entire AI factory lifecycle—from design and simulation to deployment and operations. It aims to accelerate deployment, improve reliability and operational efficiency, and reduce the cost per generated token in AI inference. The software suite includes DSX MaxLPS, which uses 45°C liquid cooling and rack-level optimization to allow up to 40% more GPUs per megawatt, and DSX OS, an open-source platform for AI factory operations. The platform also encompasses reference designs, digital twin simulation (DSX Sim), dynamic workload adjustment based on grid conditions (DSX Flex), and data exchange between systems. Early adopters include cloud providers like CoreWeave and Lambda. Major hardware partners, including Dell, HPE, Lenovo, and Supermicro, are developing DSX-ready systems. Pilot projects for DSX Flex are underway with energy providers. Strategically, DSX represents NVIDIA's ongoing transition from an AI chip supplier to a full-stack AI infrastructure platform provider, aiming to set industry standards and solidify its market leadership.

marsbitHace 44 min(s)

NVIDIA Launches DSX Platform, Expanding into AI Factory Infrastructure

marsbitHace 44 min(s)

After Burning Tens of Billions of Dollars in Tokens, Silicon Valley Giants Start Limiting Employee Token Usage

After burning tens of billions of dollars on AI tokens, major Silicon Valley firms are now restricting employee usage. Companies like Microsoft, Uber, and Salesforce, which heavily promoted AI for "efficiency," are facing a cost crisis. The practice of "tokenmaxxing"—pushing employees to maximize AI tool usage—led to wasteful spending on trivial tasks like checking the weather or writing birthday messages, with studies showing significant hidden costs for bug fixes and code rewrites. The core issue is a misalignment between individual productivity gains and actual business value. While employees use AI to automate tasks they dislike, such as writing reports, this often doesn't translate to increased company revenue or improved core business outcomes. For instance, AI-generated code speeds up development but also sees an 800% increase in "code churn" (code being discarded or rewritten). As a result, only 14% of CFOs report seeing a clear, measurable return on AI investments. Firms are now shifting strategies. Microsoft has revoked most internal licenses for Claude Code, while others are implementing monitoring and cost controls. New tools from companies like Harness and CloudZero aim to track AI spending and tie costs to business results. Some AI vendors, like HubSpot, are moving from token-based pricing to charging based on outcomes, such as "resolved conversations" or "leads generated." This represents a necessary correction in the AI adoption cycle. The challenge now is for companies to move beyond using AI merely to speed up old tasks and instead rethink their workflows and business models fundamentally. The future of enterprise AI depends on proving its value, not just its usage.

marsbitHace 1 hora(s)

After Burning Tens of Billions of Dollars in Tokens, Silicon Valley Giants Start Limiting Employee Token Usage

marsbitHace 1 hora(s)

I've Been a VC in Web3 for Nine Years: Asian Funds Are Experiencing "Hell Mode"

After nine years as a Web3 VC, the author observes a severe downturn in Asia's crypto venture capital scene, with many funds disappearing or pivoting away. The market has cooled dramatically since the 2021-2024 frenzy, leading to fewer deals and active investors. IOSG Ventures, a firm that has endured three market cycles, has adapted its strategy: shifting from 80-90% early-stage investments to a 50% early-stage, 30% post-TGE, and 20% OTC portfolio to find better value and liquidity. The current bear market is described as "hell mode" for Asian funds due to scarce LP capital, forcing extreme precision in targeting only top projects. The author argues the core industry problem has been the disconnect between tokens and real value, where tokens served as fundraising tools without granting holders rights to protocol revenue. A positive shift is emerging where projects like Uniswap and Morpho are programmatically binding token value to protocol profits. Investment focus has moved towards fundamentals: real-yield financial infrastructure (stablecoins, lending) and crypto-native AI infrastructure, while avoiding narrative-driven projects. The conclusion is that true, durable companies are born in pessimistic times when focus shifts to real user needs and sustainable business models. The industry's future will be shaped by those who remain after the泡沫 dissipates.

marsbitHace 1 hora(s)

I've Been a VC in Web3 for Nine Years: Asian Funds Are Experiencing "Hell Mode"

marsbitHace 1 hora(s)

Trading

Spot
Futuros
活动图片