Anthropic Apologized, But the Business of 'Safety' Hasn't Stopped

marsbitPublished on 2026-06-12Last updated on 2026-06-12

Abstract

On June 11, Anthropic apologized not for a model failure, but for a lack of transparency. Its new Claude Fable 5 model was found to be secretly rerouting requests from users engaged in advanced AI model development to a weaker version, Opus 4.8, without any notification. The company's response—promising future notifications for such "downgrades"—was met with user skepticism. The article argues the core issue isn't technical but commercial: Anthropic's "safety" measures are primarily a business strategy. A key feature, the "intelligent safety classifier," marketed as user protection, is described as a tool for "competitive defense" to protect Anthropic's market lead by limiting rivals' research capabilities. This covert mechanism was designed for low "false positives," precisely targeting AI researchers. Anthropic's model involves a calculated three-step process: publishing alarming security research to amplify public anxiety, offering its Fable 5 model with a "safety classifier" as a premium-priced solution, and cashing in through a planned high-value IPO. This contrasts with OpenAI's more direct "tool-and-traffic" approach. The apology, merely changing a secret downgrade to a visible one, is seen as a business "patch" rather than a principled shift. The incident risks damaging Anthropic's "safest AI" reputation among the developer community, which underpins its valuation and appeal to government and corporate clients. Ultimately, the article concludes that for Anthropic, ...

On June 11th, Anthropic apologized. The model didn't fail; the apology was for "failing to strike the right balance"—the newly released Claude Fable 5 pulled a sneaky trick. If it detected you were using Claude for cutting-edge model development, it would silently divert your request to the weaker Opus 4.8 in the backend.

After being caught red-handed, Anthropic's explanation was bizarre: from now on, they'll notify you before dumbing things down.

The netizen's retort hit the nail on the head: "With this move, are you planning to give a heads-up before changing your tune in the future?"

In reality, the core issue isn't whether the model changed, but that Anthropic's so-called "safety" has, from the start, been a business.

The algorithm's stance always sways with money.

Non-Compete Defense, Disguised as Safety Defense

The incident began when Anthropic launched Fable 5 with an "Intelligent Safety Classifier." The official spin was: it detects high-risk requests, automatically downgrades them, and protects users.

What's high-risk? Anthropic spilled the beans: "To prevent foreign adversaries from using the model to accelerate R&D and protect our own leading advantage."

Users don't need that kind of protection; the liability waiver in the terms of service is enough. What Anthropic really meant was: Using Claude for AI research is stealing their rice bowl. Safety is the packaging; the essence is non-compete defense. In short, it's all strategic knife-work.

What's even more cunning is that this defense mechanism was stealthy. Thankfully, Anthropic finally told the truth in their apology statement: "Invisible safety restrictions allow for more precise targeting of specific objectives, enabling us to deploy quickly with very low false-positive rates."

AI researchers are that precisely targeted group.

Now forced to switch to "visible," it's purely because they got caught. They even preemptively set expectations: making it visible will "inevitably lead to more false positives." Meaning, the experience of ordinary users will have to take the hit.

This rule set was never neutral; it only protects the paymasters.

The Trifecta: Hype, Monetize, Harvest

Anthropic's playbook is more meticulously calculated than their large models themselves.

On June 10th, they first released a safety research paper. They trained a model that could reverse-engineer exploit code for vulnerabilities in a matter of hours, based on security patches. What used to take hackers days or even weeks to weaponize an N-day vulnerability is now compressed to an hour scale. The research itself is solid, but releasing it on the same day as Fable 5's launch changes the flavor: proving AI is very unsafe on one hand, while selling the "safety net solution" on the other.

The "legendary model" Fable 5 is priced at $10 per million input tokens / $50 per million output tokens, a notch pricier than Opus 4.8, with the safety classifier becoming the core premium point. Capital markets played along perfectly. Anthropic's valuation hit $96.5 billion, with plans for an October IPO underwritten by Goldman Sachs and J.P. Morgan. What they're buying isn't model parameters; it's the persona of the "safest AI company."

Research amplifies anxiety, the product harvests the premium, capital cashes out. Three moves flowing with the interests, forming a seamless loop. The only problem was, this time the loop sprung a leak: In their haste to restrict competitors, they forgot the community has people who can test for it.

OpenAI Sells Tools, Anthropic Sells Anxiety

Compared to OpenAI, the approach is completely different.

OpenAI is secretly filing for an IPO, valuation nearing a trillion, pitching the "super app": ChatGPT with 900 million weekly active users, integrating with Visa to build an ecosystem. The logic is straightforward: provide tools, earn traffic. Greedy, but candid.

Anthropic doesn't compete on scale; it competes on irreplaceability. While the whole industry is anxious about safety, it plays the role of the "only responsible adult." Its patrons are governments and giants—these are the ones most afraid of incidents and most willing to throw money at "incident prevention."

Therefore, Anthropic must keep AI perpetually in a Schrödinger's cat state of "dangerous but controllable." Too safe, and the classifier doesn't sell; too dangerous, and clients run scared. The best solution? Keep the power to define "danger" firmly in their own hands.

The dumbing-down incident just exposed this logic taken too far: the boundary of "danger" was pushed to "using Claude for AI R&D." It doesn't matter if your research is harmful; threatening their lead is the original sin.

AI has no values; it's just the boss's business spreadsheet written in code.

Apology, Just After-Sales Service for the Business

What about after the apology? Changing from secretly dumbing down to giving a heads-up before dumbing down.

Netizens see right through it: "Do you really believe it won't secretly lower output quality in the future?"

Trust, once broken, stays broken. Especially when the underlying commercial motive hasn't changed: research still amplifies anxiety, the product still harvests the premium.

The Wall Street Journal reported that OpenAI is considering significant price cuts to snatch clients from Anthropic. Price wars aren't new, but this exposes a hidden truth: The ones being downgraded covertly are AI researchers, damaging reputation among the geek community. B2B clients buying Anthropic aren't buying parameters; they're buying the persona of "the industry's safety expert." Once that persona cracks within the core developer community, why should those government and enterprise clients, who sign contracts paying a "safety premium," continue to believe you're "the safest one"?

Out of that $96.5 billion valuation, how much is solid capability, and how much is performance?

Anthropic's code is honest. The safety classifier always protects the home turf; research is responsible for amplifying anxiety; the product is responsible for harvesting the premium; the IPO is responsible for cashing out. This apology is merely a patch to the system: changing "secretly dumbing down" to "overtly dumbing down."

If safety policies really worked, Anthropic wouldn't need to publish papers every year proving patches can be breached. If the classifier were truly neutral, doing AI R&D wouldn't be classified as high-risk.

The answer was already written in the business logic.

Safety is the best business. Apology is just the after-sales service.

This article is from the WeChat public account "AI Contrarian", author: Changqing

Related Questions

QWhat was the main issue with Anthropic's 'intelligent safety classifier' in the Claude Fable 5 model, according to the article?

AThe main issue was that the safety classifier would silently and automatically downgrade user requests to a weaker model (Opus 4.8) if it detected the user was conducting cutting-edge AI development or research. The article argues this was not truly about user safety but was a form of 'competitive defense' to protect Anthropic's own business advantage.

QHow does the article contrast the business strategies of Anthropic and OpenAI?

AThe article contrasts them by stating OpenAI's strategy is to 'sell tools'—focusing on building a super-app ecosystem (like ChatGPT) and monetizing scale and traffic. Anthropic's strategy is described as 'selling anxiety'—leveraging and amplifying safety concerns to position itself as the indispensable, 'most responsible' AI company for government and enterprise clients, thereby justifying premium pricing.

QWhat three-step business 'playbook' does the article attribute to Anthropic?

AThe article describes Anthropic's playbook as a three-step cycle: 1) Research that amplifies AI safety anxieties (like a paper showing models can quickly weaponize security patches). 2) Product development that harvests a price premium based on claimed safety superiority. 3) Capitalizing on this through high valuation and IPO, creating a closed financial loop.

QWhat does the article suggest is the real consequence of Anthropic's 'silent downgrade' being exposed?

AThe article suggests the real consequence is the erosion of trust, especially within the core developer and AI research community. This damage to its reputation as 'the most safety-conscious company' among technical users could ultimately undermine the 'safety premium' justification for its enterprise and government clients, threatening its business model and high valuation.

QWhat is the article's ultimate conclusion about Anthropic's concept of 'safety'?

AThe article concludes that for Anthropic, 'safety' is primarily a business strategy rather than a neutral, ethical stance. It argues that Anthropic's safety measures, such as the classifier, are designed to serve its commercial interests (like protecting its competitive lead), and that the apology was merely 'after-sales service' for this business, not a change in its underlying commercial logic.

Related Reads

Podcast Notes | Conversation with GSR Asset Management Head: To Determine if This Crypto Rally is Real, Just Watch the Lending Rates on Aave

Podcast Summary: Dialogue with GSR's Head of Asset Management: To Determine if This Crypto Rally is Real, Just Check Lending Rates on Aave Andy Baehr, Managing Director of Asset Management at GSR, discusses the current crypto market, characterizing it as stuck in a state of "ambivalence" with short-lived, unsustainable rallies. He outlines a simple framework: the market moves between "ambivalence" and "conviction" (sustained upward momentum). Currently, every rally resembles a single-stage rocket booster that quickly fizzles out. Baehr identifies three key signals to watch: 1) DeFi lending rates, 2) the potential passage of the CLARITY Act, and 3) the market forming a consensus on the "Fed hawkish peak." He emphasizes that the most immediate indicator for the sustainability of the recent CPI-triggered rally is the USDC borrowing rate on Aave, currently around 3.75%—close to U.S. Treasury yields. The absence of a credit spread indicates low leverage demand and a lack of market energy. He explains that a healthy, sustained rally requires layered buying pressure. Last year's rally progressed from an ETH short squeeze to crypto-native trader influx and finally to ETF inflows. Currently, this structure is missing. Other potential structural buyers like Digital Asset Treasury (DAT) companies are absent, and ETF flows have proven transient. Baehr notes that while small-cap crypto tokens outperformed large caps in Q2—a potential sign of capitation in major assets—capital is also flowing to more exciting opportunities like AI stocks and tech IPOs, leaving crypto sidelined. Regarding DeFi, he highlights that platforms like Aave provide a clear, real-time signal of leverage demand through their supply/demand-driven interest rates. A significant, sustained rate increase would signal genuine market conviction. He also observes the quiet emergence of fixed-income-like products and vaults in DeFi. On regulation, the probability of the CLARITY Act passing before the August 7th deadline has dropped linearly from 75% to below 40% on Polymarket. Baehr suggests its passage would be treated as a bullish surprise, a potent driver for price movement. However, political hurdles, including ethical clause debates and disclosures about the First Family's crypto profits, remain significant obstacles. Ultimately, the market awaits clarity on the Fed's terminal rate under Chair Warsh. Until the "Fed Solstice"—the point where the market collectively understands the peak of hawkish policy—sustained conviction will be difficult to achieve.

marsbit3m ago

Podcast Notes | Conversation with GSR Asset Management Head: To Determine if This Crypto Rally is Real, Just Watch the Lending Rates on Aave

marsbit3m ago

7 Months After the Collapse of Huiwang, Southeast Asia's Escrow Platforms Undergo a Major Reshuffle

Following the collapse of Huione Pay—dubbed the "Alipay of Southeast Asia"—seven months ago, the region's underground financial guarantee platform sector is undergoing a significant reshuffle. This power vacuum has been swiftly filled by emerging platforms such as XinBi, Tiger/Navigator, JinBei (renamed JinBo), Dali/Tiancheng, and FullyLight. These platforms, operating largely via Telegram and offering services like escrow for illicit transactions, have absorbed the vast user base and markets left behind by Huione. While positioning themselves as "trust intermediaries," their primary clientele consists of networks involved in online scams, money laundering, illegal gambling, and even human trafficking. For instance, the Tiger/Navigator platform explicitly provides "escrow" services for kidnapping-for-ransom operations ("强押车交易"). Data underscores the immense scale: Huione alone processed over $103 billion in cryptocurrency payments and facilitated over $31 billion through its escrow market before its downfall, linking it to Cambodia's notorious Prince Group. Since its collapse, competitors have seen explosive growth. For example, the XinBi platform has accumulated over $1.6 billion in total USDT revenue, while platforms like NewPay, OkPay (under Dali), and FullyLight Wallet collectively processed over $4.8 billion in USDT in a single year. This ecosystem thrives in regions like Cambodia and Myanmar, where regulatory gaps allow these platforms to act as critical financial infrastructure for sprawling cybercrime industries, from scam compounds to online casinos. The article concludes that the moniker "Southeast Asian Alipay" is a misnomer, obscuring the platforms' fundamental role in enabling serious criminal enterprises rather than representing legitimate financial innovation.

Odaily星球日报1h ago

7 Months After the Collapse of Huiwang, Southeast Asia's Escrow Platforms Undergo a Major Reshuffle

Odaily星球日报1h ago

The Changing Landscape: What Are Crypto VCs Experiencing?

Title: The Shifting Landscape of Crypto Venture Capital The era of dedicated crypto venture capital funds is undergoing a significant transformation. Once essential for navigating the sector's complexity and high risk, these specialized funds are now facing an identity crisis as the market matures. This shift mirrors historical patterns in other specialized investment classes like cleantech and SPACs, where initial information advantages dissipate as technologies become mainstream and integrated into existing industry frameworks. The article argues that crypto is reaching a critical inflection point, transitioning from a "building phase" to an "integration phase." Major players like Stripe, BlackRock, and Visa now engage with crypto not for its novel mechanics but as a foundational financial infrastructure. Their needs—regulatory compliance, banking partnerships, distribution channels—align with traditional fintech, a domain easily understood by large, generalist funds like Sequoia and Founders Fund. This evolution creates a "barbell effect" within the VC landscape. On one end are massive, diversified platforms that can incorporate crypto as one vertical among many. On the other are small, nimble funds focused on niche, experimental projects. The middle ground—medium-sized dedicated crypto funds—is being squeezed out. Their typical fund size makes it impossible to generate sufficient returns solely from early-stage crypto bets, yet they cannot compete with giants for later-stage deals. Consequently, leading crypto-native firms like Paradigm and Framework Ventures are expanding into AI, robotics, and other sectors, driven partly by LP pressure for better returns amid a broader VC DPI crisis. Others, like Dragonfly and a16z, have narrowed their crypto focus predominantly to financial infrastructure like stablecoins, reframing the sector's core narrative. For crypto entrepreneurs, this consolidation presents challenges. While generalist funds offer larger checks and broader resources, crypto projects now compete fiercely with AI for attention and capital within these firms. Furthermore, the long-term, non-commercial foundational work that built the ecosystem—funded by dedicated crypto VCs—is less likely to attract generalist capital focused on direct returns. The conclusion is that "crypto investor" as a standalone category is becoming obsolete, akin to "internet investor." Crypto is becoming a baseline infrastructure layer. The future will see a barbell structure: large-scale growth financing handled by generalist funds, while pioneering, speculative projects are funded by small, specialized vehicles. The dedicated crypto funds of the 2017-2021 boom, which incubated core infrastructure, are giving way to this new, bifurcated reality.

Foresight News1h ago

The Changing Landscape: What Are Crypto VCs Experiencing?

Foresight News1h ago

Trading

Spot
活动图片