OpenAI No Longer Sells Its Most Expensive Model for Profit

marsbit发布于2026-08-03更新于2026-08-03

文章摘要

OpenAI is shifting its business strategy away from promoting its most expensive, flagship models for every task. Recent price cuts—80% for GPT-5.6 Luna and 20% for Terra—signal a deeper change: the company now actively advises users that many tasks don't require the most powerful model. Instead, OpenAI recommends a tiered approach: use the high-end GPT-5.6 Sol for complex planning and analysis, then delegate execution to cheaper models like Luna. This mirrors moves by Anthropic, which recently launched Claude Opus 5 at half the price of its top model, Fable 5. Both companies are de-emphasizing flagship models as primary revenue drivers, using them instead for brand prestige and technological showcases. The industry is entering a "mass-market" phase, similar to automotive, where high-volume, cost-effective models handle daily operations and drive scale. OpenAI's price reductions are partly enabled by AI models themselves optimizing underlying code and infrastructure, creating a self-reinforcing cycle of efficiency gains and cost reduction. Competition is shifting from "who is smartest" to "who offers the best value." The goal is no longer selling individual models but fostering widespread API adoption and ecosystem lock-in. By making AI calls cheap and ubiquitous, companies like OpenAI aim to become the indispensable, utility-like infrastructure powering automated workflows—the "water and electricity" of software, quietly embedded everywhere.

If someone is still spending the most money to call OpenAI's most powerful model today.

OpenAI would instead advise them to switch to another one.

On July 30, OpenAI issued a price adjustment announcement.

The GPT-5.6 Luna model was reduced by 80%, and the Terra model by 20%.

Seeing this news, it's easy to focus on the price war starting in Silicon Valley.

However, if you carefully review the officially published technical documentation and API usage guide, you realize that the truly noteworthy action is not in the price numbers themselves.

This essentially marks the first time OpenAI has begun to tell users that for many tasks, the strongest model isn't actually necessary.

The official gave a very specific suggestion. For a complex task, first use GPT-5.6 Sol for requirement analysis and solution design, then hand it over to Luna for execution, coding, and running tests.

The most expensive model is responsible for thinking, the cheapest model is responsible for doing the work. This strategy, viewed two years ago, would have been equivalent to commercial self-denial.

After all, in the past, the entire AI industry was desperately trying to tell the market that their model was the smartest.

Yet today, OpenAI stands up and says you don't always need to buy the most expensive one.

This matter is far more important than the price cut.

1. The Tacit Understanding of Silicon Valley's Two Giants

First, look at what happened in the past two weeks.

July 30: OpenAI adjusts prices. Luna down 80%. Terra down 20%. Sol did not see a price cut; instead, a Fast mode was added, offering speeds up to 2.5x faster than the standard mode, with double the price, but with identical intelligence levels.

The top-tier model remains. What is truly starting to gain volume are the mid-to-low end models.

One week earlier.

July 24: Anthropic did almost exactly the same thing. Claude Opus 5 was released, priced at $5 for 1 million input tokens and $25 for output tokens. Exactly half the price of Fable 5.

Compared to performance breakthroughs, Anthropic emphasized its cost-effectiveness externally: with only half the price, you can obtain cutting-edge reasoning capabilities infinitely close to Fable 5.

Just one month ago, Fable 5 was Anthropic's flagship product, heavily promoted as the strongest reasoning, longest context, highest price. One month later, Anthropic personally found a half-price alternative for its flagship.

If only one company did this, it could be understood as a product adjustment. When two companies do it almost simultaneously, it's not a coincidence.

They have both begun to actively reduce the importance of their flagship models.

2. Flagships Handle the Stage, Volume Models Handle the Profit

I've been thinking, why now of all times?

The answer isn't actually complicated.

In the past, the biggest value of flagship models wasn't making money; it was proving technological leadership. After GPT-4 came out, OpenAI's valuation rose continuously. Every time Claude updated, Anthropic would redefine its technological position. Flagship models carried brand value.

But where companies actually spend their money isn't there.

A company runs millions of API calls daily. Customer service, search, approvals, code generation, Agent execution—these high-frequency tasks consume the vast majority of Tokens. What enterprise procurement cares about most isn't being first on the Benchmark, but how much a single task costs, whether it's stable enough, and the ROI.

When call volumes expand to tens of millions per day, the slight intelligence advantage of flagship models is instantly erased by the enormous compute costs.

OpenAI's action and stance mark a turning point: top-tier flagships are no longer tasked with making money.

3. AI Begins Entering the "Mass-Market Vehicle" Era

This scene has already played out in the automotive industry.

Twenty years ago, the 7 Series defined BMW's height, the S-Class upheld Mercedes-Benz's luxury appeal, the A8 established Audi's flagship image—flagship cars determined brand ceilings. Yet what truly supported brand sales and generated profits was always the BMW 3 Series, Mercedes C-Class, and Audi A4.

Later, it became even more evident. The Model S proved Tesla could build cars. What truly made it a global automaker were the Model 3 and Model Y.

Flagships prove capability; mass-market models handle scale.

AI is now beginning to enter this stage. Sol and Fable will continue to exist; they are responsible for pushing the technological boundaries and refreshing Benchmarks. The ones truly shouldering commercialization will increasingly become Luna, Terra, and Opus.

OpenAI even publicly wrote out the recommended workflow this time: Sol for planning, Luna for execution.

This is no longer just one model; OpenAI is designing a system of model division of labor. What enterprises buy in the future is not one model, but an entire suite of models. What truly determines costs isn't the chief architect, but the construction crew working every day.

4. Models Begin Optimizing Models

There's another detail I find more interesting than the price cut itself.

OpenAI mentioned in the technical notes that this price reduction is not solely due to procuring more GPUs or scaling up. The real reason is that models are starting to participate in optimizing models.

Specifically, under the guidance of human engineers, Sol autonomously rewrote and optimized the underlying production kernel. It designed hundreds of experiments itself to improve Token generation efficiency and even participated in monitoring the model training pipeline, directly intervening when problems were discovered.

The result is a 20% reduction in end-to-end operating costs and a 15% improvement in Token generation efficiency.

This information is easily overlooked, but its significance is substantial.

In the past, improving efficiency relied on engineers. After a model launched, humans would optimize the inference framework, CUDA, caching strategies, and scheduling algorithms bit by bit. Over a year, squeezing out a dozen percentage points of efficiency improvement was considered good.

Today, technological evolution has taken a new path. Models are beginning to take over the engineering optimization of underlying code and compute scheduling, iterating and running 24/7.

This is a self-accelerating cycle. The smarter the model, the stronger its ability to participate in optimization. The faster the optimization, the quicker the cost drops. The lower the cost, the larger the call volume. The larger the call volume, the more data generated, which continues to train the model.

Looking back at OpenAI's price changes over the past two and a half years. GPT-4 debuted at $30 per million input Tokens, GPT-4o dropped to $5, GPT-4o mini reached $0.15. Today, Luna is priced close to the cheap range of the earlier mini, yet its overall intelligence level has long surpassed the expensive GPT-4 from two years ago.

In just over two years, prices have dropped by nearly two orders of magnitude. If models continue to participate in optimizing themselves, this curve will most likely continue its downward trend.

The truly formidable aspect is not that the price dropped 80% today, but that cost reduction has begun to possess self-driving capability.

5. From Who Is Smartest to Who Is Most Worth It

I increasingly feel that when discussing AI competition today, people sometimes still apply the framework from the previous stage.

For example, they often still ask: Who is the smartest? GPT, Claude, Gemini, DeepSeek. Every time a new model is released, the media first looks at the leaderboard. Whoever is first, wins.

However, the new moves by OpenAI and Anthropic break this pattern; they send a new signal to the market: The appeal of single-performance champions is fading.

OpenAI mentioned a particularly crucial sentence in its announcement, suggesting developers match different models based on the task's importance, error cost, urgency, and scale.

Note, they are no longer discussing the model, but the task.

In the past, when a company deployed AI, it mostly had only one choice. Starting today, it's more like building an organization. The most critical tasks use Sol, daily execution is handed to Luna, and in the future, even lighter models might handle simpler tasks.

This is very similar to the early days of cloud computing. No one puts all data on the most expensive SSDs. Hot data goes on SSDs, ordinary data on HDDs, cold data in object storage. People never discuss which hard drive is the fastest, but how to build the entire system most cost-effectively.

In fact, DeepSeek sensed this direction earlier than Silicon Valley. Over the past six months, it has hardly emphasized being the world's smartest, repeating only a few words: cheap, fast enough, good enough.

After cache hits, the cost per million Tokens becomes almost negligible. It has been betting on one thing: what enterprises need is not a world champion, but the deployment option with the highest comprehensive ROI.

This is not a short-term price war. The current follow-up by the two Silicon Valley giants validates the inevitability of this commercial path.

6. What OpenAI Really Wants to Sell Is Not Models

By now, OpenAI's strategy is clear. What it really wants to sell is no longer models, but call volume.

In the past, model vendors relied on high technological premiums for high margins. Now, the business logic is shifting to exchanging extremely low barriers for ultra-large-scale traffic ecosystems.

Microsoft didn't make real money because Windows was expensive, but because all computers ran Windows. AWS didn't make money because individual server profits were high, but because countless applications worldwide run on it every day.

Platform revenue has always relied on penetration rate.

This explains why Sol's price hasn't moved. Sol bears the brand, proving OpenAI is still the company with the highest technological ceiling. Luna is the revenue engine.

OpenAI hopes developers form a new default habit: use Luna for writing Agents, use Luna for running workflows, use Luna for batch execution. Only when encountering truly difficult problems do they call Sol once.

Once this default is established, future competition becomes very difficult. Migrating a company's underlying model once means retesting, validating, and adapting the entire workflow. Migration costs will become increasingly higher.

The true moat is beginning to shift from capability leadership to ecosystem stickiness.

7. Flagship Models No Longer Determine the Direction

Taking a longer view, the AI industry is experiencing a classic economies-of-scale inflection point.

When the unit cost of compute drops to extremely low levels, the market's total demand for Tokens does not decrease as the unit price falls; instead, it explodes exponentially.

Flagship models no longer determine the direction because the focus of technological evolution has shifted from exploring the upper limits of intelligence to the industrial cost reduction of compute. When API call costs become low enough to be negligible, the form and boundaries of models will begin to fade.

Enterprise developers will no longer focus on the consumption of every single Token, but seamlessly embed AI into every business process.

The most profound aspect of this transformation is that the collapse of API prices is raising the migration costs of entire software engineering ecosystems. Once a company's workflows, Agent scheduling networks, and automated data pipelines are all built on combinations of low-cost models, the provider of the underlying models locks in the compute pipeline for the next decade.

Over the past few years, large model companies sold intelligence. Starting today, they are beginning to sell efficiency.

These are two completely different stories.

A note "Beyond the Layout":

In the past, we were accustomed to analogizing AI development to consumer electronics, expecting new record-breaking flagships from time to time.

But perhaps the true winning form of AI is not becoming a sensational product.

When the steam engine was first invented, people marveled at its productivity. Today, electricity flows everywhere, powering the operation of entire civilizations, yet no one specifically discusses it anymore.

When humans no longer passionately discuss which flagship model has refreshed which IQ benchmark, and AI silently embeds itself into every system, every command, becoming the water and electricity default-called behind all automated processes—

Only then will its true era have just begun.

This article is from the WeChat public account "Beyond the Layout", author: Huahua

热门币种推荐

相关问答

QWhat is the most significant change in OpenAI's recent price adjustment announcement according to the article?

AThe most significant change is not the price cuts themselves, but that OpenAI is, for the first time, explicitly advising users that for many tasks, the most powerful (and expensive) model is not necessary. They are promoting a strategy of using a cheaper model for execution after planning with a top-tier model.

QWhat analogy does the article use to describe the AI industry's shift in focus from flagship models?

AThe article uses the automobile industry as an analogy. It compares flagship AI models (like OpenAI's Sol or Anthropic's Fable) to luxury car flagships (e.g., BMW 7 Series) that define the brand's technological height. The real volume and profit, however, come from the 'mass-market' models (like OpenAI's Luna or a BMW 3 Series), which are responsible for large-scale commercialization.

QWhat is a key technical reason mentioned for OpenAI's ability to lower costs, beyond just buying more GPUs?

AA key reason is that the AI models themselves are now participating in optimizing the models. Specifically, the top-tier model Sol, under human guidance, autonomously rewrote and optimized the underlying production kernel, designed experiments, and improved token generation efficiency, leading to significant cost reductions.

QHow does the article suggest the competitive framework for AI companies is changing?

AThe article suggests the framework is shifting from 'who is the smartest' (focused on benchmark performance and flagship models) to 'who offers the best value' (focused on overall ROI, cost-effectiveness, and building an ecosystem where cheaper, 'good enough' models handle most tasks). Companies like DeepSeek are cited as emphasizing 'cheap, fast enough, and usable.'

QWhat does the article conclude is the ultimate 'real product' OpenAI wants to sell, and what historical parallel is drawn?

AThe article concludes that OpenAI's real goal is not to sell individual models, but to sell *usage/volume*—massive API call throughput. It draws a parallel to utilities like electricity. Just as we don't discuss electricity itself but rely on it seamlessly, AI's ultimate victory is to become an invisible, default infrastructure powering all automated processes, not a frequently debated flagship product.

你可能也喜欢

高盛旗帜鲜明:这是人类历史上最大规模的资本需求周期,美联储只是看客

高盛欧洲对冲基金业务负责人Mark Wilson指出,全球正经历人类历史上最大规模的资本需求周期。过去数十年的“储蓄过剩”时代正在终结,AI基础设施建设、再工业化、国防重整、电力系统重建、供应链重组以及主权债务融资等多重需求同时争夺资本,将推动资本成本持续上升,并改变投资范式。 Wilson强调,美联储在此过程中更多是“乘客”而非“司机”,收益率上行的根本驱动力是结构性资本需求,而非货币政策。近期市场指数表面平静,但内部极度分化,个股波动剧烈且动量因子出现历史性回撤,导致机构大规模去风险化。 尽管企业盈利基本面依然强劲,但Wilson认为8月市场将进入消化期,不宜轻易抄底。利率突破的影响尚未完全消化,风险模型参数因极端波动而改变,加之美国中期选举临近,市场前景复杂。 超大规模云计算企业(如亚马逊、微软)的资本开支预测惊人,但AI相关业务收入增长和回报率同样强劲,验证了投资效率。然而,Wilson指出,私营部门全力投资AI未来与国家部门面临资本约束之间的矛盾日益凸显,全球经济最终将面对资本重新定价带来的政治抉择。目前,AI超级周期仍在演进,8月可能是一个相对平静的观察窗口。

marsbit21分钟前

高盛旗帜鲜明:这是人类历史上最大规模的资本需求周期,美联储只是看客

marsbit21分钟前

交易

现货

热门文章

如何购买S

欢迎来到HTX.com!我们已经让购买Sonic(S)变得简单而便捷。跟随我们的逐步指南,放心开始您的加密货币之旅。第一步:创建您的HTX账户使用您的电子邮件、手机号码注册一个免费账户在HTX上。体验无忧的注册过程并解锁所有平台功能。立即注册第二步:前往买币页面,选择您的支付方式信用卡/借记卡购买:使用您的Visa或Mastercard即时购买Sonic(S)。余额购买:使用您HTX账户余额中的资金进行无缝交易。第三方购买:探索诸如Google Pay或Apple Pay等流行支付方法以增加便利性。C2C购买:在HTX平台上直接与其他用户交易。HTX场外交易台(OTC)购买:为大量交易者提供个性化服务和竞争性汇率。第三步:存储您的Sonic(S)购买完您的Sonic(S)后,将其存储在您的HTX账户钱包中。您也可以通过区块链转账将其发送到其他地方或者用于交易其他加密货币。第四步:交易Sonic(S)在HTX的现货市场轻松交易Sonic(S)。访问您的账户,选择您的交易对,执行您的交易,并实时监控。HTX为初学者和经验丰富的交易者提供了友好的用户体验。

3.3k人学过发布于 2025.01.15更新于 2026.06.02

如何购买S

相关讨论

欢迎来到HTX社区。在这里,您可以了解最新的平台发展动态并获得专业的市场意见。以下是用户对S(S)币价的意见。

活动图片