
If someone is still spending the most money to call OpenAI's most powerful model today.
OpenAI would instead advise them to switch to another one.
On July 30, OpenAI issued a price adjustment announcement.
The GPT-5.6 Luna model was reduced by 80%, and the Terra model by 20%.
Seeing this news, it's easy to focus on the price war starting in Silicon Valley.
However, if you carefully review the officially published technical documentation and API usage guide, you realize that the truly noteworthy action is not in the price numbers themselves.
This essentially marks the first time OpenAI has begun to tell users that for many tasks, the strongest model isn't actually necessary.
The official gave a very specific suggestion. For a complex task, first use GPT-5.6 Sol for requirement analysis and solution design, then hand it over to Luna for execution, coding, and running tests.
The most expensive model is responsible for thinking, the cheapest model is responsible for doing the work. This strategy, viewed two years ago, would have been equivalent to commercial self-denial.
After all, in the past, the entire AI industry was desperately trying to tell the market that their model was the smartest.
Yet today, OpenAI stands up and says you don't always need to buy the most expensive one.
This matter is far more important than the price cut.
1. The Tacit Understanding of Silicon Valley's Two Giants
First, look at what happened in the past two weeks.
July 30: OpenAI adjusts prices. Luna down 80%. Terra down 20%. Sol did not see a price cut; instead, a Fast mode was added, offering speeds up to 2.5x faster than the standard mode, with double the price, but with identical intelligence levels.
The top-tier model remains. What is truly starting to gain volume are the mid-to-low end models.
One week earlier.
July 24: Anthropic did almost exactly the same thing. Claude Opus 5 was released, priced at $5 for 1 million input tokens and $25 for output tokens. Exactly half the price of Fable 5.
Compared to performance breakthroughs, Anthropic emphasized its cost-effectiveness externally: with only half the price, you can obtain cutting-edge reasoning capabilities infinitely close to Fable 5.
Just one month ago, Fable 5 was Anthropic's flagship product, heavily promoted as the strongest reasoning, longest context, highest price. One month later, Anthropic personally found a half-price alternative for its flagship.
If only one company did this, it could be understood as a product adjustment. When two companies do it almost simultaneously, it's not a coincidence.
They have both begun to actively reduce the importance of their flagship models.
2. Flagships Handle the Stage, Volume Models Handle the Profit
I've been thinking, why now of all times?
The answer isn't actually complicated.
In the past, the biggest value of flagship models wasn't making money; it was proving technological leadership. After GPT-4 came out, OpenAI's valuation rose continuously. Every time Claude updated, Anthropic would redefine its technological position. Flagship models carried brand value.
But where companies actually spend their money isn't there.
A company runs millions of API calls daily. Customer service, search, approvals, code generation, Agent execution—these high-frequency tasks consume the vast majority of Tokens. What enterprise procurement cares about most isn't being first on the Benchmark, but how much a single task costs, whether it's stable enough, and the ROI.
When call volumes expand to tens of millions per day, the slight intelligence advantage of flagship models is instantly erased by the enormous compute costs.
OpenAI's action and stance mark a turning point: top-tier flagships are no longer tasked with making money.
3. AI Begins Entering the "Mass-Market Vehicle" Era
This scene has already played out in the automotive industry.
Twenty years ago, the 7 Series defined BMW's height, the S-Class upheld Mercedes-Benz's luxury appeal, the A8 established Audi's flagship image—flagship cars determined brand ceilings. Yet what truly supported brand sales and generated profits was always the BMW 3 Series, Mercedes C-Class, and Audi A4.
Later, it became even more evident. The Model S proved Tesla could build cars. What truly made it a global automaker were the Model 3 and Model Y.
Flagships prove capability; mass-market models handle scale.
AI is now beginning to enter this stage. Sol and Fable will continue to exist; they are responsible for pushing the technological boundaries and refreshing Benchmarks. The ones truly shouldering commercialization will increasingly become Luna, Terra, and Opus.
OpenAI even publicly wrote out the recommended workflow this time: Sol for planning, Luna for execution.
This is no longer just one model; OpenAI is designing a system of model division of labor. What enterprises buy in the future is not one model, but an entire suite of models. What truly determines costs isn't the chief architect, but the construction crew working every day.
4. Models Begin Optimizing Models
There's another detail I find more interesting than the price cut itself.
OpenAI mentioned in the technical notes that this price reduction is not solely due to procuring more GPUs or scaling up. The real reason is that models are starting to participate in optimizing models.
Specifically, under the guidance of human engineers, Sol autonomously rewrote and optimized the underlying production kernel. It designed hundreds of experiments itself to improve Token generation efficiency and even participated in monitoring the model training pipeline, directly intervening when problems were discovered.
The result is a 20% reduction in end-to-end operating costs and a 15% improvement in Token generation efficiency.
This information is easily overlooked, but its significance is substantial.
In the past, improving efficiency relied on engineers. After a model launched, humans would optimize the inference framework, CUDA, caching strategies, and scheduling algorithms bit by bit. Over a year, squeezing out a dozen percentage points of efficiency improvement was considered good.
Today, technological evolution has taken a new path. Models are beginning to take over the engineering optimization of underlying code and compute scheduling, iterating and running 24/7.
This is a self-accelerating cycle. The smarter the model, the stronger its ability to participate in optimization. The faster the optimization, the quicker the cost drops. The lower the cost, the larger the call volume. The larger the call volume, the more data generated, which continues to train the model.
Looking back at OpenAI's price changes over the past two and a half years. GPT-4 debuted at $30 per million input Tokens, GPT-4o dropped to $5, GPT-4o mini reached $0.15. Today, Luna is priced close to the cheap range of the earlier mini, yet its overall intelligence level has long surpassed the expensive GPT-4 from two years ago.
In just over two years, prices have dropped by nearly two orders of magnitude. If models continue to participate in optimizing themselves, this curve will most likely continue its downward trend.
The truly formidable aspect is not that the price dropped 80% today, but that cost reduction has begun to possess self-driving capability.
5. From Who Is Smartest to Who Is Most Worth It
I increasingly feel that when discussing AI competition today, people sometimes still apply the framework from the previous stage.
For example, they often still ask: Who is the smartest? GPT, Claude, Gemini, DeepSeek. Every time a new model is released, the media first looks at the leaderboard. Whoever is first, wins.
However, the new moves by OpenAI and Anthropic break this pattern; they send a new signal to the market: The appeal of single-performance champions is fading.
OpenAI mentioned a particularly crucial sentence in its announcement, suggesting developers match different models based on the task's importance, error cost, urgency, and scale.
Note, they are no longer discussing the model, but the task.
In the past, when a company deployed AI, it mostly had only one choice. Starting today, it's more like building an organization. The most critical tasks use Sol, daily execution is handed to Luna, and in the future, even lighter models might handle simpler tasks.
This is very similar to the early days of cloud computing. No one puts all data on the most expensive SSDs. Hot data goes on SSDs, ordinary data on HDDs, cold data in object storage. People never discuss which hard drive is the fastest, but how to build the entire system most cost-effectively.
In fact, DeepSeek sensed this direction earlier than Silicon Valley. Over the past six months, it has hardly emphasized being the world's smartest, repeating only a few words: cheap, fast enough, good enough.
After cache hits, the cost per million Tokens becomes almost negligible. It has been betting on one thing: what enterprises need is not a world champion, but the deployment option with the highest comprehensive ROI.
This is not a short-term price war. The current follow-up by the two Silicon Valley giants validates the inevitability of this commercial path.
6. What OpenAI Really Wants to Sell Is Not Models
By now, OpenAI's strategy is clear. What it really wants to sell is no longer models, but call volume.
In the past, model vendors relied on high technological premiums for high margins. Now, the business logic is shifting to exchanging extremely low barriers for ultra-large-scale traffic ecosystems.
Microsoft didn't make real money because Windows was expensive, but because all computers ran Windows. AWS didn't make money because individual server profits were high, but because countless applications worldwide run on it every day.
Platform revenue has always relied on penetration rate.
This explains why Sol's price hasn't moved. Sol bears the brand, proving OpenAI is still the company with the highest technological ceiling. Luna is the revenue engine.
OpenAI hopes developers form a new default habit: use Luna for writing Agents, use Luna for running workflows, use Luna for batch execution. Only when encountering truly difficult problems do they call Sol once.
Once this default is established, future competition becomes very difficult. Migrating a company's underlying model once means retesting, validating, and adapting the entire workflow. Migration costs will become increasingly higher.
The true moat is beginning to shift from capability leadership to ecosystem stickiness.
7. Flagship Models No Longer Determine the Direction
Taking a longer view, the AI industry is experiencing a classic economies-of-scale inflection point.
When the unit cost of compute drops to extremely low levels, the market's total demand for Tokens does not decrease as the unit price falls; instead, it explodes exponentially.
Flagship models no longer determine the direction because the focus of technological evolution has shifted from exploring the upper limits of intelligence to the industrial cost reduction of compute. When API call costs become low enough to be negligible, the form and boundaries of models will begin to fade.
Enterprise developers will no longer focus on the consumption of every single Token, but seamlessly embed AI into every business process.
The most profound aspect of this transformation is that the collapse of API prices is raising the migration costs of entire software engineering ecosystems. Once a company's workflows, Agent scheduling networks, and automated data pipelines are all built on combinations of low-cost models, the provider of the underlying models locks in the compute pipeline for the next decade.
Over the past few years, large model companies sold intelligence. Starting today, they are beginning to sell efficiency.
These are two completely different stories.
A note "Beyond the Layout":
In the past, we were accustomed to analogizing AI development to consumer electronics, expecting new record-breaking flagships from time to time.
But perhaps the true winning form of AI is not becoming a sensational product.
When the steam engine was first invented, people marveled at its productivity. Today, electricity flows everywhere, powering the operation of entire civilizations, yet no one specifically discusses it anymore.
When humans no longer passionately discuss which flagship model has refreshed which IQ benchmark, and AI silently embeds itself into every system, every command, becoming the water and electricity default-called behind all automated processes—
Only then will its true era have just begun.
This article is from the WeChat public account "Beyond the Layout", author: Huahua






