
The entry-level Luna model has seen a massive price cut, with input price dropping from $1 to $0.2 per million tokens, and output price falling from $6 to $1.2 per million tokens.
Converted to RMB, this is equivalent to input price dropping from about 6.8 yuan to 1.35 yuan, and output price from about 40.6 yuan to 8.1 yuan.
The Terra model, positioned as the "daily workhorse," also saw a 20% price reduction, with input and output prices dropping to $2 and $12 per million tokens respectively.
Converted to RMB, this is approximately 13.5 yuan and 81.2 yuan respectively.
Only the flagship model, Sol, remains at its original price, but it introduces a Fast mode with speeds up to 2.5 times that of the standard mode, offering acceleration at no extra cost.

Following the aggressive reset of usage in Codex, OpenAI has promptly arranged for price cuts... who are you trying to out-compete?
The price war has directly escalated to a white-hot level.
OpenAI stated that the price reduction is because GPT-5.6 Sol helped save its own costs.
GPT-5.6 Sol participated in its own improvement, achieving a leap in efficiency, hence this special feedback for our new and existing users~
So OpenAI is also playing the game of "stepping on one's own feet to ascend on the spot."

.
Two Models Get Cheaper, One Gets Faster
Currently, the API prices for the three GPT-5.6 models are:
GPT-5.6 Luna: $0.2 per million tokens input, $1.2 per million tokens output;
GPT-5.6 Terra: $2 per million tokens input, $12 per million tokens output;
GPT-5.6 Sol: $5 per million tokens input, $30 per million tokens output.

After the price cut, Luna has taken a major plunge.
Its input price is now on par with the previous generation GPT-5.4 nano, and its output price is even $0.05 cheaper.
However, the two models don't handle exactly the same tasks. GPT-5.4 nano is mainly for simple, high-frequency tasks like classification, information extraction, and ranking.
GPT-5.6 Luna is positioned by OpenAI as a model for cost-sensitive workloads, capable of calling tools, handling long contexts, and executing multi-step workflows.
Simply put, OpenAI is selling the model that powers Agents at the price point of past simple small models.
Terra's reduction is relatively more restrained, at 20%.
Sol maintains its original price but adds Fast mode, which can reach speeds up to 2.5 times the standard mode, priced at 2 times the standard mode, with no change in model intelligence.
It will also replace the previous Priority Processing. API requests originally using the priority tag can automatically switch to Fast mode.
Officials stated that this round of adjustments will also be applied to Codex and ChatGPT Work.
Subscription prices will not be lowered, and users' total quota will not increase, but when calling Terra and Luna, the quota consumed will be correspondingly reduced.
The same subscription fee can allow the model to perform more tasks.
OpenAI also casually switched the Auto-review model in ChatGPT and Codex CLI from GPT-5.4 to GPT-5.6 Luna.
Auto-review is responsible for checking code and modification results in the background, a typical high-frequency Agent task.
Combined with the model upgrade and Luna's price cut, OpenAI expects its cost to drop to about one-tenth of the original.
Cheap models are no longer only responsible for traditional "miscellaneous tasks" like classification and extraction, but are starting to enter areas like code review and backend monitoring, which require judgment and tool calling.
Currently, the positioning of the three models is:
Budget-sensitive, high-volume tasks go to Luna;
Regular daily work goes to Terra;
High-difficulty tasks continue with Sol;
If even waiting is too slow, pay double for speed.
GPT-5.6 Begins Participating in Its Own Cost Reduction
Why the sudden price drop?
OpenAI stated the reason is that GPT-5.6 Sol participated in optimizing its own production system.
The efficiency it improved then translated into discounted numbers on the price list.
Specifically, GPT-5.6 Sol rewrote and optimized some production environment GPU kernels, designed and ran hundreds of token generation experiments, and participated in monitoring training, intervening when issues arose.
The final results were also impressive: production GPU kernel optimization reduced GPT-5.6 end-to-end service costs by 20%; improved speculative decoding increased token generation efficiency by over 15%.

A GPU Kernel can be understood as the underlying program that executes specific computing tasks for the model on the GPU.
The model itself doesn't need to change; as long as these programs are written more efficiently, the same GPU can complete calculations faster, handle more requests, and the cost per call also drops.
Speculative decoding involves "guessing" the next possible tokens in bulk when generating answers, then having the main model verify them. The more accurate the guesses, the fewer steps are needed for individual generation and waiting.
Both optimizations do not reduce model capability but allow the same model to run faster and cheaper.
However, this is not GPT-5.6 fully autonomously upgrading itself.
OpenAI specifically emphasized that the entire process remains human-led.
The model can write code, run experiments, and monitor anomalies, but defining goals, deciding if results are usable, and whether code enters the production environment still require "Human in the Loop."
OMT
Finally, there's one more new development.
This price cut has a very specific focus: Agent workflows.
The most heavily discounted Luna is not just a traditional small model for classification and extraction.
According to OpenAI's positioning, it can call tools and execute multi-step tasks, mainly targeting high-frequency, cost-sensitive workloads.
Therefore, Luna's 80% price cut also serves to lower the barrier for long-running Agents.
It makes review, verification, and monitoring tasks, which were previously too costly to perform frequently, feasible as regular steps in workflows.
Combined with GPT-5.6 participating in optimizing its own production system, a new cycle emerges:
The model participates in improving operational efficiency, costs drop; after prices decrease, the model enters more high-frequency workflows; increased call volume then drives the next round of efficiency optimization.
OpenAI's price reduction flywheel is taking shape.
After this series of maneuvers, the pressure is now directly on A-She (Anthropic?):
"Your turn."

Reference links:[1]https://x.com/OpenAI/status/2082878156483219672[2]https://x.com/sama/status/2082880720989532597
This article is from the WeChat public account "QbitAI," author: Follow Frontier Technology







