For the same model, the official price is reduced by only 20%, but your bill shrinks by eighty percent.

On August 24th, OpenAI announced test results conducted jointly with AWS:
On Terminal-Bench 2.1, the cost for GPT-5.6 Terra to successfully complete a task in Kiro was reduced by approximately 82%.
Kiro is AWS's intelligent software developer agent platform, covering IDE, CLI, and Web.
The three siblings of the GPT-5.6 family—Sol, Terra, and Luna—have been running inside for over a month.
This 82% reduction is not the official price cut.
Terra's last price adjustment was on July 30th, by 20%.

OpenAI price adjustment announcement on July 30th, Terra lowered by 20%.
The unit price dropped only 20%, but the bill could be slashed by eighty percent.
The source of the remaining sixty percentage points saved in the middle is what truly deserves our attention.
GPT-5.6 landed on Kiro in July this year.
First, on the 13th, AWS announced that GPT-5.6 Sol, Terra, and Luna were officially available on Amazon Bedrock.
The very next day, Kiro published a blog post announcing the availability of the three models on IDE, CLI, and Web.

This marked the first time OpenAI models entered Kiro, coinciding with Kiro's one-year public preview anniversary.
First, put the models on the shelf, then put them to work. Over a month later, OpenAI came back with its homework:
The two companies jointly tuned the Kiro environment and OpenAI models, reducing the cost for Terra to complete a successful task by about 82%.
What's Saved
Is the Money Spent on Detours
During the price adjustment on July 30th, Terra was reduced by 20%, but in Kiro's tests, the cost per task dropped by 82%.
Where did the extra sixty percent savings come from?
The directions are limited:
The model generated fewer tokens, the number of back-and-forth tool calls decreased, and there were fewer retries and detours after failures.
Therefore, the large chunk saved is not the cost per call, but the cost of those calls that would have been wasted.
The logic is simple: if an AI agent fails a task once, the bill is still charged. If it chooses the wrong path, goes off-track three times, and then circles back, those tokens are also billed.
In real development, money often leaks out this way.
OpenAI has pointed out the same logic in its official blog: efficiency comes from three layers:
The agent framework that initiates requests and organizes context, the orchestration system that schedules requests in the middle, and finally the model itself running on GPUs.

OpenAI breaks down the sources of GPT-5.6's efficiency: requests start from the agent framework, are scheduled by the orchestration system, and finally run the model on GPU, saving at every layer.
Savings can also come from model specialization.
OpenAI also gave an example of usage: a coding workflow can first use Sol to think through the problem and define the plan, then switch to Luna to implement the well-defined changes, write tests, and run evaluations.
Same pipeline, different levels of intelligence allocated to different stages.
The Model Accounts for Only Half the Bill
Change the Framework, Change the Price
The Terminal-Bench 2.1 benchmark doesn't ask the model to answer questions alone.
It places the model in a terminal environment with a vague objective, letting it plan its own path, call tools, write scripts, handle errors, and iterate repeatedly.
So the resulting score is the performance of the "agent + model" combination.

The public Terminal-Bench 2.1 leaderboard, with cost added to the right of accuracy. (Source: Terminal-Bench)
The four lines of numbers in the leaderboard illustrate the point best:
Claude Code with Fable 5, 83.8%, $552.67;
Codex with GPT-5.5, 83.1%, $2059.19;
Codex with GPT-5.6 Terra, 78.4%, $421.15;
Codex with GPT-5.6 Luna, 75.7%, $241.45.
The scores in the first two rows differ by only 0.7 percentage points, but the bills differ by nearly 4 times.
The same model, placed into different frameworks with different context organization and tool strategies, results in completely different prices.
According to data provided by Kiro, Terra scored 77.4 on the Coding Agent Index, only slightly higher than Claude Fable 5's 77.2.
Its selling point isn't the score, but the price corresponding to that score.
This is also where Kiro's spec-driven focus lies. The core approach is simple: don't start writing code immediately.
It first breaks down the user's vague goal into a formal requirements document, technical design, and executable task list before handing it over to the model.
Thus, the model receives not a vague statement, but a well-defined job.
Those familiar with Agents will immediately realize that this step saves the most expensive part of the expenditure.
Models going off-track, reworking, and starting over often burn more tokens than doing the actual work.
Kiro also includes two checkpoints in the process: pause for human review before code is actually modified; and automatically run a round of tests after the work is done to verify correctness.
Each rework stopped by these two checkpoints saves real money.
Claude in Amazon's Territory
GPT Takes Half
A year ago, Kiro was just a spec-driven IDE, and the model selector was Anthropic's domain.
A year later, AWS placed three tiers of OpenAI models into its own developer agent platform at once.
Sol, Terra, Luna listed alongside Claude in the same dropdown menu—a scene hard to imagine a year ago.
Although GPT-5.6 is "fully deployed" this time, it's not "fully open."
The three models are released progressively and experimentally, targeting Pro, Pro+, Pro Max, and Power users. Availability is limited to two regions: US North Virginia and Europe Frankfurt, supporting cross-region inference.
There's also a point many find hard to adjust to: these models in Kiro operate with a hidden chain-of-thought; you can't see its reasoning steps, only the final result.
Those accustomed to watching the Agent reason step-by-step feel like throwing work into an opaque box.
The official statement is that this is expected behavior and doesn't affect output quality.
The three model tiers are clearly priced in Kiro.
When they first launched on July 14th, the same task cost 2.4x for Sol, 1.2x for Terra, and 0.6x for Luna.
After OpenAI's price reduction on July 30th took effect, Kiro followed the next day: Luna was slashed from 0.6x all the way to 0.1x, Terra reduced from 1.2x to 1.0x, with only Sol unchanged.

AWS's stance is clear: a development platform cannot be tied to just one model.
In the same selector, two cutting-edge models are beginning to undercut each other on price.
The evaluation criteria for models is also changing: a higher score no longer guarantees a win; spending less can also win.
For developers, the question used to be "How much per million tokens for this model?" Now it must be "How much will it actually cost me to get this thing done?"
References:
https://x.com/OpenAIDevs/status/2091966982015103068
https://openai.com/index/gpt-5-6-in-kiro/
This article is from WeChat Official Account "AI_era" (ID: AI_era), author: ASI Revelation, editor: Yuanyu






