Price Cut Just 20%, Bill Drops 80%: GPT-5.6 Steps into Claude's Turf to Recalculate the Programming Bill

marsbitXuất bản vào 2026-08-28Cập nhật gần nhất vào 2026-08-28

Tóm tắt

While the official price for GPT-5.6 Terra only dropped by 20%, developer costs for successful coding tasks have reportedly been slashed by 82% when using the model within AWS's Kiro platform. This dramatic reduction stems not from model price cuts alone, but from significant efficiency gains within the integrated "agent + model" system. By optimizing the workflow—reducing unnecessary tokens, tool calls, and failed attempts—the collaboration between OpenAI and AWS has minimized costly computational detours. Key to this efficiency is Kiro's "spec-driven" approach, which refines vague user requests into clear technical specifications before the model begins coding, preventing expensive misunderstandings and rewrites. Benchmark results highlight that the choice of AI agent framework significantly impacts cost, with different frameworks yielding vastly different bills for similar performance scores. The integration marks OpenAI's entry into Kiro, a platform previously dominated by Anthropic's Claude. AWS now offers developers a choice between GPT-5.6 models (Sol, Terra, Luna) and Claude, fostering direct competition. This shift reframes the model selection question from "cost per million tokens" to "total cost to complete the task," emphasizing end-to-end efficiency over raw benchmark scores.

For the same model, the official price is reduced by only 20%, but your bill shrinks by eighty percent.

On August 24th, OpenAI announced test results conducted jointly with AWS:

On Terminal-Bench 2.1, the cost for GPT-5.6 Terra to successfully complete a task in Kiro was reduced by approximately 82%.

Kiro is AWS's intelligent software developer agent platform, covering IDE, CLI, and Web.

The three siblings of the GPT-5.6 family—Sol, Terra, and Luna—have been running inside for over a month.

This 82% reduction is not the official price cut.

Terra's last price adjustment was on July 30th, by 20%.

OpenAI price adjustment announcement on July 30th, Terra lowered by 20%.

The unit price dropped only 20%, but the bill could be slashed by eighty percent.

The source of the remaining sixty percentage points saved in the middle is what truly deserves our attention.

GPT-5.6 landed on Kiro in July this year.

First, on the 13th, AWS announced that GPT-5.6 Sol, Terra, and Luna were officially available on Amazon Bedrock.

The very next day, Kiro published a blog post announcing the availability of the three models on IDE, CLI, and Web.

This marked the first time OpenAI models entered Kiro, coinciding with Kiro's one-year public preview anniversary.

First, put the models on the shelf, then put them to work. Over a month later, OpenAI came back with its homework:

The two companies jointly tuned the Kiro environment and OpenAI models, reducing the cost for Terra to complete a successful task by about 82%.

What's Saved

Is the Money Spent on Detours

During the price adjustment on July 30th, Terra was reduced by 20%, but in Kiro's tests, the cost per task dropped by 82%.

Where did the extra sixty percent savings come from?

The directions are limited:

The model generated fewer tokens, the number of back-and-forth tool calls decreased, and there were fewer retries and detours after failures.

Therefore, the large chunk saved is not the cost per call, but the cost of those calls that would have been wasted.

The logic is simple: if an AI agent fails a task once, the bill is still charged. If it chooses the wrong path, goes off-track three times, and then circles back, those tokens are also billed.

In real development, money often leaks out this way.

OpenAI has pointed out the same logic in its official blog: efficiency comes from three layers:

The agent framework that initiates requests and organizes context, the orchestration system that schedules requests in the middle, and finally the model itself running on GPUs.

OpenAI breaks down the sources of GPT-5.6's efficiency: requests start from the agent framework, are scheduled by the orchestration system, and finally run the model on GPU, saving at every layer.

Savings can also come from model specialization.

OpenAI also gave an example of usage: a coding workflow can first use Sol to think through the problem and define the plan, then switch to Luna to implement the well-defined changes, write tests, and run evaluations.

Same pipeline, different levels of intelligence allocated to different stages.

The Model Accounts for Only Half the Bill

Change the Framework, Change the Price

The Terminal-Bench 2.1 benchmark doesn't ask the model to answer questions alone.

It places the model in a terminal environment with a vague objective, letting it plan its own path, call tools, write scripts, handle errors, and iterate repeatedly.

So the resulting score is the performance of the "agent + model" combination.

The public Terminal-Bench 2.1 leaderboard, with cost added to the right of accuracy. (Source: Terminal-Bench)

The four lines of numbers in the leaderboard illustrate the point best:

Claude Code with Fable 5, 83.8%, $552.67;

Codex with GPT-5.5, 83.1%, $2059.19;

Codex with GPT-5.6 Terra, 78.4%, $421.15;

Codex with GPT-5.6 Luna, 75.7%, $241.45.

The scores in the first two rows differ by only 0.7 percentage points, but the bills differ by nearly 4 times.

The same model, placed into different frameworks with different context organization and tool strategies, results in completely different prices.

According to data provided by Kiro, Terra scored 77.4 on the Coding Agent Index, only slightly higher than Claude Fable 5's 77.2.

Its selling point isn't the score, but the price corresponding to that score.

This is also where Kiro's spec-driven focus lies. The core approach is simple: don't start writing code immediately.

It first breaks down the user's vague goal into a formal requirements document, technical design, and executable task list before handing it over to the model.

Thus, the model receives not a vague statement, but a well-defined job.

Those familiar with Agents will immediately realize that this step saves the most expensive part of the expenditure.

Models going off-track, reworking, and starting over often burn more tokens than doing the actual work.

Kiro also includes two checkpoints in the process: pause for human review before code is actually modified; and automatically run a round of tests after the work is done to verify correctness.

Each rework stopped by these two checkpoints saves real money.

Claude in Amazon's Territory

GPT Takes Half

A year ago, Kiro was just a spec-driven IDE, and the model selector was Anthropic's domain.

A year later, AWS placed three tiers of OpenAI models into its own developer agent platform at once.

Sol, Terra, Luna listed alongside Claude in the same dropdown menu—a scene hard to imagine a year ago.

Although GPT-5.6 is "fully deployed" this time, it's not "fully open."

The three models are released progressively and experimentally, targeting Pro, Pro+, Pro Max, and Power users. Availability is limited to two regions: US North Virginia and Europe Frankfurt, supporting cross-region inference.

There's also a point many find hard to adjust to: these models in Kiro operate with a hidden chain-of-thought; you can't see its reasoning steps, only the final result.

Those accustomed to watching the Agent reason step-by-step feel like throwing work into an opaque box.

The official statement is that this is expected behavior and doesn't affect output quality.

The three model tiers are clearly priced in Kiro.

When they first launched on July 14th, the same task cost 2.4x for Sol, 1.2x for Terra, and 0.6x for Luna.

After OpenAI's price reduction on July 30th took effect, Kiro followed the next day: Luna was slashed from 0.6x all the way to 0.1x, Terra reduced from 1.2x to 1.0x, with only Sol unchanged.

AWS's stance is clear: a development platform cannot be tied to just one model.

In the same selector, two cutting-edge models are beginning to undercut each other on price.

The evaluation criteria for models is also changing: a higher score no longer guarantees a win; spending less can also win.

For developers, the question used to be "How much per million tokens for this model?" Now it must be "How much will it actually cost me to get this thing done?"

References:

https://x.com/OpenAIDevs/status/2091966982015103068

https://openai.com/index/gpt-5-6-in-kiro/

This article is from WeChat Official Account "AI_era" (ID: AI_era), author: ASI Revelation, editor: Yuanyu

Tiền kỹ thuật số thịnh hành

Câu hỏi Liên quan

QAccording to the article, the cost of completing a successful task with GPT-5.6 Terra in Kiro decreased by 82%, but the official price reduction was only 20%. Where did the additional 60% cost saving come from?

AThe additional 60% cost saving primarily came from optimizations that reduced wasted token usage. These savings were achieved by minimizing the number of tokens generated, reducing unnecessary tool calls, and decreasing the frequency of failed attempts and inefficient detours. Essentially, the savings came from avoiding the costs associated with the AI agent making mistakes, choosing wrong paths, or having to backtrack, all of which would have incurred charges.

QWhat is the primary difference in how Kiro approaches coding tasks compared to a traditional AI agent, and how does this contribute to cost savings?

AKiro uses a spec-driven approach. Instead of letting the AI write code immediately from a vague user instruction, it first breaks down the instruction into a formal requirements document, technical design, and a list of executable tasks. This provides the model with a clear and structured job description upfront. This method saves costs by significantly reducing the expensive overhead of the model going off-track, needing rework, or starting over, which consumes a large number of tokens.

QWhat was a notable change in the Kiro platform's model offerings one year after its public preview, and what does this signify?

AA notable change was the introduction of three OpenAI GPT-5.6 models (Sol, Terra, Luna) into Kiro's model selector, where previously Anthropic's Claude models were dominant. This signifies a strategic move by AWS to avoid being tied to a single model provider on its development platform. It creates direct competition between leading models, which can drive performance improvements and price reductions for developers.

QHow does the article explain the layered efficiency improvements for AI agents like GPT-5.6 in platforms such as Kiro?

AThe article explains that efficiency gains come from three layers: 1) The agent framework that initiates requests and organizes context. 2) The orchestration system that schedules and dispatches these requests. 3) The model itself running on the GPU. Cost savings are achieved through optimizations at every one of these layers, not just the model's raw processing cost.

QAccording to the Terminal-Bench 2.1 data cited, why might a developer's choice of agent framework be as important as the choice of model itself for overall cost?

AThe Terminal-Bench 2.1 data shows that different frameworks paired with the same or similar models can result in vastly different costs. For example, Claude Code with Fable 5 achieved 83.8% accuracy at a cost of $552.67, while Codex with GPT-5.5 achieved 83.1% accuracy but at a much higher cost of $2059.19. This demonstrates that the framework's context organization, tool strategies, and workflow efficiency have a massive impact on the final bill, making the framework choice critically important.

Nội dung Liên quan

Anthropic nhắm đến "chip đào tạo"? Bị phanh phui từng muốn mua lại công ty chip AI MatX với giá 70 tỷ USD

Công ty trí tuệ nhân tạo (AI) Anthropic từng thảo luận mua lại startup chip AI MatX với giá khoảng 70 tỷ USD, nhưng thương vụ này cuối cùng không tiến triển và hai bên chuyển hướng sang đàm phán hợp tác tiềm năng. Động thái này cho thấy tham vọng tự phát triển chip, đặc biệt là chip cho quá trình huấn luyện mô hình, của Anthropic nhằm tìm kiếm năng lực tính toán nhanh hơn và rẻ hơn cho Claude. MatX, thành lập năm 2023 bởi các cựu kỹ sư Google, tập trung thiết kế chip phục vụ huấn luyện mô hình ngôn ngữ lớn (LLM). Việc Anthropic tiếp cận MatX cho thấy họ có thể hướng tới phát triển chip huấn luyện nội bộ, khác với hướng tiếp cận ban đầu nhấn mạnh chip suy luận (inference) của OpenAI. Để thúc đẩy kế hoạch chip tự thiết kế, Anthropic đang tích cực tuyển dụng nhân tài ngành bán dẫn, gần đây đã mời cựu lãnh đạo chủ chốt của dự án TPU Google là Amir Salek và cả cựu kỹ sư chip của OpenAI. Công ty cũng đang gặp gỡ nhiều startup chip AI khác để nghiên cứu các kiến trúc khác nhau. Tuy nhiên, Anthropic được cho là vẫn sẽ duy trì chiến lược đa dạng nhà cung cấp chip, tiếp tục hợp tác với các đối tác như Nvidia và Google, bên cạnh nỗ lực tự nghiên cứu. Việc tự phát triển chip là một dự án tốn kém và phức tạp, do đó việc xem xét mua lại một công ty chip có kinh nghiệm như MatX là cách tiếp cận hấp dẫn để tăng tốc lộ trình và có khả năng giảm chi phí lâu dài.

marsbit4 phút trước

Anthropic nhắm đến "chip đào tạo"? Bị phanh phui từng muốn mua lại công ty chip AI MatX với giá 70 tỷ USD

marsbit4 phút trước

Chủ tịch Fed Kevin Warsh “bồ câu diều hâu”, Goldman Sachs vẫn không tin “Fed tăng lãi suất vào tháng 9”, JPMorgan cho rằng “vẫn cần xem số liệu việc làm và CPI tháng 8”

Chủ tịch Fed Kevin Warsh đã có bài phát biểu mang tính diều hâu tại Jackson Hole, nhấn mạnh lạm phát là "mối quan tâm hàng đầu" và cho rằng xu hướng cơ bản chưa được cải thiện thực chất. Ông tuyên bố mục tiêu ổn định giá 2% là cố định và lãi suất ngắn hạn là công cụ chính để ứng phó. Thị trường phản ứng mạnh, lợi suất trái phiếu Kho bạc kỳ hạn 2 năm tăng và xác suất Fed tăng lãi suất vào tháng 9 tăng lên trên 50%. Tuy nhiên, các nhà kinh tế tại Ngân hàng JP Morgan và Goldman Sachs vẫn chưa thay đổi dự báo cơ bản của họ. JP Morgan duy trì dự báo tăng lãi suất vào tháng 12, cho rằng dữ liệu việc làm và CPI tháng 8 sắp công bố mới là yếu tố then chốt cho quyết định vào tháng 9. Goldman Sachs dự báo lạm phát cốt lõi tháng 8 tăng khoảng 0.2% và cho rằng FOMc khả năng cao sẽ giữ nguyên lãi suất, trừ khi dữ liệu CPI và PPI mạnh hơn dự kiến một cách bất ngờ. Warsh đánh giá nền kinh tế Mỹ "ấn tượng", chi tiêu tiêu dùng vẫn lành mạnh và đầu tư của doanh nghiệp tăng trưởng nhanh, với hơn một nửa mức tăng trong năm nay đến từ các khoản đầu tư liên quan đến AI. Ông cũng nhận định điều kiện tài chính hiện tại "khó có thể được mô tả là thắt chặt".

marsbit9 phút trước

Chủ tịch Fed Kevin Warsh “bồ câu diều hâu”, Goldman Sachs vẫn không tin “Fed tăng lãi suất vào tháng 9”, JPMorgan cho rằng “vẫn cần xem số liệu việc làm và CPI tháng 8”

marsbit9 phút trước

Tuần san Biên Tập Viên Chọn Lọc Weekly Editor's Picks (0822-0828)

Tóm tắt tuyển chọn biên tập hàng tuần (22-28/08): **Bối cảnh vĩ mô:** Phố Wall dự đoán Bộ Tài chính Mỹ có thể thay đổi cách thức phát hành trái phiếu, tập trung vào tín phiếu ngắn hạn và mở rộng mua lại để giảm áp lực lợi suất dài hạn. Vàng tăng vọt lên trên 4600 USD, với động lực từ ngân hàng trung ương, ETF và quyền chọn. **Đầu tư & Khởi nghiệp:** Arthur Hayes dự đoán ETH lên 30.000 USD, cho rằng tiền mã hóa là "van xả" cho chính sách tiền tệ. Cổ phiếu tiền mã hóa như MSTR, COIN tăng mạnh. Thị trường altcoin phục hồi, tổng vốn hóa vượt 1 nghìn tỷ USD. Các dự án như Circle (USDC), Hyperliquid (HYPE), Ethena (ENA) có động thái tích cực về cơ chế mua lại token. **AI & Lưu trữ:** Doanh thu NVIDIA sắp chạm 100 tỷ USD, với nhu cầu AI mở rộng từ các hãng lớn sang doanh nghiệp. SK Hynix công bố kế hoạch mua lại cổ phiếu 40 nghìn tỷ Won sau đợt điều chỉnh. **CeFi & DeFi:** Các dự án DeFi có thu nhập cao như UNI, AAVE được chú ý. Galaxy Digital ra mắt dịch vụ cho vay thế chấp bằng BTC, ETH, SOL. **Cơ hội Airdrop & Meme:** Tiếp tục các cơ hội tương tác trên PerpDEX. Thị trường meme sôi động với tin đồn quanh Trump. **Ethereum:** BitMine đang mua tích lũy ETH, có thể nắm giữ 5% tổng cung. Tom Lee (Fundstrat) nhận định điều này không có rủi ro quản trị và kỳ vọng ETH lên 10.000 USD. **Tin tức nổi bật tuần:** BTC trở lại 80.000 USD. Dự luật tiền mã hóa Clarity có thể được thông qua. Grayscale đẩy mạnh chuyển đổi trust ZEC, TAO thành ETF. Nhiều vụ lừa đảo và tranh chấp pháp lý trong ngành.

marsbit17 phút trước

Tuần san Biên Tập Viên Chọn Lọc Weekly Editor's Picks (0822-0828)

marsbit17 phút trước

Giám đốc Strive: Hiểu lại bánh xe giá của Bitcoin

Bài viết của Joe Burnett, Phó Chủ tịch Chiến lược Bitcoin tại Strive, phân tích sự phát triển của Bitcoin thông qua lăng kính "bánh xe giá". Ông lập luận rằng mô hình quy luật lũy thừa (power law) hiện tại, cho thấy lợi nhuận giảm dần khi Bitcoin trưởng thành, chỉ là Giai đoạn 2. Sự trưởng thành này dẫn đến biến động giảm, cải thiện tỷ lệ Sharpe và biến Bitcoin thành tài sản thế chấp hấp dẫn hơn. Biến động thấp hơn cho phép các nhà đầu tư hiện có phân bổ vốn nhiều hơn và quan trọng là mở rộng quy mô tín dụng có thể được hỗ trợ bởi Bitcoin. Khi rủi ro tín dụng giảm, hệ thống tài chính có thể cung cấp nhiều khoản vay hơn với chi phí thấp hơn (ví dụ: thông qua cho vay thế chấp, trái phiếu). Điều này tạo ra một vòng lặp tự củng cố: giá tăng làm tăng giá trị tài sản thế chấp, giải phóng không gian vay nhiều hơn để mua thêm Bitcoin, lại đẩy giá lên. Burnett so sánh điều này với Giai đoạn 3 trong đường cong mỏi kim loại, nơi một vết nứt tăng tốc đột ngột trước khi gãy. Ở đây, "vết nứt" là sự mở rộng tín dụng bằng USD, "vật liệu" là hệ thống tín dụng USD, và "sự gãy" là thời điểm giá Bitcoin tăng tốc vượt khỏi đường quy luật lũy thừa, do vòng xoáy kết hợp giữa vốn tự có và tín dụng mới đua nhau tranh giành nguồn cung Bitcoin cố định.

marsbit37 phút trước

Giám đốc Strive: Hiểu lại bánh xe giá của Bitcoin

marsbit37 phút trước

Đột ngột, OpenAI cắt nguồn cung hoàn toàn cho Cursor

Theo thông báo từ OpenAI, việc hợp tác cung cấp mô hình trực tiếp cho Cursor sẽ chấm dứt hoàn toàn vào ngày 12/11. Quyết định này được đưa ra sau khi công ty công cụ lập trình AI này được Elon Musk mua lại với giá 60 tỷ USD vào giữa tháng 8. Lý do chính được OpenAI viện dẫn là lo ngại về lịch sử vi phạm điều khoản dịch vụ của các công ty thuộc sở hữu Musk. Đáng chú ý, thế hệ mô hình mạnh nhất tiếp theo của OpenAI, Astra, với khả năng bảo mật cấp độ "nghiêm trọng", sẽ không bao giờ được cung cấp cho Cursor. Đối với hàng triệu nhà phát triển sử dụng Cursor, họ có 75 ngày để chuyển đổi. Sau thời hạn này, cách duy nhất để tiếp tục dùng các mô hình của OpenAI trong Cursor là sử dụng API Key cá nhân, đồng nghĩa với việc chuyển từ hình thức trả phí định kỳ sang tự thanh toán theo lưu lượng sử dụng, có thể dẫn đến chi phí cao hơn. Phản ứng lại, CEO của Cursor, Michael Truell, cho biết chỉ khoảng 5% lưu lượng người dùng sử dụng mô hình OpenAI và họ đang đàm phán để giải quyết vấn đề, đồng thời nhấn mạnh mong muốn nền tảng AI duy trì tính trung lập. Sự kiện này là một bước phát triển mới trong xu hướng phân mảnh hệ sinh thái AI, nơi các gã khổng lồ công nghệ ngày càng siết chặt kiểm soát đối với việc phân phối các mô hình trí tuệ nhân tạo tiên tiến của mình.

marsbit1 giờ trước

Đột ngột, OpenAI cắt nguồn cung hoàn toàn cho Cursor

marsbit1 giờ trước

Giao dịch

Giao ngay

Bài viết Nổi bật

Làm thế nào để Mua BILL

Chào mừng bạn đến với HTX.com! Chúng tôi đã làm cho mua Billions Network (BILL) trở nên đơn giản và thuận tiện. Làm theo hướng dẫn từng bước của chúng tôi để bắt đầu hành trình tiền kỹ thuật số của bạn.Bước 1: Tạo Tài khoản HTX của BạnSử dụng email hoặc số điện thoại của bạn để đăng ký tài khoản miễn phí trên HTX. Trải nghiệm hành trình đăng ký không rắc rối và mở khóa tất cả tính năng. Nhận Tài khoản của tôiBước 2: Truy cập Mua Crypto và Chọn Phương thức Thanh toán của BạnThẻ Tín dụng/Ghi nợ: Sử dụng Visa hoặc Mastercard của bạn để mua Billions Network (BILL) ngay lập tức.Số dư: Sử dụng tiền từ số dư tài khoản HTX của bạn để giao dịch liền mạch.Bên thứ ba: Chúng tôi đã thêm những phương thức thanh toán phổ biến như Google Pay và Apple Pay để nâng cao sự tiện lợi.P2P: Giao dịch trực tiếp với người dùng khác trên HTX.Thị trường mua bán phi tập trung (OTC): Chúng tôi cung cấp những dịch vụ được thiết kế riêng và tỷ giá hối đoái cạnh tranh cho nhà giao dịch.Bước 3: Lưu trữ Billions Network (BILL) của BạnSau khi mua Billions Network (BILL), lưu trữ trong tài khoản HTX của bạn. Ngoài ra, bạn có thể gửi đi nơi khác qua chuyển khoản blockchain hoặc sử dụng để giao dịch những tiền kỹ thuật số khác.Bước 4: Giao dịch Billions Network (BILL)Giao dịch Billions Network (BILL) dễ dàng trên thị trường giao ngay của HTX. Chỉ cần truy cập vào tài khoản của bạn, chọn cặp giao dịch, thực hiện giao dịch và theo dõi trong thời gian thực. Chúng tôi cung cấp trải nghiệm thân thiện với người dùng cho cả người mới bắt đầu và người giao dịch dày dạn kinh nghiệm.

Tổng lượt xem 748Xuất bản vào 2026.05.07Cập nhật vào 2026.06.02

Làm thế nào để Mua BILL

Thảo luận

Chào mừng đến với Cộng đồng HTX. Tại đây, bạn có thể được thông báo về những phát triển nền tảng mới nhất và có quyền truy cập vào thông tin chuyên sâu về thị trường. Ý kiến ​​của người dùng về giá của BILL (BILL) được trình bày dưới đây.

活动图片