AI Agent Outputs Garbage? The Problem Is You're Not Willing to Burn Enough Tokens

marsbitXuất bản vào 2026-03-23Cập nhật gần nhất vào 2026-03-23

Tóm tắt

The core argument is that the quality of an AI Agent's output is directly proportional to the number of tokens invested in the process. More tokens lead to fewer errors, as they allow for deeper reasoning, multiple independent attempts, self-critique from fresh contexts, and verification through testing. This approach can solve problems of scale and complexity but fails when facing novel problems not present in the model's training data. For such novel challenges, human domain knowledge and guidance are essential. Two practical, immediate solutions are proposed: implementing an automatic review cycle (WAIT) for the Agent to repeatedly critique and fix its work, and establishing frequent verification checkpoints (VERIFY) where a separate Agent validates outputs to catch errors early. The key takeaway is that insufficient token investment is often the primary reason for poor Agent performance, not the underlying framework.

Author: Systematic Long Short

Compiled by: Deep Tide TechFlow

Deep Tide Intro: The core argument of this article is just one sentence: The quality of an AI Agent's output is directly proportional to the number of Tokens you invest.

The author isn't speaking in general theoretical terms; instead, they provide two specific methods you can start using today and clearly define the boundary where throwing more Tokens won't help—the "novelty problem."

For readers currently using Agents to write code or run workflows, the information density and practicality are very high.

Introduction

Alright, you have to admit the title is quite eye-catching—but seriously, it's no joke.

In 2023, when we were using LLMs to run production code, people around us were stunned because the common belief at the time was that LLMs could only produce unusable garbage. But we knew something others didn't: the output quality of an Agent is a function of the number of Tokens you invest. It's that simple.

Run a few experiments yourself and you'll see. Have an Agent complete a complex, somewhat niche programming task—for example, implementing a convex optimization algorithm with constraints from scratch. First, run it at the lowest thinking level; then switch to the highest thinking level and have it review its own code to see how many bugs it can find. Try the medium and high levels. You'll visually see: the number of bugs decreases monotonically as the number of Tokens invested increases.

This isn't hard to understand, right?

More Tokens = Fewer errors. You can take this logic a step further; this is essentially the (simplified) core idea behind code review products. In a completely new context, invest a massive number of Tokens (for example, have it parse the code line by line, judging whether each line has a bug)—this can basically catch the vast majority, if not all, bugs. This process can be repeated ten times, a hundred times, each time examining the codebase from a "different angle," and you can eventually unearth all the bugs.

There's another empirical support for the view that "burning more Tokens improves Agent quality": those teams that claim to use Agents to write code from start to finish and push it directly to production are either the foundational model providers themselves or extremely well-funded companies.

So, if you're still struggling to get production-level code from your Agent—to be blunt, the problem lies with you. Or rather, with your wallet.

How to Tell If You're Burning Enough Tokens

I wrote an entire article saying the problem definitely isn't your framework (harness), that "keeping it simple" can still produce excellent results, and I still stand by that view. You read that article, followed the advice, but were still greatly disappointed by the Agent's output. You sent me a DM, saw I read it but didn't reply.

This article is the reply.

Your Agent performs poorly and can't solve the problem, most of the time, simply because you're not burning enough Tokens.

How many Tokens are needed to solve a problem depends entirely on the problem's scale, complexity, and novelty.

"What's 2+2?" doesn't require many Tokens.

"Write me a bot that scans all markets between Polymarket and Kalshi, finds markets that are semantically similar and should settle around the same event, sets no-arbitrage boundaries, and automatically trades with low latency whenever an arbitrage opportunity arises"—this requires burning a huge pile of Tokens.

We found something interesting in practice.

If you invest enough Tokens to handle problems caused by scale and complexity, the Agent *will* solve them, no matter what. In other words, if you want to build something extremely complex, with many components and lines of code, as long as you throw enough Tokens at these problems, they will eventually be completely resolved.

There is one small but important exception.

Your problem cannot be too novel. At this stage, no amount of Tokens can solve the "novelty" problem. Enough Tokens can reduce errors from complexity to zero, but they cannot make an Agent invent something it doesn't know out of thin air.

This conclusion actually came as a relief to us.

We spent enormous effort, burned—a lot, a lot, a whole lot—of Tokens, trying to see if an Agent could reconstruct an institutional investment process with almost no guidance. This was partly to figure out how many years we (as quantitative researchers) have before being completely replaced by AI. It turned out the Agent couldn't get anywhere close to a decent institutional investment process. We believe this is partly because they have never seen such a thing—meaning, institutional investment processes simply don't exist in the training data.

So, if your problem is novel, don't count on solving it by stacking Tokens. You need to guide the exploration process yourself. But once you've defined the implementation plan, you can confidently stack Tokens for execution—no matter how large the codebase or how complex the components, it's not a problem.

Here's a simple heuristic: the Token budget should grow proportionally with the number of lines of code.

What Are the Extra Tokens Actually Doing?

In practice, additional Tokens typically improve the Agent's engineering quality in the following ways:

Allowing it to spend more time reasoning in the same attempt, giving it a chance to discover flawed logic itself. Deeper reasoning = better planning = higher probability of success on the first try.

Allowing it to make multiple independent attempts, exploring different solution paths. Some paths are better than others. Allowing more than one attempt lets it choose the best one.

Similarly, more independent planning attempts allow it to abandon weak directions and keep the most promising ones.

More Tokens allow it to critique its previous work with a fresh context, giving it a chance to improve instead of being stuck in a certain "reasoning inertia."

And, of course, my favorite: more Tokens mean it can use tests and tools for verification. Actually running the code to see if it works is the most reliable way to confirm the answer is correct.

This logic works because engineering failures of Agents are not random. They are almost always due to choosing the wrong path too early, not checking if this path actually works (early on), or not having enough budget to recover and backtrack after discovering a mistake.

That's the story. Tokens are literally the decision quality you buy. Think of it like research work: if you ask a person to answer a difficult question on the spot, the quality of the answer decreases as time pressure increases.

Research, at its core, is what produces the foundational "knowing the answer." Humans spend biological time to produce better answers; Agents spend more compute time to produce better answers.

How to Improve Your Agent

You might still be skeptical, but there are many papers supporting this, and honestly, the very existence of the "reasoning" adjustment knob is all the proof you need.

One paper I particularly like: researchers trained on a small, carefully curated set of reasoning examples, then used a method to force the model to keep thinking when it wanted to stop—specifically by appending "Wait" where it wanted to stop. This single change raised a certain benchmark from 50% to 57%.

I want to be as clear as possible: if you've been complaining that the code written by your Agent is mediocre, the single highest thinking level is likely still not enough for you.

I'll give you two very simple solutions.

Simple Method One: WAIT

The simplest thing you can start doing today: set up an automatic loop—after building, have the Agent review its work N times with a fresh context, fixing any issues found each time.

If you find this simple trick improves your Agent's engineering results, then you at least understand that your problem is just a matter of Token quantity—welcome to the Token burning club.

Simple Method Two: VERIFY

Have the Agent verify its own work early and often. Write tests to prove that the chosen path actually works. This is especially useful for highly complex, deeply nested projects—a function might be called by many other downstream functions. Catching errors upstream can save you a lot of subsequent compute time (Tokens). So, if possible, set up "verification checkpoints" throughout the entire build process.

Finished writing a piece? The main Agent says it's done? Have a second Agent verify it. Unrelated thought streams can cover sources of systematic bias.

That's basically it. I could write a lot more on this topic, but I believe just realizing these two things and implementing them well can solve 95% of your problems. I firmly believe in doing simple things extremely well, then adding complexity as needed.

I mentioned that "novelty" is a problem that can't be solved with Tokens, and I want to emphasize it again because you will eventually hit this pitfall and come crying to me saying stacking Tokens didn't work.

When the problem you want to solve isn't in the training set, *you* are the one who really needs to provide the solution. Therefore, domain expertise remains extremely important.

Câu hỏi Liên quan

QWhat is the core argument of the article regarding AI Agent output quality?

AThe core argument is that the quality of an AI Agent's output is directly proportional to the number of tokens you are willing to invest in the process.

QAccording to the article, what is the one type of problem that cannot be solved by simply using more tokens?

AProblems that are 'novel' or not present in the model's training data cannot be solved by any amount of tokens; they require human guidance and domain expertise.

QWhat are the two simple methods suggested in the article to immediately improve an AI Agent's performance?

AThe two simple methods are: 1. WAIT - Set up an automatic loop for the Agent to review its work multiple times with a fresh context and fix any issues found. 2. VERIFY - Have the Agent (or a second one) verify its work early and often by writing tests to prove the chosen path works.

QHow does the article suggest thinking about the relationship between tokens and decision quality?

AThe article suggests thinking of tokens as literally 'buying' decision quality, analogous to how human researchers spend biological time to produce better answers, an AI Agent spends computational time (tokens) to produce better answers.

QWhat heuristic does the article provide for determining a sufficient token budget for a task?

AThe article provides a simple heuristic: the token budget should grow proportionally with the number of lines of code required for the task.

Nội dung Liên quan

9,42 triệu nhà đầu tư nhỏ lẻ tranh mua cổ phiếu Trường Tân Công nghệ, ai đã trúng thưởng?

Kết quả phát hành cổ phiếu lần đầu ra công chúng (IPO) của Trường Tân Công nghệ đã chính thức được công bố. Tổng cộng có 942,88 nghìn nhà đầu tư nhỏ lẻ và 285 tổ chức với 10.907 tài khoản đã tham gia đăng ký mua. Tỷ lệ trúng đấu giá trực tuyến cuối cùng là khoảng 0,4714%, mức cao kỷ lục đối với cổ phiếu mới trên sàn Sci-Tech Innovation Board (STAR Market), với khoảng 770 nghìn số trúng thưởng. Sau khi cơ chế chuyển bán lại được kích hoạt, số cổ phiếu phát hành trực tuyến đã tăng lên 3,851 tỷ cổ, chiếm 50,07% tổng số cổ phát hành. Mỗi lô trúng thưởng (500 cổ) dự kiến cần nộp 4.330 nhân dân tệ. Lợi nhuận kỳ vọng từ việc mua cổ phiếu mới dao động từ khoảng 10.600 đến 25.600 nhân dân tệ tùy thuộc vào định giá thị trường sau niêm yết. Về phân bổ cho các tổ chức, 285 tổ chức đã được phân bổ 2,173 tỷ cổ với tỷ lệ khoảng 0,1756%. Tài khoản Bảo hiểm Tai Kang là tổ chức được phân bổ nhiều nhất. Trong số các quỹ công chúng, Yi Fang Da, Nan Fang Fund và Công Thương Ngân hàng Ruixin dẫn đầu. Đáng chú ý, Lương Văn Phong, người sáng lập công ty mô hình lớn DeepSeek, thông qua hai thực thể định lượng tư nhân do ông kiểm soát, đã nhận được phân bổ lớn nhất trong số các quỹ tư nhân với tổng số tiền khoảng 175 triệu nhân dân tệ. Nếu vốn hóa thị trường của Trường Tân Công nghệ đạt 3 nghìn tỷ nhân dân tệ, khoản đầu tư này có thể mang lại lợi nhuận khoảng 730 triệu nhân dân tệ. 30 nhà đầu tư chiến lược, bao gồm Quỹ An sinh Xã hội Quốc gia, đã được phân bổ tổng cộng 14,437 tỷ nhân dân tệ với thời gian khóa từ 12 đến 36 tháng. Thị trường kỳ vọng Trường Tân Công nghệ sẽ chính thức lên sàn vào ngày 27/7 và có khả năng trở thành cổ phiếu công nghệ có vốn hóa thị trường lớn nhất trên sàn A. Các ước tính định giá chủ yếu ở mức trên 1 nghìn tỷ nhân dân tệ, một số dự báo lạc quan lên tới 3-4 nghìn tỷ. Tuy nhiên, đợt điều chỉnh mạnh gần đây của cổ phiếu công nghệ toàn cầu có thể ảnh hưởng đến biểu hiện giá cổ phiếu sau niêm yết.

marsbit48 phút trước

9,42 triệu nhà đầu tư nhỏ lẻ tranh mua cổ phiếu Trường Tân Công nghệ, ai đã trúng thưởng?

marsbit48 phút trước

Gần 3 triệu fan hâm mộ, "nữ thần thiện nguyện" đều do AI tổng hợp, làm giả trại mồ côi xuyên biên giới, sự nghiệp "làm từ thiện giả" sụp đổ chỉ sau một đêm

Cô gái người Úc có tên Lily Jay (tên thật Lily Jay Hinson) đã gây chấn động mạng xã hội khi bị phát hiện sử dụng AI để tạo ra một vụ lừa đảo từ thiện quy mô lớn. Với gần 3 triệu người theo dõi trên Instagram, cô xây dựng hình ảnh một tín đồ Hồi giáo ngoan đạo và tích cực thông qua "Quỹ Lily Jay", tuyên bố xây dựng nhà thờ Hồi giáo, cứu trợ trẻ mồ côi ở Uganda, Sudan, Nepal và phân phát bánh mì cho người tị nạn ở Gaza. Tuy nhiên, điều tra từ ABC News Verify đã vạch trần hàng loạt bằng chứng giả mạo: video khánh thành trại trẻ mồ côi ở Uganda với những đứa trẻ cầm kẹo, biểu ngữ và chính người phụ nữ trong video đều do AI tạo ra, kể cả chi tiết lỗi chính tả trên áo phông. Hình ảnh nhận giải thưởng nhân đạo cũng chứa watermark của ChatGPT. Quỹ này thậm chí không đăng ký hoạt động hợp pháp tại Uganda và không có tên trong sổ đăng ký từ thiện ở Úc, đồng thời đã giấu một dòng tuyên bố "không phải là tổ chức từ thiện" trên website. Hoạt động của quỹ đầy nghi vấn: trụ sở đặt ở Kosovo, cá nhân Lily Jay sống ở Cyprus và không phải là giám đốc quỹ. Sau khi bị ABC chất vấn, trang web đã gỡ video giả và nút quyên góp đối với truy cập từ Úc, nhưng vẫn để mở cho người dùng quốc tế. Chuyên gia cảnh báo đây là kiểu lừa đảo nguy hiểm, lợi dụng lòng trắc ẩn và sự tin tưởng. Vụ việc gióng lên hồi chuông cảnh tỉnh về việc AI có thể bị lạm dụng để tạo ra những câu chuyện giả tưởng hoàn hảo, đánh cắp sự thiện nguyện của công chúng trong thời đại số.

marsbit1 giờ trước

Gần 3 triệu fan hâm mộ, "nữ thần thiện nguyện" đều do AI tổng hợp, làm giả trại mồ côi xuyên biên giới, sự nghiệp "làm từ thiện giả" sụp đổ chỉ sau một đêm

marsbit1 giờ trước

L2 'Tái hiệu chuẩn': Khi L1 Trở Thành Rollup Của Chính Mình, Cục Diện Cuối Cùng Của Ethereum Là Gì?

"L2 Hiệu chỉnh lại": Khi L1 trở thành Rollup của chính nó, tương lai cuối cùng của Ethereum là gì? Cộng đồng Ethereum từng lo lắng về việc L2 làm xói mòn giá trị của L1 và phá vỡ khả năng kết hợp toàn cầu. Bài viết phân tích sự điều chỉnh mối quan hệ giữa L1 và L2 trong bối cảnh mở rộng quy mô của Ethereum. Định vị mới của L2: Với việc L1 tự nâng cao khả năng xử lý (tăng Gas Limit, zkEVM...), vai trò chính của L2 không còn đơn thuần là cung cấp không gian giao dịch rẻ hơn. Thay vào đó, L2 sẽ chuyển sang cung cấp các chức năng khác biệt mà L1 khó đáp ứng thống nhất, như tối ưu hóa ứng dụng cụ thể, tính riêng tư và mô hình quản trị linh hoạt. L2 sẽ trở thành một dải phổ liên tục từ các Rollup kế thừa bảo mật tối đa của Ethereum đến các môi trường thực thi độc lập hơn. Tương tác & Kết hợp lại: Sự phân mảnh giữa các L2 gây ra vấn đề về thanh khoản và trải nghiệm người dùng. Giải pháp nằm ở việc cải thiện khả năng tương tác, không chỉ là "một nút chuyển chuỗi", mà là làm cho các trạng thái giữa các môi trường thực thi có thể tin cậy lẫn nhau nhanh hơn, thông qua các khung công việc Intent, Lớp Tương tác Ethereum (EIL) và việc rút ngắn đáng kể thời gian xác nhận cuối cùng (finality). L1 như "Rollup của chính nó": Với sự phát triển của hệ thống chứng minh (như zkEVM), trong tương lai, các trình xác thực L1 có thể xác minh trạng thái thông qua bằng chứng mật mã thay vì thực thi lại mọi giao dịch. Điều này chia sẻ kiến trúc "thực thi-tách biệt-xác minh" với Rollup, làm mờ ranh giới truyền thống giữa L1 và L2. Các L2 tiên tiến (Native Rollup) có thể kế thừa trực tiếp hơn khả năng xác minh từ giao thức L1. Tóm lại, tương lai của Ethereum không phải là L1 thay thế L2 hay ngược lại, mà là một hệ sinh thái gồm nhiều môi trường thực thi (L2) đa dạng về chức năng và hiệu suất, nhưng có thể chia sẻ nền tảng bảo mật, thanh khoản và quan hệ trạng thái, từ đó tái tạo lại trải nghiệm "một chuỗi" thống nhất cho người dùng.

marsbit1 giờ trước

L2 'Tái hiệu chuẩn': Khi L1 Trở Thành Rollup Của Chính Mình, Cục Diện Cuối Cùng Của Ethereum Là Gì?

marsbit1 giờ trước

Giao dịch

Giao ngay
活动图片