From TPU to Self-Evolving Agents: How Jeff Dean Predicts the Next Step in AI

marsbitPublished on 2026-08-03Last updated on 2026-08-03

Abstract

At the 2026 YC Startup School, Jeff Dean outlined his vision for AI's next phase, shifting focus from simply scaling models to building intelligent, autonomous systems. He believes AI's progress is no longer just about creating smarter models, but about integrating them into systems capable of long-term, iterative work, automated experimentation, and continuous learning. This evolution moves the competition from "who has the bigger model" to "who can best organize intelligence." Dean suggests AI capabilities are now comparable to a junior engineer, enabling the automation of complex workflows. However, the true challenge and opportunity lie in managing these AI "workers" at scale. He emphasizes the importance of **context engineering**—structuring tools, memory, and feedback loops—over raw model power. For startups, this means building deep expertise in niche domains where general models currently fail (near 0-1% success rates), leveraging proprietary data, specialized tools, and domain-specific evaluators. A recurring theme is re-examining fundamental constraints. Dean's past work, like moving Google's search index to memory or creating the TPU, stemmed from questioning outdated assumptions about hardware and cost. He sees similar inflection points today, particularly in **specialized inference hardware** to drastically reduce latency and energy consumption for real-time Agent operation. Notably, he points out that in modern AI systems, the dominant cost is often not compu...

At the 2026 YC Startup School, Jeff Dean’s voice sounded a bit hoarse.

Right at the beginning of the interview, he explained that he had lost his voice and sounded different than usual. But this didn’t affect the audience’s attention. Sitting opposite him, YC partner Diana Hu listed a series of names that could easily be written into the history of computing: MapReduce, BigTable, TensorFlow, TPU, Gemini.

Any one of these projects could be the career-defining work of an engineer. Yet they all appear on the resumes of Jeff Dean and a group of Google engineers around him.

Diana didn’t turn the interview into a review of achievements. She was more concerned with another question: now that generative AI has swept through the software industry, what exactly is someone like Jeff Dean, who is best at rearchitecting systems from the ground up, looking at today?

The answer isn’t bigger models.

In this nearly hour-long conversation, Jeff Dean repeatedly talked about inference hardware, energy, data movement, context engineering, long-running agents, automated experiment systems, and how startups can avoid head-on competition with general-purpose models. What he discussed seemed scattered, but there was a very clear thread running through it: the next stage of AI isn’t just about training smarter models, but about placing models into systems that can work long-term, continuously trial-and-error, automatically validate, and constantly accumulate capabilities.

This also means that AI competition is shifting from "who has the bigger model" to "who can better organize intelligence".

I. AI is Already Like a Junior Engineer, But That's Not the Most Important Change

In May 2025, Jeff Dean made a widely discussed judgment: AI’s capability is already close to that of a junior engineer.

A year later, Diana asked him, how did that prediction turn out?

Jeff Dean’s answer was straightforward. He believed the judgment was "quite accurate." Progress in agentization, long-sequence coding, and complex tasks was even faster than he had anticipated at the time.

"The ability of models to complete increasingly complex tasks is growing faster than I expected," he said.

More notably, this capability is no longer limited to writing code. More and more agent systems are entering scientific, engineering, and other professional fields. They don't just answer questions; they break down tasks, use tools, run experiments, read results, and then act based on feedback.

Comparing AI to a junior engineer easily draws attention to labor replacement. But Jeff Dean is more concerned with another layer of change: when a "junior engineer" can be replicated dozens or hundreds of times, working in parallel for days or even weeks, how will the organization of production change?

In traditional teams, junior engineers need to get up to speed on the business, understand the tools, and receive constant feedback. The same goes for agents. Except their training material is no longer just documentation, but prompts, tool specifications, skill files, testing frameworks, evaluators, and the entire context environment.

This creates a new division of labor in AI engineering.

In the past, engineers were mainly responsible for writing code. In the future, more engineers will be responsible for defining problems, setting up environments, writing specifications, designing feedback loops, and then orchestrating a group of agents to complete tasks.

Jeff Dean’s prediction for 2027 is exactly that. He believes machine learning systems will increasingly participate in improving machine learning systems themselves. They will break goals into sub-problems, automatically run numerous experiments, compare results, and combine effective solutions to form stronger new systems.

"Whenever a field has a measurable objective, there's an opportunity to make significant progress."

This sentence is the first key to the entire interview.

The first areas AI automation will invade aren’t necessarily the ones with the most knowledge, but those with the clearest feedback. Does the code pass the test? Can the chip layout reduce area? Can the model architecture improve accuracy? Does the material property meet requirements? These questions all have relatively clear evaluation criteria. As long as the evaluator is reliable enough, machines can experiment with extremely high frequency.

Therefore, the truly important unit in the AI era may no longer be a single answer, but a complete closed loop: propose a solution, execute it, measure the result, adjust direction.

II. What Changed Google Search Was an Arithmetic Problem

Many of Jeff Dean’s representative works stem from a very simple starting point: first calculate the order of magnitude.

In 2001, Google Search still heavily relied on hard drives. Hard drives had large capacity but slow access speeds. Jeff Dean and Sanjay Ghemawat did an estimate and found that Google’s entire search index at the time could already fit into the memory of all its servers.

Today, this sounds like just an upgrade in storage media. But back then, it meant a completely different system design.

If the index mainly resided on hard drives, queries had to wait for mechanical seek times. By moving the index into memory, access latency could plummet. The two quickly wrote a new version and put it into production within days. Google Search became noticeably faster as a result.

This story is most easily packaged as a flash of genius inspiration. Jeff Dean’s telling, however, is more like an engineer stating common sense: the system conditions changed, a previously unworkable solution suddenly became viable, so it should be recalculated.

Many industry innovations happen at such moments.

An old problem persists for a long time, and people get used to patching around it. Later, hardware prices, memory capacity, network bandwidth, or model capability cross a certain threshold, and the old constraints disappear. Yet most people still use the old architecture because it has become common sense.

What Jeff Dean is good at is turning common sense back into a hypothesis.

He asks: Why must it be this way? Are today’s orders of magnitude still the same as yesterday’s? If we replace the most expensive step, could the entire system take on a completely different shape?

This is also his advice to entrepreneurs. Don't just look at where current solutions fall short, but re-examine the problem from first principles. Can performance be improved by an order of magnitude? Can cost be reduced by two orders of magnitude? Can we stop following the industry's default implementation path?

"Sometimes, you just need to squint at a problem, not be anchored by today's solutions, but think from first principles about how it should be solved."

This doesn’t sound mysterious. The real difficulty is that most people, upon entering an industry, quickly learn all its default answers. Experience helps people become more efficient, but it can also make them lose the ability to ask questions anew.

III. Why Three Minutes of Speech Gave Birth to the TPU

In 2013, Google’s deep learning speech recognition started significantly outperforming the old system. The error rate was cut in half, equivalent to twenty years of progress in speech recognition concentrated into a few months.

The product team was, of course, excited. Jeff Dean first did the math.

If speech recognition truly got better, users would be more willing to use it. Assuming each Google user only used three minutes of speech recognition per day, how many servers would Google need to support that?

The result wasn’t optimistic. Based on the efficiency of CPUs at the time, Google might need to double its server fleet.

This was the origin of the TPU.

It wasn’t because a research team suddenly wanted to build a chip, nor to prove Google could do hardware. It was because a successful model was about to create an unsustainable service cost.

This history reveals an often-overlooked pattern in AI products: improving model performance doesn’t always reduce cost. On the contrary, the better the performance, the greater the usage, and the heavier the system pressure.

When speech recognition wasn’t useful, users rarely invoked it. System cost wasn’t an issue. When the error rate plummeted, demand was suddenly unleashed, and the previously hidden compute constraint surfaced.

The path the TPU chose was to build specialized hardware for the most central computing patterns of machine learning. It didn’t need to run a browser or handle all general-purpose programs. It was primarily good at low-precision, dense linear algebra. This type of computation happened to be at the heart of modern machine learning.

The first-generation TPU ultimately delivered order-of-magnitude benefits. According to Jeff Dean, it was 30 to 80 times more energy-efficient and had 20 to 30 times lower latency than CPUs and GPUs of the time.

There’s another easily overlooked design consideration here.

The TPU was specialized, but not so specialized it could only run one fixed model. The team knew machine learning algorithms would continue to evolve quickly, so they designed the chip as a somewhat general-purpose linear algebra system. It sacrificed the ability to run Chrome or Word but preserved the space to support future algorithmic changes.

This is a difficult balance to strike. Not specialized enough, and the benefits aren’t obvious. Too specialized, and the hardware becomes obsolete when the algorithm changes.

Jeff Dean’s view on today’s inference hardware clearly echoes the TPU’s story. He believes the next wave of important opportunities still lies in specialization, but the focus will shift further toward low-latency, low-energy inference.

"Imagine what you could do if latency improved by 50 times."

When model replies take over ten seconds, people treat it as an occasional consultation tool. When latency is near-instantaneous, it can truly enter interactive interfaces, robots, real-time video, operating systems, and continuous decision-making processes.

Waiting isn’t just a minor UX issue. Waiting changes product forms.

IV. The Cost Center of AI Isn't Computation, It's Moving Data

If we were to update "Latency Numbers Every Programmer Should Know" for AI engineers in 2026, Jeff Dean believes the focus should shift from hard drive seek times, cache misses, and intercontinental network latency to data flow inside the chip.

Engineers need to know: What is the bandwidth from main memory to on-chip memory? From on-chip memory to the multiplier unit? How much energy does one multiplication consume? How do chips interconnect? When scaling from 500 chips to 10,000, how does network efficiency degrade?

These numbers seem far from products but actually determine which products are viable.

Jeff Dean gave a striking ratio. Performing a single mathematical multiplication requires about one picojoule of energy. Moving data from high-bandwidth memory to the compute unit can cost about 1000 times more energy.

In other words, the expensive action in today’s AI systems often isn’t "computing" but "moving the stuff to be computed."

This also explains why batching is so important.

After a set of model weights is moved from memory into the compute unit, if it only processes one token, the entire data movement cost is borne by that single token. If a larger batch is processed simultaneously, the same set of weights can serve more computations, amortizing the energy and bandwidth cost.

But batching inherently conflicts with low latency. To gather enough requests for a batch, the system often has to wait. Throughput increases, but individual user responses may slow down.

Therefore, many problems that seem to belong to the model layer are actually hardware and system problems. Why use large batches for training? Why does inference need KV Cache? Why do models pursue low precision? Why do systems need quantization? The answers all lie in data movement and energy constraints.

Jeff Dean’s recent focus on inference is precisely because inference is extremely sensitive to latency. If a training task runs a bit slower, it often just means the experiment ends later. If an inference task waits an extra second, it directly impacts user experience and agent efficiency.

If an agent needs to call a model 1000 times consecutively, a 50% reduction in single-call latency could make a huge difference in the total task completion time. Not to mention future agents running for days or weeks.

Therefore, AI’s "energy problem" isn’t a distant environmental issue. It directly determines whether models can serve more people cheaply, whether agents can run continuously, and whether a startup’s gross margin can be healthy.

V. The Model is Just a Component; Context is the Agent's Workspace

Over the past few years, the AI industry has grown accustomed to measuring progress by parameter count, training data, and benchmark scores. In 2026, Jeff Dean emphasizes everything around the model more.

A truly useful AI system, besides the model, needs retrieval, tools, memory, historical information, execution environments, and feedback mechanisms. The model needs to know what tools are available, when to call them, how to break down complex problems into a sequence of actions, and be able to compare multiple plans to judge which is more likely to succeed.

This is why "context engineering" is starting to take center stage.

Jeff Dean says the information a model sees during training is ultimately "stirred" into hundreds of billions or trillions of parameters. It’s like a thick soup—knowledge is present but not necessarily clear. Information actually placed in the current context is more direct and easier for the model to use accurately.

This leaves an important opportunity for small teams.

Training foundation models requires massive capital, data, and compute. Context engineering can start with an API. Entrepreneurs can organize domain knowledge, tool workflows, customer data, and evaluation standards around a specific business, making a general-purpose model perform more reliably in a narrow scenario.

Jeff Dean gave a personal example.

He and Sanjay Ghemawat often optimize Google’s internal low-level libraries. These data structures might run across millions of processes; tiny performance differences are amplified by scale. The traditional approach is for an engineer to write microbenchmarks, measure current performance, modify the code, rerun benchmarks, observe cache usage and performance changes, and iterate.

The two encoded this workflow into an agent skill. The model learned how to run benchmarks, modify code, compare results, and continue optimizing based on measurements.

"We just took the method a human would use and gave it to the model in a form it could use."

This sentence could almost serve as a plain definition of context engineering.

It’s not a mysterious prompt technique or piling on more background material. It’s about answering three questions: What steps would an expert take? What reliable tools does the system have? How should results be verified?

When this content is structured, what the model gains isn’t more knowledge, but a repeatable methodology.

This is also why "skills" are becoming key assets in the agent ecosystem. A good skill file might encapsulate years of a team’s tacit experience. It tells the model what to do first when encountering a certain type of problem, what mistakes are most common, which tools are trustworthy, and what outcome constitutes completion.

The differentiation of future companies likely won’t exist only in model weights, but also in this experience encoded into workflows.

VI. Why Agents Start to Go Astray Around Step 30

Almost every team that has seriously worked on agents has encountered the same scenario.

The first few steps go smoothly. The model can read requirements, call tools, write code. By step 30 or 50, it starts forgetting goals, misinterpreting states, repeating actions, or heading down a wrong path further and further.

Jeff Dean attributes one cause to out-of-distribution problems.

The model has seen many common tasks during training. As long as the task remains on the familiar "bright path," performance is usually fine. Once sequential operations take it to unfamiliar states, performance can suddenly drop. The further from the comfort zone, the more errors accumulate.

One solution is to provide skills and prompts to constrain the model as much as possible to familiar paths. Another method is to use multi-agent systems.

Multiple agents can try different plans, with another model acting as an evaluator judging which directions are more promising. Failed branches are discarded; successful ones continue. This is essentially performing search during inference.

It’s not unfamiliar to how human teams work. Faced with a complex problem, one person proposes a plan, another reviews risks, a third runs experiments. The team doesn’t bet everything on the first idea but reduces single-point failures through division of labor and feedback.

The longer an agent runs, the less the system design can rely on being correct the first time.

Truly reliable long-running agents need checkpoints, state management, rollback, branch exploration, external evaluation, permission control, and error recovery. It’s more like a distributed system than an extra-long chat window.

This is precisely where Jeff Dean’s background becomes relevant again.

One of the core problems MapReduce solved was how to have a large number of unreliable machines perform reliable computation. Today’s agent systems face a similar contradiction: a single model call isn’t perfect, tools can fail, but the overall task still needs to complete as reliably as possible.

Future excellent agent platforms might inherit many distributed systems ideas. Tasks can be split, results verified, failures retried, state recovered; local errors shouldn’t destroy the entire workflow.

When Jeff Dean says agents will run for days or even weeks, he’s not describing a longer chat. He’s describing a new computing infrastructure.

VII. How Two or Three People Can Beat Google: Find Problems Where Model Success Rate is Only 1%

In the context of Startup School, the most watched question is naturally entrepreneurial opportunities.

Google can co-design chips, data centers, models, and products. General-purpose models like Gemini are still rapidly expanding their capabilities. How can a two- or three-person team possibly win?

Jeff Dean’s answer isn’t romantic.

The opportunity for small teams usually exists in specific domains that general-purpose models haven’t fully focused on. Entrepreneurs can combine product interfaces, proprietary data, workflows, and domain skills to provide higher accuracy and better experience in a narrow scenario.

But he immediately gave a warning: general-purpose models are getting stronger quickly. What seems like an independent product feature today might be directly covered by foundation models in six or twelve months.

Therefore, entrepreneurs need to judge whether their advantage is durable.

Jeff Dean offered a very specific screening criterion: Look for tasks where the current general-purpose model success rate is close to 0% or 1%, not those it can already do 20% of the time.

"If the model completely fails, that might be a good sign. If it can already do part of it, just not very well, that might actually not be a good sign."

The reason is simple. 20% means the capability is already starting to emerge. More data, bigger models, and longer reasoning could quickly push it to usability. 0% or 1% suggests the task might lack key data, special tools, domain feedback, or require a capability general-purpose models can’t easily acquire in the short term.

This could be called Jeff Dean’s "1% Rule".

It’s not suggesting entrepreneurs pick the hardest problems, but look for problems where general-purpose models have a structural blind spot.

These blind spots fall into roughly three categories.

The first is proprietary data. General-purpose models can organize the world’s information but might not access a user’s full personal profile, a company’s internal processes, or real-time data from a specific device. Startup products that gain this data can form a perspective different from foundation models.

The second is professional evaluation. Many industries don’t lack generation capability but lack reliable judgment. Healthcare, materials, chips, manufacturing, and scientific research all need high-quality validators. Whoever defines "what is correct" can have agents continuously optimize.

The third is narrow and deep models. AlphaFold isn’t a general chat model; it builds highly specialized capability for protein structure. Similar opportunities might emerge in materials science, chip design, and other specialized fields.

This judgment isn’t easy for entrepreneurs. It requires teams to understand both the boundaries of model capabilities and the deep problems within an industry. Knowing only AI leads to building features quickly absorbed by platforms. Knowing only the industry might underestimate the speed of model progress.

The real opportunity lies at the intersection.

VIII. When Code is No Longer Scarce, Specifications, Taste, and Problem Selection Become More Valuable

Diana posed a hypothetical: If in the future every founder could manage 50 or 100 agents simultaneously, and all code was written by agents, what capability would become scarce?

Jeff Dean’s answer was "taste."

More precisely, the judgment of what agents should be tasked to do.

He believes most of the value in research work isn’t in executing experiments beautifully, but in whether one chooses a problem worth researching. A team can use the most exquisite methods to complete irrelevant research. Or they can seize a key problem that, once solved, changes the entire field.

As the cost of execution drops with agents, the importance of problem selection will rise further.

In the past, a vague idea might naturally die due to high development costs. In the future, with enough agents mobilized, many ideas can be rapidly prototyped. The world won’t automatically produce more good products; it will just produce more products.

Specifications will also become more important.

Jeff Dean said that when collaborating with virtual agents, the clearer the goal, the higher the success rate. In the past, vague requirements given to a senior engineer could be clarified through questioning, and shared context helped fill in intent. Agents, though they can also ask questions, are more prone to guessing on their own when context is missing.

A typical high-success-rate task is migrating software from one programming language to another. The reason isn’t that migration is simple, but that the specification is extremely complete. The old code defines behavior, tests define boundaries, and the agent can check item by item until the new version behaves consistently.

"Now agents can write software for you, but specifying what you actually want becomes more important."

This sentence has direct implications for so-called AI-native organizations.

Future managers won’t just assign tasks; they’ll need to write clearer goals and acceptance criteria. Design documents won’t just be team communication materials; they’ll also become input for machine execution. Tests, metrics, constraints, and examples will move from the end of the development process to the task definition stage.

As for how to train "taste," Jeff Dean’s method is pragmatic.

Write down a list of things you think will become important in the next 12 months. You don’t have to work on all of them. Check back in 12 months: which predictions came true, which were built by others, which made no progress. By accumulating prediction samples, people gradually calibrate their judgment.

Taste isn’t entirely innate. It can also be trained through reflection.

IX. A Good Thought Experiment First Removes the Industry's Most Solid Premises

In the latter part of the interview, Jeff Dean shared a rather wild thought experiment.

For the past 60 years, the chip industry has pursued smaller, more stable transistors with lower error rates. It’s assumed that chips from the same design should be as identical as possible, with bit flips as rare as possible.

But in large distributed systems, engineers long ago accepted that individual components fail. Hard drives die, machines crash, switches malfunction. System reliability doesn’t come from each component never failing, but from replication, checks, redundancy, and recovery.

So Jeff Dean asked: What if transistors had 20 errors per day, instead of one error every few million years?

This isn’t an actual product plan. He’s just trying to remove a taken-for-granted premise. Perhaps extremely unreliable transistors could be manufactured in a completely different way, with the system guaranteeing results through multiple paths and high-level redundancy.

Most thought experiments don’t become products. Many industry practices persist for decades for good reasons. But Jeff Dean believes we should still periodically re-examine those reasons.

MapReduce came from a similar process.

Early Google’s crawler and indexing systems contained lots of manual parallel code, checkpoints, and fault recovery logic. The actual business computations were often simple, like reading all web pages to determine language. But the simple intent was drowned in system code.

Jeff Dean and Sanjay Ghemawat drew inspiration from functional programming. They abstracted many tasks into Map and Reduce, pushing parallelization, scheduling, fault tolerance, and retries down into a unified framework. Business developers only needed to express the computation itself.

This design didn’t make machines infallible. It made errors absorbable by the system.

Today’s agent engineering might be at a similar stage. Many teams are still manually orchestrating prompts, retry logic, and tool calls for each task. In the future, could a concise abstraction like MapReduce emerge, making decomposition, validation, recovery, and parallel exploration for long-running agents a foundational capability?

This might be the opportunity for the next batch of infrastructure companies.

X. AI Starts Building Better AI, The Scientific Method Compressed into High-Speed Loops

Jeff Dean’s most exciting direction for the future is automating the scientific method itself.

The traditional research process is to propose a hypothesis, design an experiment, run it, analyze results, and generate the next hypothesis. The speed of this loop has long been constrained by experiment cost and verification latency.

AI can change two parts.

One part is automatically proposing and executing more experiments. The other is turning expensive validators into cheap approximate models.

Jeff Dean gave the example of quantum chemistry. To determine the properties of a molecular configuration, researchers can run density functional theory simulations. One simulation might take all night. Google researchers trained a neural network approximator using lots of simulation inputs and outputs. It approached the accuracy of the original simulator but was about 300,000 times faster.

When verification speed changes, the shape of scientific problems changes too.

Screening 10 million candidate solutions in the past might have been a project requiring months of compute. Now, while a researcher eats lunch, the system can do the initial screening. Experiments are no longer precious single bets but high-frequency searches.

This is also the common logic behind systems like AlphaEvolve and AlphaChip. Models propose solutions, tools execute them, evaluators filter results, and promising results go into the next round. As long as the loop is fast enough, the system can continuously explore a vast solution space.

Machine learning itself will become an object of this automated science.

Today, large research teams typically have humans propose new architectures or training methods, run small-scale experiments first, then scale promising ones. Jeff Dean believes there’s no fundamental barrier preventing models from taking over more and more of these steps. Humans give high-level direction; the system automatically explores structures, data recipes, training strategies, and combines successful experiments into new models.

A future metric for research efficiency might not just be FLOPS per second, but "how many effective discoveries per unit of compute."

Compute is important. How to turn compute into discovery is more important.

XI. The Distillation Paper Rejected by NeurIPS, and How to View Failure

In 2014, Jeff Dean, Geoff Hinton, and Oriol Vinyals submitted a paper on knowledge distillation. Today, knowledge distillation is a foundational method in model compression and capability transfer. Large models act as teachers, transferring their capabilities to smaller, faster, cheaper student models.

This later influential paper was rejected by NeurIPS that year.

One reviewer thought it was "unlikely to have a significant impact." Interested readers can visit "Rejected ≠ Failure! These High-Impact Papers Were Also Rejected by Top Conferences."

Jeff Dean spoke about this experience without anger. He said the reviewer might not have understood the real-world problems facing large-scale AI services. For Google, transforming expensive large models into small models that could serve hundreds of millions of users was clearly very important. For reviewers focused only on theoretical novelty, it might not have seemed sufficiently "fundamental."

After the paper was rejected, the team posted it on arXiv. The field read it anyway and started using it.

Today, distillation is an important method enabling Gemini’s Flash models to maintain strong capabilities at smaller sizes and lower latencies.

This story isn’t just inspirational material about "perseverance leads to success." It shows that evaluation systems always have blind spots. The value of a solution is sometimes immediately apparent only to those who have truly felt that system bottleneck.

This is also important for entrepreneurs.

Rejection from the market, investors, or peers might mean the direction is wrong, or it might just mean they aren’t in the same problem space. The difference lies in whether the team has specific enough evidence about why the problem is important and why it can be solved now.

Jeff Dean didn’t encourage blind persistence. He encouraged: understand the problem, keep validating, and don’t treat one review as the world’s final judgment.

XII. What Would the Young Jeff Dean Do Today

As the interview neared its end, Diana asked an imaginative question.

If the young Jeff Dean from 1999, when he joined Google, were transported to 2026, would he join a cutting-edge lab or start a company with two or three friends?

Jeff Dean didn’t give a standard answer.

Large organizations have structure, platforms, and many excellent colleagues. One can access knowledge they don’t understand and leverage mature products to impact global users. Small teams are freer but carry greater risk. Founders must truly believe in a problem and be willing to bear uncertainty for years.

The criteria he offered were more fundamental than "join big tech or start up."

"If I solve this problem, and the best possible outcome actually happens, will the world be noticeably better for it? Or will people just say, 'Huh, cool,' and that’s it?"

If the answer is just "cool," it might not be worth investing the most precious time.

He also emphasized the importance of companions. Find people with complementary skills, but also those with low ego, willing to collaborate, and enjoyable to be around. Truly hard problems often require long-term collaboration. Team members should ideally each have tools others lack and continue expanding their own "tool belts" through shared work.

This talk had a kind of old-school engineer's simplicity.

The AI industry likes to talk about exponential growth, superintelligence, and massive funding. Jeff Dean still brought the choice back to three small things: Work on a problem you truly care about, with people you enjoy working with, and try to make the world a bit better.

Conclusion: The Scarcest Thing in the AI Era is Still Seeing the Problem Clearly

Throughout Jeff Dean’s career, there are many oft-told legends.

He and Sanjay Ghemawat rewrote the search system in days, moving the index into memory. An estimate about three minutes of speech pushed Google to build the TPU. MapReduce hid massive parallelism and fault tolerance in a unified abstraction. Knowledge distillation went from a rejected paper to a foundational industry technique.

These stories easily paint him as a genius constantly receiving inspiration.

But from this interview, his method is actually highly consistent.

First calculate the order of magnitude. Find the real bottleneck. Then question default assumptions and build a simpler abstraction. Finally, use measurement and feedback to drive system iteration.

Today’s AI industry is undergoing a similar transition.

Models are already strong enough to handle junior engineer-level tasks. Next, what determines practical productivity isn’t just model IQ, but inference cost, context organization, tool quality, verification speed, and long-run reliability.

Agents will become more like team members. But they need clear specifications, skills, checkpoints, evaluators, and a system that can accommodate failure.

Startup opportunities won’t disappear, but they’ll become more demanding. Best not to work on things general-purpose models can already do 20% of the time, but to look for problems where success rates are still close to 0% or 1%. There might lie proprietary data, professional evaluators, narrow-domain models, or entirely new system abstractions.

When code generation becomes cheap, what becomes truly expensive is the problem itself.

What is worth doing? Which constraints are obsolete? What change just crossed a threshold? What system, if made 50 times faster, would become a completely different product?

Jeff Dean didn’t give the 6000 entrepreneurs a list of opportunities. He offered a more durable way of thinking.

Don’t rush after the hottest answers.

Calculate the problem first.

References

https://x.com/ycombinator/status/2082938685071491219

https://www.ycrootaccess.com/p/jeff-dean-the-1-rule-for-building

This article is from the WeChat public account "Almost Human" (ID:almosthuman2014), author: Panda

Trending Cryptos

Related Questions

QAccording to Jeff Dean, what is the key shift in AI competition from the past to the next stage?

AAccording to Jeff Dean, the key shift is from 'who has the bigger model' to 'who can better organize intelligence.' AI's next stage is not just about training smarter models, but about placing models into systems that can work long-term, continuously experiment, self-validate, and accumulate capabilities.

QWhat does Jeff Dean refer to as the '1% rule' for startup opportunities in the AI era?

AJeff Dean's '1% rule' advises startups to focus on tasks where the current general model's success rate is close to 0% or 1%, not tasks where it already achieves around 20%. A 0-1% rate suggests the task has a structural blind spot for general models, possibly requiring proprietary data, specialized tools, or domain-specific feedback, offering a more durable advantage.

QWhat major insight led to the development of Google's TPU?

AThe development of Google's TPU was driven by a calculation Jeff Dean made when deep learning for speech recognition significantly improved. He estimated that if every Google user used just three minutes of voice recognition daily, serving this demand with existing CPUs would require doubling Google's server footprint. This impending cost and scale crisis prompted the creation of specialized hardware for core ML computations.

QWhat are the three key elements Jeff Dean highlights as becoming more critical than the model itself in a useful AI system?

AJeff Dean highlights that beyond the model, a truly useful AI system critically needs organized context. This includes retrieval mechanisms, tools, memory, historical information, execution environments, and feedback loops. Structuring this context—defining steps, trusted tools, and verification methods—enables models to perform reliably in specific workflows.

QWhat fundamental engineering principle does Jeff Dean consistently apply, as illustrated by examples like improving Google Search and developing the TPU?

AJeff Dean consistently applies the principle of first-principles thinking and recalculating the order of magnitude. He questions default assumptions, identifies the true bottleneck (like data movement energy costs vs. computation), and asks if constraints have fundamentally changed (e.g., memory capacity, model capability). This leads to re-architecting systems for simplicity and efficiency based on current realities.

Related Reads

From South Korea to the United States: Blue-Collar Jobs Are Becoming Increasingly Popular, Thanks to AI

AI is reshaping the labor market's value proposition. The traditional four-year college degree is losing its appeal as a guaranteed career path, while skilled blue-collar trades like electricians, welders, and plumbers are experiencing historic demand and wage premiums. This shift is driven by dual pressures: AI's displacement of certain white-collar roles and a booming need for physical infrastructure and data center construction. Data confirms the trend. In the U.S., vocational school revenue surged, and a significant portion of recent layoffs are AI-related. Surveys show a majority of Gen Z adults plan to pursue blue-collar work, citing better job security against AI automation. Vocational education interest has exploded recently. Experts cite a psychological shift as younger generations seek tangible, AI-resistant careers and avoid high student debt. In many cases, salaries for skilled trades now match or exceed those requiring a bachelor's degree. In South Korea, semiconductor vocational high schools boast near-total employment, with graduates securing high-paying roles at companies like Samsung. The shortage is structural, exacerbated by a retiring baby boomer workforce and massive infrastructure projects. Companies like JPMorgan Chase, Meta, and Lowe's are investing heavily in training programs. However, overcoming historical stigma and a "perception gap" around trade careers remains a key challenge to closing the talent gap.

marsbit58m ago

From South Korea to the United States: Blue-Collar Jobs Are Becoming Increasingly Popular, Thanks to AI

marsbit58m ago

Qualcomm: AI Hype Subsides, When Will Smartphones Emerge from the Gloom?

Qualcomm reported its Q3 FY2026 results (ending June 2026), with revenue of $9.95B, down 4% YoY but above expectations. Gross margin declined to 53.1%, pressured by rising costs across manufacturing and memory. Key business segments showed mixed performance: Handset revenue fell 19.6% YoY to $5.09B, dragged by an 11% decline in non-Apple Android shipments and weaker high-end mix. Conversely, Automotive revenue surged 61% to $1.59B, and IoT grew 9% to $1.83B. Core operating profit dropped 41% YoY due to margin compression and higher expenses. Management's Q4 FY2026 guidance projects revenue of $9.7B-$10.5B, in line with consensus, but Non-GAAP EPS guidance of $2.05-$2.25 fell short of expectations. Amidst persistent weakness in its core handset market, Qualcomm is pursuing growth in AI, focusing on Edge AI (smartphones, PCs, automotive) and Data Center AI. Its data center strategy includes four pillars: AI accelerators (e.g., AI200), commercial CPUs (Dragonfly C1000), custom silicon, and connectivity solutions. While these initiatives initially boosted its stock, concerns over AI capital expenditure sustainability have since erased those gains. The company targets $5B in data center revenue for FY2027 and $15B for FY2029. The report concludes that with the traditional handset business still under pressure, the data center opportunity is currently viewed as a longer-term option, and a more conservative valuation based on core operations may be warranted until AI contributions materialize.

marsbit1h ago

Qualcomm: AI Hype Subsides, When Will Smartphones Emerge from the Gloom?

marsbit1h ago

Trading

Spot

Hot Articles

Discussions

Welcome to the HTX Community. Here, you can stay informed about the latest platform developments and gain access to professional market insights. Users' opinions on the price of AI (AI) are presented below.

活动图片