He Personally Built o1 and o3, but Suddenly Announced: Humans Can Retire Forever

marsbitPublished on 2026-08-27Last updated on 2026-08-27

Abstract

AI researcher Jerry Tworek, former lead of OpenAI's reasoning models (o1/o3), predicts human AI researchers will be largely obsolete within two years. He argues that while AI agents currently excel at execution (reducing experiment cycles from a month to a day), their creative idea generation remains low-quality. True frontier AI development, he claims, is driven by only 30-50 people globally, with everyone else in supporting roles. A core bottleneck is the Transformer architecture, which he believes is suboptimal but has seen only 3-4 serious replacement attempts at OpenAI in seven years due to massive computational and validation hurdles. He highlights a key paradox: you need significant compute to prove a new direction, but you need proof to get the compute. A breakthrough came when he received dedicated GPU resources, leading to scaling reinforcement learning and the o1 model. Now, through his company Core Automation, he uses AI agents to automate costly R&D tasks. In one case, a $100k AI effort optimized a GPU kernel by 60x, matching rare expert-level performance. This acceleration, he argues, will soon automate even the remaining creative and architectural work of the few top researchers. Looking ahead, Tworek envisions a post-work future where humans pursue excellence for its own sake—like ancient philosophers or dedicated athletes—free from the necessity of labor for societal function. He acknowledges the personal challenge of deriving self-worth from work but is a...

Hey, did you hear?

The good days for human researchers are, at best, only two years left.

Two years from now, humans doing AI research will be like humans playing chess today.

Of course you can still play, but no one will care how well you play.

This is the latest bombshell dropped by former OpenAI inference model lead Jerry Tworek.

He joined OpenAI in 2019 and stayed for seven years. Back when reinforcement learning refused to scale, he was the one who stubbornly pushed it forward. He personally led the teams that hammered out the o1 and o3 generations of major breakthroughs.

"We want to build AGI." This was the entire roadmap Ilya gave at an all-hands meeting when he first joined OpenAI in 2019. Seven years later, he went out on his own.

What's even more chilling is that he's not the only one saying this. The most popular inside joke among AI researchers right now goes like this—

We only have a few days of work left, so let's get it done while we still can, then we can all retire and rest.

Everyone treats it as dark humor, passing it around. But everyone who tells the joke knows perfectly well in their hearts.

This joke is probably reality.

And that's just the appetizer. Throughout the interview, he delivered one hard-hitting statement after another:

  • Agents do have creativity, but it's an extremely low-quality, massively produced creativity. Ideas are diverse, but they are also generally terrible.
  • In the whole world, there might only be thirty to fifty people who truly understand, end-to-end, how to train and deploy a frontier model. Everyone else is supporting them.
  • The assumption that Transformer is the optimal solution is almost certainly false.
  • In the seven years at OpenAI, there were only three or four serious attempts to replace the underlying architecture.
  • The premise that the world needs us to go to work in order to function simply doesn't hold.

Below is the edited version. Enjoy.

Half the Work is Already Done Without Humans

The two-year figure isn't about benchmark scores or computing power.

He's counting how much of the *research* work is still being done by humans.

And that work was already split into two parts long ago.

One part is generating ideas, figuring out which direction to go. The other part is the labor, turning ideas into code that runs and fetching the data.

And the labor part has basically been handed over to Agents.

Tworek started a company called Core Automation in April this year, with the slogan of building the world's most automated AI lab.

There, the full cycle of an experiment has been compressed from a month down to a single day. Efficiency improved thirtyfold!

However, the idea-generation part is not something Agents can handle yet. The reason is—

They are absolutely a high-creativity species.

But they squander creativity in an extremely low-quality, massively produced way. The ideas are indeed diverse, but the ideas are also generally terrible.

Only Thirty to Fifty People in the World

Listening to this, the idea-generation part seems secure.

But there aren't many people who can actually generate ideas to begin with.

During the interview, the host revealed that an OpenAI internal researcher had told him in confidence—

In the whole world, there might only be thirty to fifty people who truly understand, end-to-end, how to train and deploy a frontier model.

Everyone else is supporting these few dozen minds, including the vast majority of full-time employees at that top company.

Tworek didn't refute this.

He also believes that's how any top-tier team operates. A handful of people set the course, followed by an entire roaring execution machine.

To understand just how sought-after these few dozen people are, look at his own hiring list.

Core Automation's co-founder Rohan Anil is from Anthropic, and before that, Google DeepMind.

Anmol Gulati, who worked on Gemini at DeepMind, was also recruited. Even Julia Villagra, OpenAI's former head of people, followed him over.

All the labs are fishing from the same pool. And this pool only has a few dozen fish.

So the two-year figure doesn't relate to all of humanity.

It's about how long these few dozen people can hold on.

Seven Years, the Architecture Was Only Seriously Challenged Three or Four Times

Since there are only a few dozen people left, what's the final wall standing in front of AI?

Tworek's answer is one word: Transformer.

In his view, it almost certainly isn't the optimal solution.

First, models can't continue learning after deployment. You can chat with them all you want, they won't get stronger; the next conversation will be with the same old model. The context window also can't hold up. He said after using Codex for about twenty minutes, he'd have to compress once.

Second, if you try to patch it with continuous fine-tuning, you'll find it's not only extremely inefficient but also leads to catastrophic forgetting—learning new things makes it forget the old.

So most of the work the entire industry has done around Transformer in recent years has essentially been about making it cheaper, not actually making it stronger.

What Tworek really wants isn't actually in the architecture itself.

He wants models that can continue learning at test time, models that can keep growing from user interaction, from user data. Changing the architecture is just a means to that end.

Everyone in the industry understands these principles. But after all these years, why is Transformer still standing firm?

The reason behind it is absurdly simple—they barely even tried.

In his seven years at OpenAI, there were only three or four serious attempts to replace the underlying architecture.

The process was like this.

A researcher first writes a small-scale validation experiment, which takes at least three months to run after writing. If the results look promising, only then dare they scale it up.

And by scaling up, you have to convince about ten people with decision-making power using your silver tongue. Then these ten people pour three to six months into your bottomless pit.

In the end, it's either crushed by the momentum Transformer has already snowballed, or partially absorbed, becoming a screw on its body.

Seven years, three or four times. This is the total output in architectural innovation from the world's strongest AI lab.

"Jare, Take This Compute and Burn It"

He himself ran into exactly the same dead-end at OpenAI. What pried it open was one sentence.

At that time, that batch of half-dead experiments had struggled to show a faint "sign of life." Not great, but at least there was a sprout.

Right at this critical juncture, Jakub Pachocki, who later became OpenAI's Chief Scientist, found him.

"Jare, all these GPUs are for you. See if you can push the results in your hands a bit bigger, a bit harder."

"Now you have these GPUs in your hands." When he recounted this sentence, o1 didn't even have a name yet.

This sentence pulled him out of a paradox—

You must first deliver results to be qualified for compute.

But you clearly need that compute first to hammer out the results.

Tworek says most frontier directions are trapped in this paradox. The secret wars fought tooth and nail for compute in labs are ultimately about finding a way out.

And the exit sometimes is just a tiny bit of confidence from leadership.

Just having someone from above say, I want you to have a decent compute quota to go hard at it. That's enough.

Once the gate opened, there was no holding back.

Reinforcement learning crossed several orders of magnitude from this point on.

Then the emergence of o1 broke through; the path of reasoning models, which countless people had declared dead, was stubbornly brought to life by him.

$100,000, Replacing an Expert

And this time, he doesn't have to wait for anyone to say anything.

The most expensive link in the whole chain is translating abstract ideas into code that can actually run. And that's precisely what today's Agents excel at.

For example, Core Automation used this approach to tackle GPU kernels.

Specifically, they threw a QR decomposition kernel at a programming Agent for four weeks, burning about $100,000 in API call costs, and finally pushed the speed of this kernel to sixty times its original.

This work belongs to low-level performance engineering, tweaking how a piece of matrix computation runs faster on graphics cards, usually requiring a handful of experts to manually tune line by line.

And there aren't many such experts in the world.

Now $100,000 can replace one. It doesn't need to be convinced, nor does it need three to six months.

The barrier that was stuck for seven years was just washed away like that.

So What Can We Do Then?

Even the hardest-to-crack architecture is loosening up, so the steering wheel in the hands of those few dozen people won't be held for much longer.

At the end of the interview, the host made it clear and asked him: when this day really comes, what is left for humans to do?

For this, Tworek described two scenes.

The first is Ancient Greece.

People meet in the square, leisurely chat philosophy all day, then go exercise, eat olives, and drink wine.

He laughed himself right after describing it, admitting that this is probably just projecting his own wishes.

"We meet in the square and then chat." His exact words describing post-AGI era human daily life. He laughed first after saying it.

The second scene is high school, or rather university.

He thinks humans should retain a pursuit of "excellence" itself. Keep greedily learning new things, training both body and mind to the limit.

It's a bit like professional sports; there's no economic reason forcing you to bleed and sweat, but the pride in human bones is to reach for that thing called "greatness."

We need to find various ways to do this.

Because in the next world, there won't be things like "if you don't do it, the world will collapse." The world will run on its own on the infrastructure we've already built.

"We should keep learning forever." He described this as a kind of obligation, not a pastime.

Then he calmly revealed his cards.

The assumption that we must work for the world to function is fundamentally unnecessary.

"Many people's self-worth comes from work." He admits this is the hardest hurdle to cross, including for himself.

He Himself Is Also on This List

To be honest, what sends the biggest chill down the spine after listening to the whole thing isn't that two-year figure.

It's that he himself knows he is stepping on the accelerator for this countdown.

The thing he wants to build is defined as something that can learn and strengthen itself, no longer needing people to feed it while standing by. If built, those few dozen positions will disappear even faster.

And he has an even harsher statement; the first one he sentences is himself.

If you're not the lab with the most terrifying compute reserves and the largest scale, you will die a very ugly death.

When he said this, Core Automation was only four months old, with zero revenue on the books.

He said since starting the company, almost every week someone comes to tell him, Jare, it's too late, that ladder to heaven was already pulled up long ago.

His response is only one of his company's mantras.

People at our company love to say, everything is a skill issue.

Translated, it means, when you fail in the end, don't blame the environment; blame your own lack of skill.

You know what? In the context of this interview, it really has that flavor.

Looking back at that popular joke from the beginning.

We only have a few days of work left, so let's get it done while we still can.

Is the tone in that excitement, or desolation? We cannot know.

Because for the people saying it, those two emotions are one and the same.

Tworek says the pace of this field is extremely draining. After so many years of doing it, he is really, really tired.

"But if we truly believe this is the most important period in our entire careers."

"Then it is probably, incredibly worth it."

References:

https://x.com/MTSlive/status/2092387349623935322

This article is from the WeChat public account "New Zhiyuan", author: ASI Revelation, editor: Moshe

Related Questions

QWhat is the core prediction made by Jerry Tworek regarding the future of human AI researchers?

AJerry Tworek predicts that in about two years, human AI researchers will become largely irrelevant, similar to how human chess players are viewed today, as AI agents will take over most research tasks.

QAccording to Tworek, what is a major limitation of current large models like those based on the Transformer architecture?

AA major limitation is their inability to learn continuously after deployment (at test time), and their tendency to suffer from catastrophic forgetting during fine-tuning.

QWhat is the purpose of Tworek's company, Core Automation, and what key achievement does he mention?

ACore Automation aims to build the world's most automated AI lab. He mentions it reduced a full experimental cycle from one month to one day, achieving a 30x efficiency gain, and used a programming agent to accelerate a GPU kernel by 60x.

QHow many people does the article suggest are capable of end-to-end development and deployment of cutting-edge AI models?

AThe article suggests only about thirty to fifty people worldwide truly understand how to train and deploy a cutting-edge AI model from end to end.

QWhat are the two visions Tworek describes for human life in a post-AGI world where work is no longer necessary?

AHe describes two visions: first, a life reminiscent of ancient Greece, focused on philosophy, exercise, and leisure; second, a life akin to high school or university, where people pursue excellence and learning for its own sake, driven by personal pride and ambition.

Related Reads

NVIDIA Earnings Report Quick Read: Quarterly Revenue on the Verge of Breaking the $100 Billion Mark, Can Still Grow 70% Next Year

NVIDIA Q2 FY2027 Earnings Report: Record Revenue Nears $100 Billion, Projects 70% Growth for Next Year NVIDIA reported exceptional Q2 FY2027 results, with revenue reaching $96.22 billion (up 106% YoY) and non-GAAP net income of $53.95 billion (up 118% YoY). Q3 revenue guidance of $108 billion signals the company's imminent entry into a "quarterly $100 billion revenue" era. The most striking projection came from the CFO, who forecasted approximately 70% revenue growth for FY2028, significantly above Wall Street's 45% expectation, and notably stated this outlook excludes any data center revenue from China. The data center segment remained the core driver, generating $89 billion (up 117% YoY), representing over 90% of total revenue. Growth continues to be primarily fueled by capital expenditures from hyperscale cloud providers like Amazon, Microsoft, Google, and Meta. A key product milestone was the confirmation that the Vera Rubin platform has entered full-scale production and begun shipments. Management stated that every 1 GW of Vera Rubin compute deployed represents a roughly $40 billion revenue opportunity, with the platform expected to contribute about 20% of data center revenue in Q3. CEO Jensen Huang noted its ramp-up is the fastest in company history, providing confidence for the strong future outlook. Management emphasized that AI demand is not slowing but accelerating and broadening into new areas like inference, enterprise AI, and robotics. The primary constraint on growth is now supply, not demand. Huang indicated that without supply limitations, the FY2028 outlook would be "much higher." Bottlenecks include components like HBM, advanced packaging, and data center power. In response, NVIDIA is expanding its role, collaborating with financial firms to mobilize over $500 billion in third-party capital for AI infrastructure and securing key resources like land and power. Crossing the $100 billion quarterly revenue mark represents a new scale for NVIDIA, while the 70% growth projection for next year suggests it may just be the starting point for the next phase.

Odaily星球日报2m ago

NVIDIA Earnings Report Quick Read: Quarterly Revenue on the Verge of Breaking the $100 Billion Mark, Can Still Grow 70% Next Year

Odaily星球日报2m ago

Inflation Has Not Improved. Will Warsh Support a Rate Hike on Friday?

U.S. inflation remained stubbornly high in July, with the PCE price index holding at a year-on-year increase of 3.7%, unchanged from June and still far above the Federal Reserve's 2% target. The core PCE index also stayed flat at 3.3%. While inflation did not worsen, the fact that it did not improve either has increased market expectations for further interest rate hikes. Futures pricing now indicates a higher probability of a rate increase in September and fully prices in one hike by year-end. The economic backdrop is mixed. Second-quarter GDP growth was revised to 1.5%, but underlying components like consumer spending and business investment were robust. However, inflation-adjusted consumer spending stalled in July, and real incomes have barely grown over the past year, eroding purchasing power. The data provides arguments for both sides of the policy debate. The "wait-and-see" camp points to the lack of acceleration in inflation and upcoming methodological changes that may lower reported figures. The "pro-hike" camp highlights that inflation remains hotter than forecasts, sticky services prices, rising diesel and chip costs, and renewed trade tensions with Canada. All eyes are now on Fed Chair Kevin Warsh's upcoming speech at Jackson Hole for clarity on his policy stance. With inflation persistently above target for over five years and midterm elections approaching where prices are a key issue, the pressure for decisive action is mounting. The speech carries significant two-way risk for markets.

marsbit47m ago

Inflation Has Not Improved. Will Warsh Support a Rate Hike on Friday?

marsbit47m ago

Trading

Spot
活动图片