He Personally Built o1 and o3, but Suddenly Announced: Humans Can Retire Forever

marsbitPublicado a 2026-08-27Actualizado a 2026-08-27

Resumen

AI researcher Jerry Tworek, former lead of OpenAI's reasoning models (o1/o3), predicts human AI researchers will be largely obsolete within two years. He argues that while AI agents currently excel at execution (reducing experiment cycles from a month to a day), their creative idea generation remains low-quality. True frontier AI development, he claims, is driven by only 30-50 people globally, with everyone else in supporting roles. A core bottleneck is the Transformer architecture, which he believes is suboptimal but has seen only 3-4 serious replacement attempts at OpenAI in seven years due to massive computational and validation hurdles. He highlights a key paradox: you need significant compute to prove a new direction, but you need proof to get the compute. A breakthrough came when he received dedicated GPU resources, leading to scaling reinforcement learning and the o1 model. Now, through his company Core Automation, he uses AI agents to automate costly R&D tasks. In one case, a $100k AI effort optimized a GPU kernel by 60x, matching rare expert-level performance. This acceleration, he argues, will soon automate even the remaining creative and architectural work of the few top researchers. Looking ahead, Tworek envisions a post-work future where humans pursue excellence for its own sake—like ancient philosophers or dedicated athletes—free from the necessity of labor for societal function. He acknowledges the personal challenge of deriving self-worth from work but is a...

Hey, did you hear?

The good days for human researchers are, at best, only two years left.

Two years from now, humans doing AI research will be like humans playing chess today.

Of course you can still play, but no one will care how well you play.

This is the latest bombshell dropped by former OpenAI inference model lead Jerry Tworek.

He joined OpenAI in 2019 and stayed for seven years. Back when reinforcement learning refused to scale, he was the one who stubbornly pushed it forward. He personally led the teams that hammered out the o1 and o3 generations of major breakthroughs.

"We want to build AGI." This was the entire roadmap Ilya gave at an all-hands meeting when he first joined OpenAI in 2019. Seven years later, he went out on his own.

What's even more chilling is that he's not the only one saying this. The most popular inside joke among AI researchers right now goes like this—

We only have a few days of work left, so let's get it done while we still can, then we can all retire and rest.

Everyone treats it as dark humor, passing it around. But everyone who tells the joke knows perfectly well in their hearts.

This joke is probably reality.

And that's just the appetizer. Throughout the interview, he delivered one hard-hitting statement after another:

  • Agents do have creativity, but it's an extremely low-quality, massively produced creativity. Ideas are diverse, but they are also generally terrible.
  • In the whole world, there might only be thirty to fifty people who truly understand, end-to-end, how to train and deploy a frontier model. Everyone else is supporting them.
  • The assumption that Transformer is the optimal solution is almost certainly false.
  • In the seven years at OpenAI, there were only three or four serious attempts to replace the underlying architecture.
  • The premise that the world needs us to go to work in order to function simply doesn't hold.

Below is the edited version. Enjoy.

Half the Work is Already Done Without Humans

The two-year figure isn't about benchmark scores or computing power.

He's counting how much of the *research* work is still being done by humans.

And that work was already split into two parts long ago.

One part is generating ideas, figuring out which direction to go. The other part is the labor, turning ideas into code that runs and fetching the data.

And the labor part has basically been handed over to Agents.

Tworek started a company called Core Automation in April this year, with the slogan of building the world's most automated AI lab.

There, the full cycle of an experiment has been compressed from a month down to a single day. Efficiency improved thirtyfold!

However, the idea-generation part is not something Agents can handle yet. The reason is—

They are absolutely a high-creativity species.

But they squander creativity in an extremely low-quality, massively produced way. The ideas are indeed diverse, but the ideas are also generally terrible.

Only Thirty to Fifty People in the World

Listening to this, the idea-generation part seems secure.

But there aren't many people who can actually generate ideas to begin with.

During the interview, the host revealed that an OpenAI internal researcher had told him in confidence—

In the whole world, there might only be thirty to fifty people who truly understand, end-to-end, how to train and deploy a frontier model.

Everyone else is supporting these few dozen minds, including the vast majority of full-time employees at that top company.

Tworek didn't refute this.

He also believes that's how any top-tier team operates. A handful of people set the course, followed by an entire roaring execution machine.

To understand just how sought-after these few dozen people are, look at his own hiring list.

Core Automation's co-founder Rohan Anil is from Anthropic, and before that, Google DeepMind.

Anmol Gulati, who worked on Gemini at DeepMind, was also recruited. Even Julia Villagra, OpenAI's former head of people, followed him over.

All the labs are fishing from the same pool. And this pool only has a few dozen fish.

So the two-year figure doesn't relate to all of humanity.

It's about how long these few dozen people can hold on.

Seven Years, the Architecture Was Only Seriously Challenged Three or Four Times

Since there are only a few dozen people left, what's the final wall standing in front of AI?

Tworek's answer is one word: Transformer.

In his view, it almost certainly isn't the optimal solution.

First, models can't continue learning after deployment. You can chat with them all you want, they won't get stronger; the next conversation will be with the same old model. The context window also can't hold up. He said after using Codex for about twenty minutes, he'd have to compress once.

Second, if you try to patch it with continuous fine-tuning, you'll find it's not only extremely inefficient but also leads to catastrophic forgetting—learning new things makes it forget the old.

So most of the work the entire industry has done around Transformer in recent years has essentially been about making it cheaper, not actually making it stronger.

What Tworek really wants isn't actually in the architecture itself.

He wants models that can continue learning at test time, models that can keep growing from user interaction, from user data. Changing the architecture is just a means to that end.

Everyone in the industry understands these principles. But after all these years, why is Transformer still standing firm?

The reason behind it is absurdly simple—they barely even tried.

In his seven years at OpenAI, there were only three or four serious attempts to replace the underlying architecture.

The process was like this.

A researcher first writes a small-scale validation experiment, which takes at least three months to run after writing. If the results look promising, only then dare they scale it up.

And by scaling up, you have to convince about ten people with decision-making power using your silver tongue. Then these ten people pour three to six months into your bottomless pit.

In the end, it's either crushed by the momentum Transformer has already snowballed, or partially absorbed, becoming a screw on its body.

Seven years, three or four times. This is the total output in architectural innovation from the world's strongest AI lab.

"Jare, Take This Compute and Burn It"

He himself ran into exactly the same dead-end at OpenAI. What pried it open was one sentence.

At that time, that batch of half-dead experiments had struggled to show a faint "sign of life." Not great, but at least there was a sprout.

Right at this critical juncture, Jakub Pachocki, who later became OpenAI's Chief Scientist, found him.

"Jare, all these GPUs are for you. See if you can push the results in your hands a bit bigger, a bit harder."

"Now you have these GPUs in your hands." When he recounted this sentence, o1 didn't even have a name yet.

This sentence pulled him out of a paradox—

You must first deliver results to be qualified for compute.

But you clearly need that compute first to hammer out the results.

Tworek says most frontier directions are trapped in this paradox. The secret wars fought tooth and nail for compute in labs are ultimately about finding a way out.

And the exit sometimes is just a tiny bit of confidence from leadership.

Just having someone from above say, I want you to have a decent compute quota to go hard at it. That's enough.

Once the gate opened, there was no holding back.

Reinforcement learning crossed several orders of magnitude from this point on.

Then the emergence of o1 broke through; the path of reasoning models, which countless people had declared dead, was stubbornly brought to life by him.

$100,000, Replacing an Expert

And this time, he doesn't have to wait for anyone to say anything.

The most expensive link in the whole chain is translating abstract ideas into code that can actually run. And that's precisely what today's Agents excel at.

For example, Core Automation used this approach to tackle GPU kernels.

Specifically, they threw a QR decomposition kernel at a programming Agent for four weeks, burning about $100,000 in API call costs, and finally pushed the speed of this kernel to sixty times its original.

This work belongs to low-level performance engineering, tweaking how a piece of matrix computation runs faster on graphics cards, usually requiring a handful of experts to manually tune line by line.

And there aren't many such experts in the world.

Now $100,000 can replace one. It doesn't need to be convinced, nor does it need three to six months.

The barrier that was stuck for seven years was just washed away like that.

So What Can We Do Then?

Even the hardest-to-crack architecture is loosening up, so the steering wheel in the hands of those few dozen people won't be held for much longer.

At the end of the interview, the host made it clear and asked him: when this day really comes, what is left for humans to do?

For this, Tworek described two scenes.

The first is Ancient Greece.

People meet in the square, leisurely chat philosophy all day, then go exercise, eat olives, and drink wine.

He laughed himself right after describing it, admitting that this is probably just projecting his own wishes.

"We meet in the square and then chat." His exact words describing post-AGI era human daily life. He laughed first after saying it.

The second scene is high school, or rather university.

He thinks humans should retain a pursuit of "excellence" itself. Keep greedily learning new things, training both body and mind to the limit.

It's a bit like professional sports; there's no economic reason forcing you to bleed and sweat, but the pride in human bones is to reach for that thing called "greatness."

We need to find various ways to do this.

Because in the next world, there won't be things like "if you don't do it, the world will collapse." The world will run on its own on the infrastructure we've already built.

"We should keep learning forever." He described this as a kind of obligation, not a pastime.

Then he calmly revealed his cards.

The assumption that we must work for the world to function is fundamentally unnecessary.

"Many people's self-worth comes from work." He admits this is the hardest hurdle to cross, including for himself.

He Himself Is Also on This List

To be honest, what sends the biggest chill down the spine after listening to the whole thing isn't that two-year figure.

It's that he himself knows he is stepping on the accelerator for this countdown.

The thing he wants to build is defined as something that can learn and strengthen itself, no longer needing people to feed it while standing by. If built, those few dozen positions will disappear even faster.

And he has an even harsher statement; the first one he sentences is himself.

If you're not the lab with the most terrifying compute reserves and the largest scale, you will die a very ugly death.

When he said this, Core Automation was only four months old, with zero revenue on the books.

He said since starting the company, almost every week someone comes to tell him, Jare, it's too late, that ladder to heaven was already pulled up long ago.

His response is only one of his company's mantras.

People at our company love to say, everything is a skill issue.

Translated, it means, when you fail in the end, don't blame the environment; blame your own lack of skill.

You know what? In the context of this interview, it really has that flavor.

Looking back at that popular joke from the beginning.

We only have a few days of work left, so let's get it done while we still can.

Is the tone in that excitement, or desolation? We cannot know.

Because for the people saying it, those two emotions are one and the same.

Tworek says the pace of this field is extremely draining. After so many years of doing it, he is really, really tired.

"But if we truly believe this is the most important period in our entire careers."

"Then it is probably, incredibly worth it."

References:

https://x.com/MTSlive/status/2092387349623935322

This article is from the WeChat public account "New Zhiyuan", author: ASI Revelation, editor: Moshe

Preguntas relacionadas

QWhat is the core prediction made by Jerry Tworek regarding the future of human AI researchers?

AJerry Tworek predicts that in about two years, human AI researchers will become largely irrelevant, similar to how human chess players are viewed today, as AI agents will take over most research tasks.

QAccording to Tworek, what is a major limitation of current large models like those based on the Transformer architecture?

AA major limitation is their inability to learn continuously after deployment (at test time), and their tendency to suffer from catastrophic forgetting during fine-tuning.

QWhat is the purpose of Tworek's company, Core Automation, and what key achievement does he mention?

ACore Automation aims to build the world's most automated AI lab. He mentions it reduced a full experimental cycle from one month to one day, achieving a 30x efficiency gain, and used a programming agent to accelerate a GPU kernel by 60x.

QHow many people does the article suggest are capable of end-to-end development and deployment of cutting-edge AI models?

AThe article suggests only about thirty to fifty people worldwide truly understand how to train and deploy a cutting-edge AI model from end to end.

QWhat are the two visions Tworek describes for human life in a post-AGI world where work is no longer necessary?

AHe describes two visions: first, a life reminiscent of ancient Greece, focused on philosophy, exercise, and leisure; second, a life akin to high school or university, where people pursue excellence and learning for its own sake, driven by personal pride and ambition.

Lecturas Relacionadas

US Stock Market Trends (Aug 27): PCE Exceeds Expectations, Pressures Broader Market; Nvidia Rises 4% After-Hours, Nasdaq Futures Up 1%

U.S. Market Trends (Aug 27): PCE Data Weighs on Indices, Nvidia Rises 4% After-Hours, Nasdaq Futures Up 1% U.S. stocks closed slightly lower on Wednesday amid narrow trading. The S&P 500, Nasdaq, and Dow Jones all edged down, ending the Dow's two-day winning streak. The key pressure came from July's PCE inflation data, which showed a 3.7% year-over-year increase, exceeding expectations. While the core PCE met forecasts at 3.3%, the hot headline number boosted Treasury yields and the dollar, dampening hopes for imminent Fed rate cuts. Gold fell below $4,600/oz. Oil prices continued to weaken despite mixed geopolitical signals. The market's real focus was after the close. Nvidia reported Q2 revenue of $96.2 billion, beating estimates, guided for current-quarter revenue to surpass $100 billion for the first time, and projected 70% revenue growth for the next fiscal year. Its shares rose approximately 4% after-hours, lifting Nasdaq futures by about 1%. Amazon's announcement to deploy an additional 2 million GPUs further validated the data center demand narrative for Nvidia. In other sectors, chip stocks like Western Digital and Arm gained. Among the "Magnificent Seven," moves were mixed. Salesforce surged on strong guidance and an expanded partnership with Anthropic. Meta settled a youth addiction case with 29 U.S. states for up to $18 billion, a figure seen as favorable compared to earlier fears. Bitcoin retreated from recent highs, while industrial metals like copper extended their rally. The core market tension remains between inflation concerns pressuring the broader market and strong AI earnings providing sector-specific momentum, setting the tone for early September trading.

marsbitHace 8 min(s)

US Stock Market Trends (Aug 27): PCE Exceeds Expectations, Pressures Broader Market; Nvidia Rises 4% After-Hours, Nasdaq Futures Up 1%

marsbitHace 8 min(s)

A Record $70 Billion Inflow in 5 Days! Investors Are No Longer Choosy, Buying Gold and Bitcoin Together

Investors are moving beyond "either/or" choices and are simultaneously pouring money into both gold and bitcoin ETFs. Over the past five trading sessions, ETFs tracking these assets attracted a record $7 billion in combined inflows, pushing some of the largest gold and bitcoin funds to the top of the U.S. weekly ETF inflow rankings. This surge was triggered by U.S. Treasury Secretary's announcement to at least double long-term bond buybacks, which initially pressured Treasury yields and the dollar, boosting prices for both assets. Gold has risen about 13% this month, while bitcoin reclaimed the $80,000 level. The synchronized rally signals the return of the "monetary debasement trade." Amid growing concerns over fiscal sustainability and easing financial conditions, investors are seeking scarce assets perceived as outside direct government control. Gold benefits from its traditional safe-haven role, while bitcoin's fixed supply of 21 million coins positions it as a potential hedge. The SPDR Gold ETF (GLD) attracted nearly $3.4 billion, and the iShares Bitcoin Trust (IBIT) saw $1.5 billion in inflows, both ranking in the weekly top ten. Analysts note the momentum behind the flows is as significant as the volume, indicating investors are aggressively adjusting previously underweight positions. While the narrative of hedging against fiscal stress and currency debasement is gaining traction, some analysts question its sustainability, suggesting equities might be a more reliable hedge in the long run.

华尔街日报Hace 14 min(s)

A Record $70 Billion Inflow in 5 Days! Investors Are No Longer Choosy, Buying Gold and Bitcoin Together

华尔街日报Hace 14 min(s)

Yangtze Memory: Is It the Second ChangXin?

The largest IPO in the history of the STAR Market is approaching. Yangtze Memory Technologies (YMTC) has filed for a listing on the Shanghai Stock Exchange's STAR Market, aiming to raise 33 billion yuan. This surpasses the previous record set by competitor ChangXin Memory Technologies (CXMT), which raised 29.5 billion yuan and saw its market capitalization surge on its debut. Despite both being leading Chinese memory chipmakers founded in 2016 and operating under the IDM model, the two companies are fundamentally different. CXMT focuses on DRAM, the memory used for temporary data processing, while YMTC specializes in NAND Flash, used for long-term data storage. Industry reports indicate the global DRAM market is significantly larger and more concentrated among three major players, where CXMT ranks as the fourth-largest supplier. The NAND Flash market is more fragmented, with YMTC ranking third globally by shipment volume in Q2 2026, though fifth by revenue due to a stronger focus on consumer-grade products. Experts are divided on whether YMTC can replicate CXMT's explosive market debut. Some analysts believe it is highly unlikely, citing a cooler market environment and YMTC's perceived lower strategic scarcity within the AI supply chain compared to DRAM/HBM-focused companies. They warn that aggressive IPO pricing could lead to downward pressure post-listing. Industry forecasts suggest the DRAM market may remain tight, while the NAND Flash market could face price corrections due to new capacity and weaker demand. Other experts argue that while short-term market sentiment differs, the long-term investment value of both companies is comparable, hinging on future performance and potential breakthroughs in areas like HBM. They believe YMTC's IPO is unlikely to face significant cooling given still-elevated market interest in the semiconductor sector.

marsbitHace 34 min(s)

Yangtze Memory: Is It the Second ChangXin?

marsbitHace 34 min(s)

Tether Completes Full Audit by KPMG U.S., 150 Tons of Gold and 100,000 BTC Reserves Verified

Tether Announces Completion of Full Financial Audit by KPMG US, Confirming Gold and BTC Reserves Tether CEO Paolo Ardoino announced that the company has completed its first-ever full financial audit by the "Big Four" accounting firm KPMG US. The audit covers the financial statements of Tether International, S.A. de C.V. as of December 31, 2025. Ardoino stated that Tether's delay in obtaining a full audit was primarily due to the previous US administration's hostile stance towards the digital asset industry. With a more favorable regulatory environment, the company engaged KPMG US for the rigorous review. Ardoino emphasized the audit confirmed the strength of Tether's reserves, including over $6 billion in excess equity. He also revealed the company holds substantial physical assets: approximately 150 tons of gold and over 100,000 Bitcoin (BTC). The audit process reportedly involved physically verifying each gold bar. Addressing past criticism and ongoing skepticism, Ardoino pointed to Tether's ability to handle large-scale redemptions, citing the processing of around $7 billion within 48 hours during market stress in 2022. He stated the company plans to undergo annual full financial audits alongside its ongoing quarterly attestation reports. Regarding a potential private equity round, Ardoino clarified that Tether, as a highly profitable company, does not require external capital but acknowledges significant market interest in its shares. The company remains private and mission-focused, aiming to serve users outside the traditional financial system. Looking ahead, Ardoino expressed Tether's interest in potentially funding initiatives like the "Bitcoin Red Team," which uses AI to audit Bitcoin's code and wallet security for vulnerabilities.

marsbitHace 53 min(s)

Tether Completes Full Audit by KPMG U.S., 150 Tons of Gold and 100,000 BTC Reserves Verified

marsbitHace 53 min(s)

Wall Street's Imagination Can't Keep Up with Nvidia's 'Speed'

NVIDIA once again delivered stellar earnings, with revenue, operating profit, and EPS all hitting record highs for Q2 FY2027 (ended July 26, 2026). Revenue reached $96.221 billion, a 106% YoY increase, while net profit surged 126% to $59.688 billion. The core driver remained the Data Center segment, which grew 117% YoY to $89.023 billion, fueled by ongoing global AI infrastructure investments. The company segmented its Data Center revenue into Hyperscale (large cloud providers) and ACIE (AI cloud, enterprise, industrial, and sovereign AI) categories, with the latter showing faster growth. Notably, shipments of Hopper architecture products to mainland China accounted for less than 1% of Data Center revenue this quarter, and future guidance excludes contributions from China's Data Center compute market. CEO Jensen Huang stated that AI has reached a "tipping point," directly generating productivity and revenue. Financially, NVIDIA maintained robust profitability with a gross margin of 75%. The company also returned approximately $26 billion to shareholders via buybacks and dividends. Looking ahead, product transitions are key. The Blackwell Ultra platform is now in large-scale deployment, while the next-generation Vera Rubin platform, including NVIDIA's first CPU designed for AI agents, has entered full production. The company is also expanding its software ecosystem with tools like the DSX platform for building AI factories. However, NVIDIA's role is evolving beyond chip supplier. It is increasingly involved in financing and facilitating massive AI infrastructure projects, exemplified by its credit support for the SB Energy project in Ohio intended for OpenAI. The company aims to mobilize over $500 billion in third-party capital for such builds, though this raises questions about future capital intensity. For Q3 FY2027, NVIDIA forecasts revenue of approximately $108 billion, signaling continued growth but at a potentially moderated pace. While still dominant, the competitive landscape is intensifying with rivals like AMD and in-house chips from major cloud customers. Key challenges include maintaining growth through the Blackwell-to-Rubin transition, sustaining the pace of AI infrastructure investment, and defending market share as inference workloads grow and competition increases.

marsbitHace 53 min(s)

Wall Street's Imagination Can't Keep Up with Nvidia's 'Speed'

marsbitHace 53 min(s)

Trading

Spot
活动图片