OpenAI Researcher Exposes ASI Timeline: Most Have Become Reality

marsbit2026-08-17 tarihinde yayınlandı2026-08-17 tarihinde güncellendi

Özet

In April 2025, a group of former OpenAI researchers published a 71-page document titled "AI 2027," outlining a timeline for Artificial Superintelligence (ASI). Their predictions, now being tracked by an independent project, show 51% are already confirmed, ahead of schedule, or on track. Notably, alarming predictions are arriving faster than anticipated. The forecast that AI would achieve top-tier human-level capabilities in cyber offense and defense by early 2027 was realized in April 2026, nine months early. Similarly, major Pentagon contracts with leading AI labs were signed 18 months earlier than predicted. The core mechanism for an intelligence explosion—Recursive Self-Improvement (RSI), where AI accelerates its own development—has not yet closed its loop. While AI, like Anthropic's Claude, now writes most new code, the bottleneck has shifted to human review and high-level research direction. A July 2026 study indicates the current AI-driven productivity gain in R&D is about 9%, below the estimated 15% threshold needed for a self-sustaining RSI feedback loop. However, underlying capabilities continue to accelerate rapidly. The "time horizon" metric for AI to autonomously handle tasks is doubling every three months, suggesting monthly-scale autonomous operation could be feasible by early 2027. Consequently, the original authors have revised their median prediction for fully automated AI programming forward to around mid-2028.

In April 2025, a group of researchers who resigned from OpenAI published a 71-page document, outlining a chilling future month by month:

By the end of 2026, AI coding agents begin to replace junior programmers;

By 2027, superhuman coding agents automate AI R&D itself, triggering an intelligence explosion;

Before the end of that year, humanity might face a superintelligence (ASI) it created but cannot control.

This forecast, named "AI 2027," reads like science fiction.

Its lead author, Daniel Kokotajlo, also wrote "What 2026 Looks Like" in 2021, accurately describing the emergence of chain-of-thought reasoning and agents before ChatGPT existed, with over half of the specific predictions eventually coming true.

https://www.youtube.com/watch?v=_g4l7YkDQwA

Now it's time for "AI 2027" to face scrutiny.

Johannes Haus, an independent tracker from Hamburg, Germany, extracted 53 verifiable predictions from it and built the AI 2027 Tracker to score them one by one.

https://ai2027tracker.com/timeline

The scorecard as of now: 51% of the predictions have been confirmed, ahead of schedule, or are on track.

Reality is unfolding at 70% of the predicted speed.

But what truly makes one uneasy is far more than just these two numbers.

The Scariest Predictions Are Arriving Early

Out of the 53 predictions, 3 are "ahead of schedule." The most unsettling one: AI gains cyber offense/defense capabilities close to top-tier human hackers.

"AI 2027" placed this event in early 2027.

It actually occurred in April 2026, 9 months early.

After Anthropic released Claude Mythos Preview, they deployed it to several open-source projects under the Project Glasswing framework. This model autonomously discovered thousands of zero-day vulnerabilities, some of which had remained hidden for ten to twenty years under the review of human security experts.

The key point: Mythos Preview wasn't even trained for cyber offense/defense.

It's just a general-purpose model that learned to code and reason; discovering vulnerabilities was a side effect.

Anthropic's internal evaluation stated that AI, in terms of coding capability, can already surpass most humans in finding and exploiting software vulnerabilities.

In July 2026, OpenAI's System Card disclosed more direct evidence: during an evaluation, a model exploited a zero-day vulnerability, reached Hugging Face's production infrastructure, bypassed sandboxes, and obfuscated authentication tokens.

A report from the UK's AISI in the same month stated that GPT-5.5 completed end-to-end multi-step cyber-attack simulations.

Another early prediction: The Pentagon draws AI labs into defense contractor relationships.

"AI 2027" predicted this would happen in early 2027. In reality, the Pentagon signed four contracts worth $200 million each in June 2025, awarded to Anthropic, OpenAI, xAI, and Google, 18 months ahead of the script.

Tracker maintainer Haus summarized the pattern behind these cases on LessWrong: Risks are arriving faster than the original capabilities that generate these risks.

This finding holds systematically across the 53 predictions and is the most cautionary insight in the entire scorecard.

The Only Good News: The Final Step Hasn't Been Taken

The core mechanism leading to ASI is a feedback loop, known recently as RSI (Recursive Self-Improvement): AI accelerates AI R&D, the results make the next-generation AI stronger, and the stronger AI accelerates R&D again, and so on.

The entire second half of the "AI 2027" plot is built on the assumption of this loop closing.

However, this loop hasn't closed yet.

This is the biggest gap in the scorecard and, in a sense, the only good news.

Anthropic disclosed in May 2026 that Claude had already written over 80% of the company's new code; a year prior, that number was in the single digits.

In an internal test, Mythos Preview optimized an ML training code to 52 times the baseline; human engineers' results were about 4 times.

But Anthropic itself admitted the bottleneck has shifted: The speed of generating code is now fast enough that there's a growing backlog of code waiting for human review.

The faster AI writes code, the greater the workload for human reviewers.

At the research level, the bottleneck is "taste," the judgment to decide the direction of research. The first half of the feedback loop is turning, but the second half hasn't connected yet.

A paper published in July 2026 by the Elasticity Institute (members include Tom Cunningham from METR) gave the precise threshold needed for "connection": Each generation's model capability improvement must yield at least a 15% increase in AI R&D productivity for RSI to become self-sustaining.

Based on System Card data, the paper estimated the current number is about 9%, below the critical point.

https://x.com/AnikaSomaia/status/2087408169660064218

The 6 percentage points between 9% and 15% constitute a buffer between humanity and the acceleration loop.

But the underlying curve supporting this buffer continues to accelerate.

METR's Time Horizon metric tracks the duration of tasks an AI can handle autonomously.

Data from January 2026 showed this metric doubles every 3 months, and the rate of doubling itself is accelerating.

Claude Opus 4.6's time horizon reached about 12 hours, while Mythos Preview hit the measurement ceiling.

Extrapolating the 3-month doubling: 12 hours to 24 hours to 48 hours to a week to two weeks, monthly-level tasks arrive by early 2027.

The authors of "AI 2027" are also adjusting their own expectations.

Kokotajlo moved the median prediction for fully automated programming from the end of 2029 to mid-2028; Lifland's estimate is around mid-2030.

The RSI cycle, like stepping on one's own foot, is beginning to bring about the exponential acceleration of ASI's arrival.

References:

https://ai2027tracker.com/?p=timeline

https://github.com/elasticity-ai/elasticity/raw/main/paper/elasticity-rsi-paper.pdf

This article is from the WeChat public account "New Zhiyuan," author: Ma Ke

İlgili Sorular

QAccording to the AI 2027 prediction tracker, what percentage of the 53 verifiable predictions have been confirmed, are ahead of schedule, or are progressing as planned?

A51% of the predictions have been confirmed, are ahead of schedule, or are progressing as planned.

QWhat is the name of the feedback loop mechanism considered central to the potential arrival of ASI (Artificial Superintelligence), and what is its current status according to the article?

AThe mechanism is called RSI (Recursive Self-Improvement). According to the article, this feedback loop has not yet closed. While AI is accelerating AI research, the resulting productivity gains (estimated at 9%) are below the critical threshold (15%) needed for the cycle to become self-sustaining.

QWhich specific prediction from the 'AI 2027' report arrived 9 months earlier than forecasted, and what event in 2026 demonstrated it?

AThe prediction that AI would acquire cyber offense and defense capabilities near the level of top human hackers arrived 9 months early. This was demonstrated in April 2026 when Anthropic's Claude Mythos Preview model autonomously discovered thousands of zero-day vulnerabilities in open-source projects as a byproduct of its general capabilities.

QWhat does the article identify as the 'only piece of good news' regarding the timeline to ASI?

AThe 'only piece of good news' is that the core RSI (Recursive Self-Improvement) feedback loop, where AI accelerates AI research in a self-sustaining cycle, has not yet closed. This creates a buffer, as current AI-driven productivity gains in research are below the critical threshold needed for the cycle to become autonomous and exponential.

QAccording to the Elasticity Institute's July 2026 paper cited in the article, what is the precise productivity increase threshold needed for RSI to become self-sustaining, and what is the current estimated figure?

AThe Elasticity Institute's paper states that each generation of AI models must deliver at least a 15% increase in AI R&D productivity for RSI to become self-sustaining. The current estimated figure is approximately 9%, which is below this critical threshold.

İlgili Okumalar

Is RWA Still Meaningful Without DeFi?

The article "Would RWA Still Matter Without DeFi?" argues that tokenizing real-world assets (RWA) alone, like putting a barcode on a container, is not transformative. True value emerges when these tokenized assets are integrated into decentralized finance (DeFi) ecosystems, enabling valuation, financing, hedging, trading, and loss management in a programmable, automated manner. Tokenization provides digital representation, but DeFi provides utility through leverage, liquidity, and composability. The core challenge lies in aligning the different "time clocks" of blockchain (fast, 24/7), traditional markets (limited hours), and asset redemption (slow processes), which creates liquidation risks and gaps. Effective RWA integration requires more than a token; it needs a full stack: legally enforceable rights, reliable data oracles, clear transfer rules, executable secondary liquidity, appropriate collateral parameters, and credible loss resolution paths. Liquidity is defined not by total value locked (TVL) but by the ability to exit a position under stress within a required timeframe. Risk management for RWAs must be modeled as a dependency graph, monitoring interconnected nodes like issuers, custodians, oracles, and liquidity pools for early warning signs beyond just price data. While tokenized government bonds serve as an initial "ping test," the future lies in more complex assets like computing power and energy, which require bespoke risk models. Tokenized stocks paired with perpetual futures present a major test, combining global equity ownership with crypto-native leverage, necessitating robust architectural safeguards like isolation and dynamic collateral rules. The conclusion is that without DeFi, RWA tokenization offers limited value—improving distribution and transparency. The significant opportunity arises when tokenized assets become functional components within open, programmable capital markets, where they can be used as collateral and facilitate complex financial strategies. The token is merely the barcode; the market operating system is the real machine.

marsbit9 dk önce

Is RWA Still Meaningful Without DeFi?

marsbit9 dk önce

Government Intervention in the Bond Market: What Does It Mean?

On August 19, 2026, the U.S. Treasury unexpectedly announced it would at least double the size of its long-term Treasury buyback operations, starting September 9. The move triggered an immediate market reaction, with the 30-year yield falling 9 basis points. This intervention came against a backdrop of the 30-year yield hitting a 19-year high of 5.34% the previous day, driven by persistent inflation, high oil prices, deteriorating U.S. fiscal health with public debt surpassing $40 trillion, and a global "buyers' strike" for long-dated bonds. Treasury buybacks involve the government repurchasing older, less liquid bonds from the market via reverse auctions to improve market functioning, not to reduce overall debt. While the action provided tactical relief and signaled the Treasury's willingness to intervene, analysts caution it does not address core structural issues: massive fiscal deficits, high borrowing needs, and fading demand from traditional buyers like foreign central banks. For investors, the announcement offered short-term support for long-duration bond ETFs and boosted assets like gold, which rose 2.7%. However, it's unlikely to significantly lower mortgage rates or alter the challenging environment for long-term bonds. The key takeaway is the distinction between a tactical market operation and the unresolved structural pressures that continue to push yields higher, requiring close monitoring of upcoming economic data, Treasury auctions, and potential signals of more aggressive policy measures.

marsbit13 dk önce

Government Intervention in the Bond Market: What Does It Mean?

marsbit13 dk önce

After 'Bessent Put', How Far Is the U.S. from Restarting QE?

Following an unscheduled announcement from the U.S. Treasury Department on August 19th, investor discussions have intensified regarding the potential for future policy interventions in long-term bond markets. The Treasury increased the maximum size of its regular buyback operations for 10-to-30 year bonds from $2 billion to $4 billion, citing a desire to improve liquidity. This move came shortly after a surge in long-term yields, with the 30-year Treasury yield briefly touching 5.34%. Market analysts, rather than focusing on the modest operational size, have interpreted the timing—outside the normal quarterly communication window—as a significant signal. The move has been dubbed the "Bessent Put," implying the market's growing expectation that Treasury officials, led by Deputy Secretary Josh Bessent, may act to prevent a disorderly rise in long-term borrowing costs. This perception represents a potential shift in the market's view of the government's "policy reaction function." The underlying pressures on long-term bonds are multifaceted, including large fiscal deficits, increased Treasury supply, a rise in corporate debt issuance for AI infrastructure (creating a "crowding out" effect), and uncertainty around foreign holdings, particularly from Japan. Geopolitical risks in the Middle East further complicate the policy landscape, potentially creating conflicting pressures between fighting inflation and managing financing costs. However, the article clarifies that this Treasury buyback program is distinct from Quantitative Easing (QE). It is a debt management operation, not a Federal Reserve balance sheet expansion. While the announcement opens the door for market speculation about more forceful tools like Yield Curve Control (YCC) or a return to QE, analysts note that conditions would need to deteriorate significantly for such measures to be implemented. For now, the "Bessent Put" reflects a change in market expectations about possible policy boundaries, not an imminent launch of new monetary stimulus.

marsbit14 dk önce

After 'Bessent Put', How Far Is the U.S. from Restarting QE?

marsbit14 dk önce

Is Stablecoin Really Necessary for Cross-border Payments?

"Is Stablecoin Truly Necessary for Cross-Border Payments? While stablecoins are often praised as superior for cross-border transfers—especially if the recipient desires crypto—their advantage is less clear in traditional cross-currency scenarios (e.g., USD to Mexican Peso). The existing correspondent banking system is slow and costly (≈15% fees) due to multiple intermediaries. Modern fintech solutions like Wise have dramatically improved this by using a netting model: they hold local currency pools and settle payments domestically, avoiding actual cross-border fund movement. This offers near-instant transfers with low, transparent fees (averaging ~0.52%). The 'stablecoin sandwich' model (convert fiat to stablecoin, transfer on-chain, convert to local fiat) offers similar user experience but doesn't inherently provide major cost or speed advantages over fintech. Its on-chain transfer is cheap, but fiat conversion spreads remain. Stablecoin's real innovation is 'unbundling' the cross-border payment stack. Instead of requiring a proprietary global network like Wise, businesses only need reliable on- and off-ramps in specific corridors. This lowers market entry barriers, fosters competition among local providers, and can drive down costs, particularly for niche or underserved corridors (e.g., US to Africa). While vertical integration may reoccur, the open, permissionless nature of the underlying blockchain layer makes monopolistic pricing difficult. The long-term benefit is the potential redistribution of value—previously captured as intermediary rents—to consumers through lower costs."

marsbit18 dk önce

Is Stablecoin Really Necessary for Cross-border Payments?

marsbit18 dk önce

Ethereum Finally 'Bounces Back'! ETH Returns to the Golden Line, What's Different About This Rally?

Ethereum has surged above $2,300, a level not seen in over three months, marking a significant recovery and technical breakthrough by reclaiming key resistance levels, including the weekly EMA50 for the first time in the current bear market. The rally has been fueled by a combination of factors: improved market risk appetite, positive regulatory sentiment, and a major short squeeze where over 88% of recent ETH liquidations were short positions. The ETH/BTC ratio has also broken its long-term downtrend, indicating renewed strength for Ethereum relative to Bitcoin. Fundamentally, Ethereum spot ETFs have shown strong inflows, outperforming their Bitcoin counterparts in recent months. Wall Street institutions like Morgan Stanley and JPMorgan have notably increased their ETH exposure, while several banks have added or expanded positions in ETH ETFs. On-chain fundamentals remain robust, with the proportion of staked ETH reaching a new all-time high of nearly 33.7%. Data indicates long-term holding sentiment, as small wallet holdings increased while large outflows were often directed towards staking or contracts, not selling. However, potential headwinds include declining staking yields and community debate over proposals to adjust staking rewards at high participation levels. The upcoming Glamsterdam upgrade, featuring improvements to validator exit efficiency, could enhance liquidity and further attract institutional stakers. While the rebound is supported by sentiment, capital flows, and fundamentals, its sustainability hinges on continued positive developments.

marsbit18 dk önce

Ethereum Finally 'Bounces Back'! ETH Returns to the Golden Line, What's Different About This Rally?

marsbit18 dk önce

İşlemler

Spot
活动图片