Ilya Delivers, Huang Jen-Hsun Bets 50 Billion on SSI, First Model Suspected to Launch This Week

marsbitPubblicato 2026-08-25Pubblicato ultima volta 2026-08-25

Introduzione

Ilya Sutskever's Safe Superintelligence (SSI) lab, backed by a rumored $5 billion investment from NVIDIA granting exclusive access to next-gen systems, may release its first groundbreaking model as early as this week. Industry insiders are hinting at a major breakthrough from a non-major player, with some claiming AI will be "utterly transformed." The model is speculated to fundamentally differ from current large language models like ChatGPT. Instead of relying on massive, static pre-training and ever-larger context windows, SSI's approach reportedly centers on **Test-Time Training (TTT)**. This would allow the model to **update its own internal weights in real-time** as it processes information, effectively "learning on the fly" and internalizing new knowledge like a human, rather than just referencing it from a large prompt. This aligns with Ilya Sutskever's long-stated vision of moving beyond the "pre-training" and "scaling" era. He envisions superintelligence not as a monolithic, pre-loaded database, but as an adaptive, "infinitely curious 15-year-old genius" capable of continual learning through interaction. SSI's strategy has been focused solely on this goal of "safe superintelligence" through efficient continual learning and alignment, explaining its prior silence. If successful, this model could disrupt the entire AI industry, challenging the current reliance on vast compute resources for training and the business models built around context windows. It suggests th...

Ilya Sutskever's Safe Superintelligence (SSI) lab may release its first groundbreaking model as early as this week!

Industry insiders are heavily hinting that SSI will launch its first disruptive large model.

Silicon Valley Bigwigs Crazy Teasers: "AI Will Be Completely Upended!"

Well-known investor and a16z partner Martin Casado was so excited he directly hinted that he recently saw "the most important new model of the year."

The industry immediately erupted, with attention instantly focusing on Ilya's SSI.

The reason is simple: a16z is one of the core investors behind SSI. Although some speculate that what Casado saw might be OpenAI's next-generation Astra model.

But rumors of a breakthrough from a "non-mainstream" player are rampant.

Notable tech observer Andrew Curran was even more excited to point out that this major breakthrough doesn't seem to come from a mainstream giant (like OpenAI or Google), but from an independent lab, which feels "very real."

The most exaggerated claim comes from AI watcher Dan McAteer, who outright stated: "Ilya has truly created superintelligence, the game has changed!"

Earlier this month, investor Gavin Baker flatly stated on a podcast:

SSI says they will release their model in August.

If you think this is just talk, look at where the real money is flowing.

On July 27th of this year, NVIDIA suddenly announced a long-term strategic partnership with this mysterious company with no revenue and no product, and will insanely increase SSI's computing power scale by 10 times in the next 12 months.

There's even hotter gossip: the investment is as high as $50 billion, and NVIDIA has granted SSI exclusive access to the next-generation Vera Rubin system!

Why is Huang Jen-Hsun willing to bet such a huge sum?

In the press release, NVIDIA stated that they decided to make this heavy bet only after "gaining rare access to its closely guarded research results."

What did Huang Jen-Hsun see? Combined with today's leaks, the answer seems obvious.

Upending Common Sense: AI Learned to "Think and Learn On the Fly"

To understand how shocking SSI's new model is, we must first understand how current AI (like ChatGPT, Claude) works.

Current AI models, regardless of how many parameters they have, essentially solidify knowledge into their weights during the "pretraining" phase. Once training is complete, its "brain" is frozen.

To allow models to process new information, major companies are frantically expanding the "Context Window," from 100k to millions of tokens.

This is like an open-book exam—the model's brain doesn't get smarter; you're just allowing it to bring a thicker and thicker "cheat sheet" (context) into the exam.

Even reasoning models like OpenAI's o1, which focus on "Test-time Compute," are essentially just spending more time thinking on scratch paper (consuming more tokens). The connections in its neural network, its "brain cells," are still static.

But what SSI has reportedly achieved is a new architecture based on Test-Time Training (TTT).

What does this mean?

When this model reads a long document you provide, it doesn't stuff the document into a "cheat sheet." Instead, it truly "learns" it, generates gradient updates, and changes the structure of its own brain.

After finishing reading, it has already become a minutely but entirely new, evolved AI.

It no longer needs a massive context window because it has turned what it read into its true "internalized knowledge."

It is no longer limited by the computing power monopoly of pretraining. A small, refined model that can continuously adapt and evolve on the job could easily surpass those massive, inflexible behemoths built with vast computing power.

While the industry currently competes on "how long a model can think," Ilya is betting on "whether a model can change itself."

Ilya's Crazy Foresight

Looking back over the past two years, SSI's "zero products, zero papers" status once made outsiders wonder if they were stuck.

But if you connect the timeline, you'll find Ilya has been playing a long game.

Ilya has repeatedly conveyed a core idea to the public: the era of Pretraining as we know it is coming to an end.

At the 2024 NeurIPS conference, he made this astonishing prediction.

By November 2025, on Dwarkesh Patel's podcast, he went further: "We are moving from the Scaling era to the Research era."

Ilya believes the entire industry has been misled by the terms "AGI" and "pretraining." Humans are not born omniscient.

In his mind, true superintelligence is not a massive machine that memorized the entire internet at birth; it is an "extremely smart, infinitely curious 15-year-old genius."

This 15-year-old might initially know nothing, but if you place them in any role, through continuous trial, error, and learning, they can quickly master programming, medicine, law, or even any unknown skill.

He further explained: Humans themselves are not "AGI out of the box," but rely on continual learning.

True superintelligence should be the same: deployment itself is a learning process, evolving through real-world feedback, not being "finished" after a one-time pretraining.

This philosophy directly determines SSI's strategy: no pursuit of short-term products, no release of intermediate models, only aiming for a "straight shot to safe superintelligence," focusing research on efficient continual learning and alignment.

To achieve this, "Meta-learning" is the only solution—the model must not only master skills but also master "the method of acquiring skills."

This perfectly explains why SSI has been unusually low-key. If you're building a new species that completely overturns the existing paradigm, you wouldn't be blogging about it before the paradigm is even fully built.

But the truth can't stay hidden forever; clues were already planted.

In July 2024, scholar Yu Sun et al. published the original TTT paper.

Most crucially, Stellar co-founder and SSI investor Jed McCaleb co-authored a paper, bluntly stating:

Long-context language modeling is not an architectural problem at all, but a continual learning problem!

Consistent research direction, investors personally co-authoring papers, and now today's leaks—all clues point to the same fact: SSI has turned TTT from an academic concept in the lab into a real commercial weapon.

Conclusion: The Second Half of AI Has Just Begun

Now, all eyes are on this August.

Whether it's a test version accessible only to a small group of geeks or a stunning public release, if SSI's first model truly possesses "Test-Time Training" and "real-time weight updating" capabilities, the entire logic of the AI industry will be overturned.

The computing power moats piled up in data centers by major companies, the business model charging by the million tokens, even the debate over open-sourcing weights—all will face a dimensional reduction strike.

This proves one thing: the AI race is far from the final stage where "money and computing power can guarantee victory." Real technological leaps are still hidden in the top minds daring to break common sense.

This time, Ilya Sutskever stands again at history's crossroads. Back then, it was his line of code that brought deep learning back into the light; today, perhaps it is him again who will personally end the "Pretraining era" of large models.

Do you think Ilya can ascend to legend status again this time? If AI can truly "evolve as it's used," how far is humanity from completely losing control?

References:

https://x.com/hakmgpt/status/2091855200713638146

https://x.com/daniel_mac8/status/2091891607641440598

https://x.com/AndrewCurran_/status/2091890441465499995

https://x.com/martin_casado/status/2091650951736361073

https://aimidday.com/ssis-first-model-reportedly-trains-itself-while-it-thinks/

This article is from the WeChat public account "New Zhiyuan," author: ASI Revelation; editor: David

Domande pertinenti

QWhat is the core technology that SSI's rumored new model is said to be based on, and how does it fundamentally differ from current models like ChatGPT?

ASSI's new model is rumored to be based on Test-Time Training (TTT). Unlike current models that work with fixed, pre-trained weights and use large context windows to process new information, a TTT-based model can update its internal parameters (weights) in real-time as it processes data. This means it truly 'learns' and adapts itself during usage, rather than just referencing a larger 'cheat sheet' of context.

QWhy did NVIDIA decide to invest heavily in SSI according to the article, and what specific advantage did they gain from the deal?

ANVIDIA decided to invest heavily in SSI (reportedly $5 billion) after gaining 'rare access' to SSI's closely guarded research. As part of the strategic deal, NVIDIA will significantly scale SSI's compute power and has granted SSI exclusive access to its next-generation Vera Rubin system.

QHow does Ilya Sutskever's vision for a true superintelligence differ from the current industry concept of AGI?

AIlya Sutskever believes the industry is misguided by the concepts of 'AGI' and 'pretraining.' He envisions a true superintelligence not as a massive machine pre-trained on all internet data, but as an 'extremely smart, infinitely curious 15-year-old genius.' This intelligence wouldn't be born omniscient but would possess a fundamental capability for continual learning, rapidly mastering any skill or domain through real-world interaction and feedback, much like humans do.

QAccording to the article, what is a key strategic difference between SSI and other major AI labs, and how does this relate to their research focus?

AA key strategic difference is that SSI has pursued a 'zero product, zero paper' approach, avoiding interim model releases and short-term products. Their stated strategy is a 'straight shot to safe superintelligence,' focusing all research on efficient continual learning and alignment. This aligns with their goal of creating a foundational breakthrough rather than iterating on the existing pretraining paradigm.

QWhat potential impact could a successful SSI model with Test-Time Training capabilities have on the current AI industry landscape?

AA successful TTT-based model could radically disrupt the current AI industry. It could render massive compute power investments (used for pretraining) less of a decisive advantage, challenge business models based on charging per token for long contexts, and make debates about open-sourcing static model weights less relevant. It would shift competition from 'how long a model can think' to 'how well and quickly a model can adapt and learn.'

Letture associate

Demand Test: Bitcoin Whales Earn Record $1.2 Billion, Ethereum Holders Return to Profitability

In just three days after Bitcoin's price recovery, new Bitcoin whales have realized over $1.2 billion in profit, marking the largest profit-taking event for this cohort on record according to CryptoQuant. The peak occurred on August 20 with roughly $614 million, setting a daily record. Analysts note this selling pressure began after Bitcoin rose above the realized price of short-term whales, which was around $68,900. With Bitcoin trading near $77,700 on August 23, these whales were sitting on an average profit of about 12.8%. The market recovery allowed investors who were previously at breakeven or at a loss to lock in gains. CryptoQuant described the situation as a key test for Bitcoin demand; sustained prices above whale cost-basis with normalized profit-taking could signal strong new demand, while continued selling pressure could turn the rally into a mere break-even exit. Simultaneously, large Ethereum holders have also returned to an unrealized profit zone following its rally, as noted by CryptoQuant analyst Darkfost. Current profit levels, however, remain relatively low and are not seen as creating significant selling pressure. The Unrealized Profit/Loss Ratio for different whale cohorts stands at 0.075 (for 1k-10k ETH holders), 0.16 (10k-100k ETH), and 0.38 (over 100k ETH). This marks a significant improvement from June, when these whales were in substantial unrealized loss, with Ethereum having risen over 65% since then. The increased profitability is viewed as a positive sentiment signal for the Ethereum market.

cryptonews.ru51 min fa

Demand Test: Bitcoin Whales Earn Record $1.2 Billion, Ethereum Holders Return to Profitability

cryptonews.ru51 min fa

Kinetiq Team Announces Elysium L2 Network for Hyperliquid

On August 24, the liquid staking protocol Kinetiq announced Elysium, a new L2 network for the Hyperliquid ecosystem. It aims to increase HyperEVM's throughput and simplify the launch of spot markets, tokens, and DeFi applications. Elysium will use $HYPE for gas fees and plans direct integration with HyperCore's trading engine, giving apps access to its liquidity and orderbook data. Technical details and partners will be revealed later, with a launch date set for "soon." A primary reason for Elysium's development is to overcome HyperEVM's limitations, such as low throughput and rising fees during high load. Kinetiq claims the L2 will start with significantly higher block production and transaction speeds, later aiming to approach HyperCore's performance. This targets applications needing frequent state updates like high-frequency spot trading and AMMs, and will provide them deeper access to HyperCore orderbook data. The network also proposes to streamline the process of launching new assets within Hyperliquid, allowing a token to progress from AMM liquidity to HyperCore's spot orderbook and eventually to perp markets via HIP-3 in a unified flow. Regarding revenue, Kinetiq's model allocates 50% of sequencer fees to buy back and burn $KNTQ, 25% to developers using Elysium's block space, and 25% to the Kinetiq treasury. Elysium marks Kinetiq's expansion beyond its core liquid staking product, kHYPE.

cryptonews.ru52 min fa

Kinetiq Team Announces Elysium L2 Network for Hyperliquid

cryptonews.ru52 min fa

Trading

Spot
活动图片