AlphaGo's Creator Puts AI into a 23-Year-Old Artificial Society: All Three Toughest Challenges for AI Agents Are Here

marsbitPublished on 2026-05-25Last updated on 2026-05-25

Abstract

Demis Hassabis, CEO of DeepMind, has embarked on a new AI research venture by partnering with the long-running space MMO, EVE Online. This collaboration, announced in early May, aims to use the game's 23-year-old, player-driven persistent universe as a testbed for tackling three core challenges in AI agent research: long-horizon planning, memory, and continual learning. Unlike previous DeepMind environments like AlphaGo (Go) or AlphaStar (StarCraft II), EVE Online features no fixed end state. Its single-shard universe has fostered complex, emergent player societies with real economies, political alliances, and wars that can span months or years. These conditions naturally demand the very skills—long-term strategic planning, maintaining memories over extended periods, and adapting to constant change—that are hardest for current AI agents to master. The research will initially use an offline version of EVE, providing a controlled, complex sandbox without interfering with the live player server. This move continues DeepMind's trajectory of using increasingly complex and open-ended virtual worlds for AI training, from Atari games and Go to StarCraft II and the SIMA project. The EVE environment represents a significant step towards testing AI in a persistent, socially complex, and continuously evolving world shaped by human behavior over decades.

DeepMind CEO and AlphaGo creator Demis Hassabis has been using games for AI research for over a decade.

This time, he has thrown AI into a "living universe" that has been running for 23 years: the space-themed massively multiplayer online game EVE Online, a game whose new player tutorial alone can deter players.

Chess games have an end, but EVE does not.

In early May, DeepMind officially announced a research collaboration with EVE Online for a simple reason: EVE's complex, player-driven universe is the perfect safe sandbox to test AI memory, continual learning, and long-term planning.

DeepMind's collaboration with EVE is not about pursuing fun gameplay or enhancing game mechanics. Instead, it aims to tackle the three toughest, most widely recognized challenges in current AI agent research. Hassabis is betting on finding answers in a 23-year-old game.

Fenris Creations (formerly CCP Games) announces partnership with DeepMind

On the same day, May 6th, the company behind EVE Online announced four things:

  • Regained independence from its parent company Pearl Abyss;
  • Renamed to Fenris Creations;
  • Completed a $120 million transaction;
  • As part of this independence, Google acquired a minority stake in Fenris Creations and simultaneously initiated a research partnership with Google DeepMind.

Fenris Creations CEO Hilmar Veigar Pétursson stated in the announcement:

This transition does not involve layoffs or restructuring. The team, products, and development plans remain unchanged. EVE continues.

Looking at operational figures, this company came to the table with "real ammunition" for collaboration, not to sell assets for survival.

EVE Online's revenue in 2025 exceeded $70 million, with November setting a historical revenue record, and Q4 becoming the second-highest revenue quarter in the game's 20-year history.

Fenris Creations' independence means EVE now has a parent company that can autonomously decide on research collaborations, no longer constrained by the strategic goals of a larger game publishing company.

A box of a board game product published by Fenris in 1997. The name "Fenris" predates EVE Online by 6 years. Renaming to Fenris Creations is a look back, not a fresh start.

Why did DeepMind choose EVE?

A 23-Year "Artificial Society"

An AI Benchmark Difficult to Replicate

When many people hear "games + AI research," their first thought is of AlphaGo or AlphaStar. EVE is different from both.

Go and StarCraft share a common characteristic: a match has a beginning, an end, and clear win/lose rules.

AlphaGo's goal was to win a Go game. AlphaStar's goal was to win a StarCraft match. Both represent a "single-game intelligence" research paradigm. But EVE has no endgame.

EVE Online is famous for its "single-shard / single shared universe," where a vast number of players compete, trade, form alliances, and wage war in a persistent world over the long term.

Players here have built real economic systems, political alliances, military coalitions, trade routes, historical grudges, and warfare plans that span years.

Some campaigns take an entire year from preparation to conclusion. The rise and fall of some alliances are studied by later players as real history.

Hilmar stated in the announcement: "EVE is one of the few places where we can explore questions of intelligence in an environment that already operates like the real world."

Hassabis further explained that he has played games since childhood, his career started with designing AI simulation games, and his work on AlphaGo, AlphaStar, and SIMA has been deeply tied to games. EVE is the choice for the next stage:

I'm thrilled to partner with Fenris Creations to safely explore new game experiences and advance AI research within this player-created, uniquely complex universe.

Most AI benchmarks are like medical checkups. EVE is more like throwing AI into an "artificial society" that has been running for 23 years.

The Three Toughest Challenges for Agents

Happen to be Daily Life for EVE Players

The official announcement explicitly lists three research directions: long-horizon planning, memory, and continual learning.

These three directions are widely acknowledged as the three toughest challenges in current AI agent research.

If you know someone who has played EVE Online for over ten years, ask them to open their account and show you their friend list. You'll likely see dozens of groups and hundreds of names, with notes in the remarks field like "Debt owed from the 2018 Delve campaign," "Traitor within Goonswarm, do not cooperate," "This guy is a spy, everyone in the corp knows."

This isn't a context window; it's cross-session long-term memory spanning at least a decade.

EVE players navigate the memory challenge every day. The continual learning challenge is the same.

In January 2014, the B-R5RB battle lasted about 21 hours, involving over 7,500 characters, the destruction of 75 Titans, with losses equivalent to roughly $300,000 in real-world currency. The trigger for the entire battle was a sovereignty bill that failed to auto-pay.

After this battle, the entire game's fleet tactics were rewritten. Alliance fleet compositions and tactical systems for years after revolved around post-battle analysis and iteration. Updates were made monthly, with every failure broken down into actionable strategic updates.

As for long-horizon planning, the standard time unit for EVE alliance warfare isn't hours; it's months. From preparation to execution, a cross-regional war involves shipbuilding, logistics, diplomacy, infiltration, and counter-espionage, with hundreds of players spontaneously collaborating without any task manager to advance a common goal over months.

This collaborative system evolved organically from the players over 23 years.

The three hardest challenges recognized in current AI agent evaluation happen to be the daily life of EVE players.

Twenty-three years of player-driven evolution in EVE have produced an environment that is always changing, always complex, with no shortcuts. This level of complexity cannot be synthetically created in a lab.

DeepMind's SIMA 2, released in November 2025, has evolved from "executing instructions" to "understanding goals, reasoning about processes, and learning while playing."

From a research question perspective, the EVE project shares the same "games as a training ground for agents" path as SIMA 2. The difference is that the venue has been swapped for a real universe that has been running for 23 years.

In-game battle scene from EVE Online. These large-scale, player-organized battles, often lasting for hours, are the core reason DeepMind chose EVE as a research environment for long-horizon planning and continual learning.

DeepMind is Entering an Offline Sandbox

Not the Live Player Universe

DeepMind's collaboration method with Fenris is more conservative than one might imagine. DeepMind does not have direct access to the live player servers.

DeepMind officially stated in the announcement: Initial research will be conducted on an offline version of EVE Online, using local servers in a controlled environment to test and evaluate models, without connecting to EVE Online's live operational servers.

On one hand, the offline version means DeepMind will not consume live player PvP data or disrupt the actual server economy, avoiding any privacy and compliance complexities.

On the other hand, the offline version of EVE can still retain the complex rule systems, ship and economic mechanics, star system structure, and other core design elements.

DeepMind is getting a "complex world pressure-tested by players for 23 years" as the examination hall where its agents must survive.

From Atari to EVE

Where This Path Leads

Looking back at DeepMind's choice of training grounds over the past decade, there's a clear evolutionary line.

2013 to 2015: Atari was the starting point. DQN put agents into games like *Breakout* and *Space Invaders* with clear levels and closed rules. It tested reaction and value estimation.

2016 to 2017: AlphaGo and AlphaZero. Go has neat rules, a huge but closed action space. It tested search and long-chain reasoning.

2019: AlphaStar entered *StarCraft II*. The first entry into a real-time, imperfect-information, multi-threaded博弈 environment. It tested decision-making under partial observability.

2024: SIMA aimed to be a generalist agent across multiple games. It tested transfer and generalization.

2025: SIMA 2 upgraded: not just executing instructions, but also conversing with users, reasoning about goals, and self-improving during gameplay.

DeepMind's SIMA 2, released in 2025, has evolved from "executing instructions" to "understanding goals, reasoning processes, and learning while playing."

Each generation of environment incorporates more aspects of the "real world" than the last: from closed rules to open rules, from perfect information to imperfect information, from single-game对抗 to cross-game migration.

However, these previous environments were still relatively closed, segmentable, and repeatable task fields. For example, Atari has fixed-rule arcade games; AlphaStar faced StarCraft matches that ended one by one; SIMA tested cross-game generalization in multiple 3D virtual environments.

The difference with EVE is that it is a persistent world that has been running long-term, driven by players, with continuously evolving economic and political structures.

It has been organically evolved over 23 years by real players in an open-ruled world: a complete player-driven economy (ISK price fluctuations comparable to real financial markets), political structures across alliances (diplomacy, espionage, ceasefires), and a whole warfare ecosystem from small skirmishes to 21-hour mega-battles.

The consensus within the field on agent evaluation is increasingly clear: running point task benchmarks hasn't produced anything new for a long time, but long-term memory, planning across weeks, and learning from failure still lack decent evaluation arenas.

Therefore, DeepMind's choice this time is: rather than creating another synthetic environment, step into an "artificial society" that has already been pressure-tested by human players for 23 years.

But a bigger question then emerges:

An AI agent that can persist, continually learn, and plan within EVE—what is still missing between it and an autonomous agent operating in the real world?

References:

https://x.com/GoogleDeepMind/status/2052011542707630461

https://www.ccpgames.com/news/2026/studio-behind-eve-online-goes-independent-rebrands-as-fenris-creations-enters-research-partnership-with-google-deepmind

https://deepmind.google/blog/sima-2-an-agent-that-plays-reasons-and-learns-with-you-in-virtual-3d-worlds/

This article is from the WeChat public account "新智元" (New Zhiyuan), author: ASI启示录 (ASI Revelation), editor: 元宇 (Yuanyu).

Trending Cryptos

Related Questions

QWhy did DeepMind choose EVE Online as a research environment for AI agents?

ADeepMind chose EVE Online because it provides a complex, player-driven, and persistent universe that has evolved over 23 years. This environment is a perfect safe sandbox for testing key challenges in AI research, specifically long-horizon planning, memory, and continual learning, which are difficult to replicate in standard, closed-ended AI benchmarks.

QWhat are the three main research challenges DeepMind aims to tackle in its EVE Online collaboration?

AThe three main research challenges are long-horizon planning, memory, and continual learning. These are considered among the hardest problems in current AI agent research, and they correspond to the everyday activities and adaptations of long-term EVE Online players.

QHow is the DeepMind and Fenris Creations research collaboration structured in terms of accessing the EVE Online game world?

AThe initial research will be conducted in an offline version of EVE Online on local servers. DeepMind will not connect to the live, operational game servers. This approach provides a controlled environment for testing and evaluation without impacting the active player economy or raising privacy and compliance issues.

QAccording to the article, how does the complexity of EVE Online's environment differ from previous DeepMind research platforms like Atari games or StarCraft?

AUnlike previous platforms like Atari (closed rules, single sessions) or StarCraft (individual matches with a clear end), EVE Online is a persistent, single-shared universe without a defined end. Its complexity is not just in its rules but in the player-driven, long-term evolution of its economy, politics, and warfare, which have developed organically over 23 years.

QWhat major corporate changes happened to the studio behind EVE Online alongside the announcement of the DeepMind partnership?

AThe studio (formerly CCP Games) became independent from its parent company Pearl Abyss, rebranded as Fenris Creations, completed a $120 million transaction, and Google acquired a minority stake in the new company as part of the deal, which also initiated the research partnership with Google DeepMind.

Related Reads

A Critical Review of Base Co-founder's "Letter of Self-Criticism"

Yesterday, Base co-founder Jesse Pollak published a self-critical post on X, reflecting on Base's successes and failures over the past two years. While his candor is commendable, the author argues his analysis still misses key points. Jesse admitted his primary mistake was heavily betting that onchain social platforms (like Farcaster) and creator tokens would drive mainstream crypto adoption. The author harshly criticizes this strategy, giving it a "0/100" score. They argue onchain social offers no superior user experience, and creator tokens, like Jesse's own $JESSE, fail to solve fundamental issues about value and creator adoption, unlike successful memecoins on other chains. The author is more forgiving regarding Base lagging in areas like perpetual DEXs and prediction markets, giving this an "80/100". They note strong competitors like Hyperliquid had unique advantages, and Base has found significant success in other narratives like AI and RWA. The author gives Jesse's overall reflection a "60/100" or passing grade. While Jesse now correctly sees multiple paths to adoption (stablecoins, payments, AI Agents) beyond just social, he, along with other industry leaders like Toly and Vitalik, still underestimates the powerful role of memecoins in driving mainstream awareness and adoption. The critique concludes that leaders often leverage memes for user acquisition but fail to genuinely embrace or thoughtfully develop the space, revealing a retreat from idealism and a disconnect from grassroots, viral growth mechanisms.

marsbit18m ago

A Critical Review of Base Co-founder's "Letter of Self-Criticism"

marsbit18m ago

When Traditional Finance Fails People in Crisis, Bitcoin Succeeds

When traditional finance fails to reach people in crisis, Bitcoin succeeds. This article highlights how fundraising platforms like GoFundMe face severe limitations in delivering aid to conflict zones like Gaza due to banking regulations, sanctions, and compliance rules. For instance, Sami Jamal Al-Shannat raised over £55,000 for his family but couldn't receive the funds directly; instead, money had to be routed through a beneficiary in another country, leading to disputes and loss of access. The piece explores how compliance requirements often force reliance on intermediaries, shifting risk and responsibility away from platforms and onto individuals. In contrast, Bitcoin and blockchain-based platforms like Geyser and Agora enable direct, peer-to-peer transactions, bypassing traditional financial bottlenecks. These platforms shift trust from centralized institutions to verifiers—local partners or trusted organizations—who validate projects, allowing donors to send funds directly to recipients' wallets. While not a panacea, Bitcoin offers a way to circumvent "transnational financial repression," where sanctions and AML rules inadvertently harm legitimate aid recipients, activists, and dissidents. The conclusion emphasizes that open payment networks and decentralized trust models represent a systemic shift, empowering beneficiaries and providing more resilient humanitarian fundraising in crisis situations. However, challenges around verification, accountability, and fraud prevention remain.

marsbit45m ago

When Traditional Finance Fails People in Crisis, Bitcoin Succeeds

marsbit45m ago

Bitcoin Shifts Towards Consolidation, Long-Term Holder Selling Pressure Significantly Eases

Bitcoin's Bottoming Process Shows Signs of Shifting Dynamics Bitcoin's bottom formation is ongoing, but key characteristics are changing. The capitulation selling by long-term holders (LTHs), a primary source of selling pressure this cycle, has begun to cool from its recent peak. Buyers successfully absorbed the selling at the June lows, and price is now recovering to challenge overhead resistance. The market is testing higher resistance levels. Bitcoin reacted more strongly to soft inflation data than major equity indices, signaling sellers may be exhausted and buyers are waiting for a catalyst. Its correlation with stocks is weakening while its inverse relationship with the USD is deepening, suggesting liquidity dynamics are now more influential than risk sentiment. On-chain, price sits between the network's Realized Price (a historical bear market floor) and the Short-Term Holder (STH) cost basis near $69k, a key resistance level where recent buyers break even. LTH profit-taking has largely dried up, and losses now dominate realized on-chain volume—a typical late bear market signal. Crucially, the pace of LTH capitulation has started to decline. Derivatives markets show bearish positions are being unwound, with put/call ratios falling and crash protection costs moderating. However, this derisking hasn't been accompanied by significant spot buying, a missing link for sustained recovery. US spot ETF outflows have slowed but not reversed. In conclusion, foundational elements for a bottom are forming: LTH selling is easing, low-point selling was absorbed, and the market is responding to positive macro cues. The next critical test is whether spot-driven buying can push price through and hold above the STH cost basis near $69k. The follow-through is not yet confirmed.

marsbit1h ago

Bitcoin Shifts Towards Consolidation, Long-Term Holder Selling Pressure Significantly Eases

marsbit1h ago

Defending Champions or New Kings? World Cup Final Sees All AIs Backing the Same Side

Will the 2026 World Cup final see Argentina successfully defend their title or a new champion crowned? AI models from various platforms have made their prediction. The final in Buenos Aires pits defending champions Argentina against a resilient Spanish side that has reached this stage with a record of exceptional defensive solidity, conceding only one goal in seven matches. Argentina's path was dramatically different, filled with late comebacks and narrow victories, including a semi-final win over England secured by late goals assisted by the 39-year-old Lionel Messi. A poignant subplot adds narrative weight: a nearly 20-year-old photo shows a young Messi bathing an infant Lamine Yamal, who is now a 19-year-old key player for Spain, symbolizing a potential passing of the torch. In the semi-finals, most AI models incorrectly predicted a French victory over Spain, with only Google's Gemini correctly picking Spain's advancement and also accurately forecasting Argentina's win over England. For the final, however, all six surveyed AIs—ChatGPT, Claude, Gemini, Grok, DeepSeek, and Qwen—unanimously predict a Spanish victory. Their reasoning centers on Spain's superior defense, midfield control, and better physical preparedness after a less strenuous knockout stage journey. While consensus favors Spain as champions, five of the six AIs believe the match will be tightly contested, predicting a draw (1-1 or 0-0) within regular time, with Spain's advantage potentially telling in extra time or even a penalty shootout. Only DeepSeek forecasts a clear Spanish victory within 90 minutes. The stage is set for a clash between Argentina's legendary fighting spirit and Spain's machine-like consistency, with artificial intelligence firmly backing the latter to lift the trophy.

Odaily星球日报1h ago

Defending Champions or New Kings? World Cup Final Sees All AIs Backing the Same Side

Odaily星球日报1h ago

Trading

Spot

Hot Articles

Discussions

Welcome to the HTX Community. Here, you can stay informed about the latest platform developments and gain access to professional market insights. Users' opinions on the price of AI (AI) are presented below.

活动图片