Топовые ИИ-модели не осилили видеоигры девяностых

cryptonews.ruPublicado a 2025-03-21Actualizado a 2025-04-21

Даже самые продвинутые ИИ-модели не способны эффективно играть в классический шутер от первого лица Doom. К такому выводу пришли эксперты после проверки нейросетей в новом бенчмарке VideoGameBench.

Claude can play Pokemon, but can it play DOOM?

With a simple agent, we let VLMs play it, and found Sonnet 3.7 to get the furthest, finding the blue room!

Our VideoGameBench (twenty games from the 90s) and agent are open source so you can try it yourself now —> 🧵 pic.twitter.com/vl9NNZPBHY

— Alex Zhang (@a1zhang) April 17, 2025

Тест призван проверить способность современных нейросетей играть и побеждать в 20 популярных видеоиграх. Использовать они могут только информацию с экрана.

«Современные модели VLM с трудом справляются с видеоиграми из-за высокой задержки вывода. Когда агент делает снимок экрана и запрашивает VLM о том, какое действие ему следует предпринять, к моменту получения ответа состояние игры значительно меняется, и действие уже неактуально», — отметили исследователи.

Для теста использовались классические игры из 1990 годов из-за простых визуальных эффектов и различных стилей ввода вроде мыши, клавиатуры и игрового контроллера. Такой подход позволяет проверить у модели пространственное мышление и «зрение».

VideoGameBench разработан ученым и ИИ-исследователем Алексом Чжаном. В бенчмарк входят Warcraft II, Age of Empires, Prince of Persia и другие игры.

Список игр из бенчмарка VideoGameBench. Данные: сайт vgbench.

Sonnet 3.7 справилась с Doom лучше остальных — нейросеть нашла синюю комнату.

Исследователи подчеркнули, что задержка реакции — главная проблема в шутерах от первого лица. В быстро меняющейся обстановке враг может переместиться или даже добраться до игрока раньше его реакции на происходящее.

Помимо проблем с пониманием игрового окружения, модели также не могли выполнить основные действия.

«Мы часто наблюдали случаи, когда агент не мог понять, как его действия вроде движения вправо будут отображаться на экране. Самой распространенной ошибкой среди всех протестированных нами пограничных моделей оказалась неспособность надежно управлять мышью в таких играх, как Civilization и Warcraft II, где очень важны точные и частые движения», — отметили эксперты.

Также модели не всегда понимают игровые механики, когда нет прямой инструкции о необходимых действиях.

Напомним, в феврале ИИ-стартап Anthropic представил свою «самую интеллектуальную модель» Claude 3.7 Sonnet, которая прошла игру Pokemon.

Lecturas Relacionadas

Generating Profits for Seven Consecutive Quarters, Emerging Markets Carry Trade Outperforms Everything

For the seventh consecutive quarter, dollar-funded emerging market carry trades have delivered positive returns, marking the longest winning streak since 2008. According to Bloomberg's index, this strategy has gained approximately 22% since late 2024, outperforming U.S. Treasuries, emerging market sovereign, and corporate dollar debt. The core of the trade involves borrowing low-interest currencies like the U.S. dollar, euro, or yen to invest in high-yielding emerging market assets, such as Turkish lira bonds offering over 40% returns. Returns were amplified by favorable currency moves, with the dollar weakening against most emerging market currencies and other traditional funding currencies. For instance, the trade gained 48% on the Colombian peso in the past year. A key test came in August 2024 with a historic joint U.S.-Japan currency intervention, which caused only a modest 1% dip in the carry trade risk premium as investors shifted funding from the yen to the euro and Swiss franc. Looking ahead, the primary risk is the timing of Federal Reserve policy changes. While persistent inflation allows the Fed to hold rates, a rapid rise in long-term U.S. yields could threaten the trade. Another concern is crowding, as massive inflows increase vulnerability to a sudden reversal. High interest rates in regions like Latin America and Eastern Europe, supported by external factors like Middle East tensions and energy prices, continue to sustain the opportunity. Major investors remain engaged, favoring currencies like the Mexican peso, South African rand, and Turkish lira.

marsbitHace 6 min(s)

Generating Profits for Seven Consecutive Quarters, Emerging Markets Carry Trade Outperforms Everything

marsbitHace 6 min(s)

Unpacking the Truth Behind On-chain Assets: Leverage, Liquidity, and Risk

The article analyzes the concept of "real-world asset" (RWA) tokenization, arguing that while tokenizing assets on-chain is a useful step, it is far from transformative on its own. The author compares it to placing a barcode on a shipping container—it enables identification but does not build the necessary market infrastructure. The core argument is that true value emerges not from tokenization, but from integrating these tokens into DeFi systems where they can be valued, financed, hedged, traded, and liquidated under stress. Key challenges identified include: 1. **Multiple Time Clocks**: A fundamental tension exists between blockchain's 24/7 settlement and the slower, business-hour-dependent processes of traditional markets, custody, and redemption. This "duration mismatch" can create dangerous liquidity gaps during crises. 2. **Liquidity Misconceptions**: True liquidity is not measured by Total Value Locked (TVL) or trading pairs, but by the ability to exit a position within a required timeframe at an acceptable price. It requires analyzing multiple exit paths and stress-testing scenarios. 3. **Leverage and Risk**: Leverage unlocks economic utility (e.g., using tokenized assets as collateral) but also introduces fragility. Risk models must account for more than asset volatility, incorporating factors like legal enforceability, oracle freshness, and market structure. Paradoxically, a "safer" asset like tokenized Treasury bonds could require a higher collateral discount than ETH due to slower, less-proven liquidation mechanisms. 4. **A Risk Graph**: RWA risk should be modeled as a network of interconnected dependencies (e.g., issuers, custodians, oracles, stablecoin pools), not a single score. Failures can propagate through this graph, turning operational issues into systemic liquidity crises. The article states that tokenized government bonds are merely an entry point, while more complex frontiers like computing power and energy assets present greater challenges and opportunities. It also examines the interplay and risks between tokenized stocks and perpetual futures contracts. The conclusion is that the future lies not in "tokenizing everything," but in building robust market layers where tokenized rights become resilient financial primitives within a programmable capital system. The token is just the barcode; the market is the machine.

marsbitHace 31 min(s)

Unpacking the Truth Behind On-chain Assets: Leverage, Liquidity, and Risk

marsbitHace 31 min(s)

The End of Mathematics: 40 Top Mathematicians Gather at Secret OpenAI Meeting

In August 2026, OpenAI hosted a closed-door summit with approximately 40 leading mathematicians, including recent Fields Medalist Jacob Tsimerman and OpenAI researcher Sébastien Bubeck. The meeting, spurred by a series of recent AI breakthroughs in mathematics, grappled with the potential existential threat AI poses to the field. The backdrop includes several high-profile AI achievements: OpenAI models disproving long-standing conjectures like the unit distance problem, generating 10 new mathematical discoveries, and Anthropic's Claude aiding in constructing complex multidimensional objects. These results often bypass traditional academic pipelines, appearing directly on social media. The mathematical community is divided. Over 3000 researchers signed the "Leiden Statement," advocating for responsible AI use and verification. Others resist AI entirely to preserve human-centric mathematics. A key concern is AI's current inability to *explain* proofs, particularly the difficult steps, which is central to mathematical understanding. At the summit, Bubeck outlined four potential futures: mathematics becoming like collaborative software engineering, a compute-driven field like physics, a curatorial exercise where humans interpret AI output, or a mass transition of mathematicians into AI safety. While emphasizing that "mathematics only makes sense when mathematicians learn from it," no consensus was reached. The event highlights a profound moment of reflection. As AI demonstrates increasing competence in solving complex problems, mathematicians are forced to question their future role and the very meaning of their discipline.

marsbitHace 36 min(s)

The End of Mathematics: 40 Top Mathematicians Gather at Secret OpenAI Meeting

marsbitHace 36 min(s)

Trading

Spot
活动图片