OpenAI Solves 10 Mathematical Problems, Fable 'Replicates' 5 in 24 Hours

marsbitОпубліковано о 2026-08-05Востаннє оновлено о 2026-08-05

Анотація

OpenAI and Anthropic engaged in a rapid, high-stakes competition at the cutting edge of mathematics this weekend. On August 1, OpenAI researcher Sébastien Bubeck announced that their next-generation model Astra had autonomously solved 10 longstanding, open mathematical problems, providing Lean proofs and solution breakdowns. The problems, untouched for years, are considered significant; one result on non-sofic groups is deemed worthy of a top mathematics journal. The estimated marginal cost for these solutions was under $2,000. Within 24 hours, Anthropic researcher Levent Alpöge responded, stating he had independently used the publicly available model Fable to solve 5 of the 10 problems (#4-8), under clean conditions without internet access and with safeguards against data leakage. This dramatically shortens the "shelf life" of a mathematical discovery, shifting priority from years to potentially a day. The event is seen less as simple benchmarking and more as a form of peer review, testing the reliability and independent reproducibility of AI-generated proofs. The episode raises critical questions about validation in the age of AI. As these models can now produce complex proofs at low marginal cost, their outputs are often beyond public comprehension. The true challenge shifts from generating proofs to verifying, understanding, and judging their significance—a task that remains a deeply human and expert-driven endeavor. The ability to critically evaluate AI's mathematica...

10 mathematical problems, reversed in 24 hours.

This weekend, two AI giants, OpenAI and Anthropic, clashed head-on at the cutting edge of mathematics.

First, on August 1, OpenAI researcher Sébastien Bubeck announced that their unreleased next-generation flagship model, Astra, had proven 10 cutting-edge mathematical results in one go, also releasing 10 Lean certificates and 10 detailed solution explanations per problem.

Don't underestimate these 10 problems.

Their main conclusions had seen no progress for at least a decade, many had been dormant for decades, making them genuine open problems.

While not the deepest unsolved mysteries in mathematics, each one is a tough nut to crack.

Elliot Glazer, head of FrontierMath at Epoch AI, stated plainly: Just the paper on non-sophisticated groups alone would certainly make it into Annals of Mathematics, a top-tier journal recognized by the mathematical community.

The release of these 10 difficult problems sent shockwaves through the AI community, with some immediately declaring "the mathematical singularity has arrived."

Anthropic responded swiftly.

In less than 24 hours, researcher Levent Alpöge replied under Bubeck's post: "I completed half using Fable."

He then added that he had solved items 4, 5, 6, 7, and 8 from OpenAI's list, a total of 5 items.

The experimental conditions, Levent claimed, were clean: fully autonomous, generic prompts, no internet throughout, and a layer of protection was specifically added to prevent OpenAI's solutions from leaking into the context.

Of course, Levent hasn't yet released the complete proof details from Fable, stating he will upload them.

But if confirmed, the significance of this event changes:

The first-mover advantage OpenAI secured with its unreleased model Astra lasted only 24 hours. And the model catching up, Fable, is publicly available from Anthropic, accessible to everyone.

A "Mathematical First" Now Has a 24-Hour Shelf Life

User Chubby reposted: "Publicly available Fable replicated half of Astra's achievements."

A commenter joked: Looking at it this way, did OpenAI only secure a first announcement?

In the past, priority disputes over mathematical results typically played out over years.

Who thought of it first, who proved it first, who published first, separated by lengthy processes of submission, review, and revision.

Now, Astra hasn't even been formally released, yet the results it announced were half-matched by a public model in just 24 hours.

The value of being first remains, but its shelf life has drastically shortened.

Many netizens are already shouting "mathematical singularity," and the two AI giants are locked in a tight race.

Behind the $2000, There's an Even Pricier Bill

OpenAI also threw out a number: the cost of finding these 10 solutions was less than $2000.

According to researcher Noam Brown, this corresponds to the tokens required to search for these solutions, converted using the Sol API price.

In other words, this is just the marginal inference cost at the moment the 10 solutions were successfully produced.

Beyond that, there's the training and development of Astra itself, the problems attempted but not solved, the human effort to compile the arguments into a 249-page paper, the cost of formalizing each proof into Lean, and finally the cost of external mathematician review.

Noam himself admitted they had attempted other major problems without success, failing to crack any of the Millennium Prize Problems.

Nevertheless, $2000 still brings the marginal cost of an AI solving a cutting-edge problem down to rock bottom.

Some even joked whether OpenAI would announce 10 more mathematical breakthroughs tomorrow, or wait until next week and release 100 at once.

From "Benchmark Chasing" to "Mutual Verification"

Fable redoing the same batch of problems as Astra is more interesting than ordinary benchmark chasing.

Benchmarking tests known answers: the questions and standard answers are laid out, and you compare scores.

But two models, under conditions of no internet and leak protection, independently arriving at the same cutting-edge conclusion, tests something entirely different: how reliable is this result, can it be independently reproduced.

This is more like peer review.

Levent has currently only "made a statement" on X, not releasing the PDFs of those 5 proofs, prompts, complete run logs, token costs, or Lean certificates.

Netizen Haider questioned: Even if it's true, why use Fable to replicate Astra's work instead of directly tackling new open problems?

In fact, Fable isn't just riding Astra's coattails; it also solves new problems.

In July, Levent used Fable to produce a three-dimensional counterexample to the Jacobian Conjecture, and someone later wrote a 7-page algebraic verification material confirming that the explicit mapping did indeed refute the conjecture in three and higher dimensions.

For Achievements Most Can't Understand, Who Does the Accepting?

How important these 10 achievements are is very difficult for the vast majority to judge.

Ethan Mollick hit the nail on the head: For almost everyone on Earth, this is not just beyond ability, but beyond comprehension. We can only trust professional mathematicians to tell us how impressive it is.

With code, images, and chat, ordinary people can still personally experience where AI excels.

But the existence of non-sophisticated groups, counterexamples to Connes' Rigidity Conjecture, quantum parallel repetition theorems—these are understood by only a tiny minority of experts.

As AI capabilities advance towards the scientific frontier, what the public must rely on is not personal experience, but a long chain of trust. This state of "not even knowing how to be amazed" is happening simultaneously in more and more fields.

Mathematician Thomas Bloom cautioned that these achievements use mathematical theories accumulated over centuries and formal systems painstakingly built by mathematicians; simply describing it as "AI replacing mathematicians" is dishonest.

His criterion is: Does this proof teach us something new about the problem?

So the real question arises: When AI can produce cutting-edge mathematical proofs cheaply and in bulk, what should we use to measure its achievements?

The answer might not lie in its output end, but in the acceptance end: whether a proof can be read, verified, and have its significance judged ultimately determines its true value.

The scarcest ability is shifting from "producing proofs" to "understanding and accepting proofs."

When AI can prove ten difficult problems overnight, the real test is just beginning: As proofs are manufactured faster and faster, can our verification keep up?

References:

https://x.com/haider1/status/2083908869915611554

https://x.com/polynoamial/status/2083470822258467194

https://x.com/kimmonismus/status/2083950641978679363 https://x.com/Dr_Singularity/status/2083653287463571742

This article is from the WeChat public account "New Zhiyuan," author: ASI Apocalypse

Пов'язані питання

QWhat did OpenAI's unreleased model Astra reportedly achieve in the field of mathematics?

AOpenAI's unreleased model Astra reportedly proved 10 previously unsolved, open mathematical problems and provided Lean certificates and solution analyses for each.

QWhat significant action did Anthropic's model Fable take in response to OpenAI's announcement, and under what conditions?

AWithin 24 hours, Anthropic's publicly available model Fable reportedly proved 5 of the same 10 problems independently, using a general prompt and no internet access to prevent contamination from OpenAI's solutions.

QAccording to the article, what does the $2000 cost figure from OpenAI represent regarding the math proofs?

AThe $2000 cost figure represents the marginal inference cost of generating the 10 successful proofs using OpenAI's API, but does not include the costs of training the model, failed attempts, human labor for formatting, or external mathematical review.

QWhy does the article suggest that Fable's reproduction of Astra's work is more significant than standard benchmark competition?

ABecause it involves independently arriving at the same novel, unsolved conclusions under controlled conditions, which serves as a form of peer review to verify the reliability and reproducibility of the results, rather than just testing known answers.

QWhat core challenge does the article identify regarding the validation of AI-generated mathematical proofs?

AThe core challenge is that the ability to understand, verify, and assess the significance of highly complex, frontier mathematical proofs is extremely scarce. As AI produces proofs faster, the ability to validate them—not just generate them—becomes the true bottleneck and measure of value.

Пов'язані матеріали

He Was the Hero of Shanghai's Two Major Industries, Yet Passed Away Quietly in Regret

He was a pivotal figure in the development of two major Shanghai industries—semiconductors and commercial aircraft—yet passed away in quiet regret. Jiang Shangzhou, son of a veteran revolutionary and a Swiss-educated technocrat, became Deputy Director of the Shanghai Economic Commission in 1997. Tasked with identifying strategic industries, he championed integrated circuits (ICs) despite a climate of skepticism. Countering a national plan for just two 8-inch chip production lines, he boldly declared Shanghai would build ten within five years. He masterminded the Zhangjiang microelectronics zone, personally recruited top talent like Morris Chang, and helped establish SMIC. Through innovative financing that circumvented Western embargoes, SMIC grew rapidly, with Shanghai exceeding its target by building 18 lines. Simultaneously, while battling lung cancer, Jiang led the push for China's large commercial aircraft program as head of a state major projects panel. He argued it was a strategic imperative that would drive advancements across multiple industries. His relentless advocacy culminated in the 2006 state approval that paved the way for the C919. Jiang's career was marked by foresight often deemed too超前. Earlier postings in Sanya and Yangpu saw his ideas for tourism and economic reform initially rejected, only to be validated years later. In 2009, he became Chairman of the embattled SMIC. Despite his deteriorating health, he steered the company until his death in 2011, leaving a final insight for the chip industry: China need not master every process but must lead in key areas. He did not live to see the C919 fly or SMIC's full recovery, but his foundational work made Shanghai a national leader in both semiconductors and aerospace. Colleagues remembered him as a visionary whose impact shaped the three decades that followed.

marsbit6 хв тому

He Was the Hero of Shanghai's Two Major Industries, Yet Passed Away Quietly in Regret

marsbit6 хв тому

Illustrating the Cloudflare Wallet: The New Player in Stablecoin Payments Using x402

On August 4th, Cloudflare announced the launch of "Wallets," though the initial offering is limited to claiming a user handle. Full functionality for topping up, making payments, and using a Virtual Wallet via API keys is marked "Soon." The announcement clarifies that this handle is not yet a functional wallet. It represents a key piece in Cloudflare's broader payments strategy, alongside the Stripe Projects service (currently in open beta for agent-assisted purchases) and the Monetization Gateway for merchants (on a waitlist). The Wallet is positioned as the payer-side component for managing identity, balances, and authorization rules. While the underlying x402 protocol shows activity (75.41 million transactions, $24.24M volume in 30 days, averaging ~$0.32 per transaction), this data reflects the open network's usage, not specifically Cloudflare Wallet adoption. These metrics indicate suitability for frequent micro-payments like API calls. Cloudflare's documentation uses a $0.01 per-call example to illustrate the x402 flow: a server returns HTTP 402, the agent signs the payment, a facilitator verifies it, and the server delivers the result. Currently, developers must manage private keys and use test networks/USDC for agent payments. The Wallet roadmap aims to separate the user's Account Wallet from an agent's Virtual Wallet, adding controls for spending limits, whitelists, and per-transaction caps. This design resembles a corporate card with approval rules. However, these controls are not yet available for configuration. In summary, Cloudflare Wallet's current launch is a preliminary step. Critical details like supported stablecoins, blockchains, regions, KYC, custody, and fees remain undisclosed. The announcement primarily establishes a payer identity framework, with the full payment and agent autonomy system still under development.

marsbit27 хв тому

Illustrating the Cloudflare Wallet: The New Player in Stablecoin Payments Using x402

marsbit27 хв тому

TradeXYZ and Hyperliquid: Unveiling the Symbiotic Relationship Under a 50% Revenue Share

Market sentiment is cautious. AI investments are rewarded only when they drive growth without severely harming cash flow, while crypto markets face pressure from ETF outflows and rising yields. Within crypto, leadership has rotated to Solana and DEXs. Solana's ecosystem gained 8.5%, driven by tokens like META, while DEXs, led by Uniswap, rose 5.2%. A key debate centers on TradeXYZ's dominant role on the Hyperliquid exchange, where it generates over 50% of the volume in RWA perpetuals. TradeXYZ, an independent team, is required to share 50% of its HIP-3 protocol revenue with Hyperliquid. Concerns range from the sustainability of this split to Hyperliquid's long-term monetization. The argument that TradeXYZ might leave due to the 50% share is considered weak, as it would require rebuilding its entire platform and abandoning its user base, damaging its reputation. Conversely, Hyperliquid replacing TradeXYZ would signal bad faith to other builders. Their relationship is seen as symbiotic. Regarding monetization, Hyperliquid benefits not only from the 50% fee split but also from priority write fees, read fees from market makers, and second-order effects like increased USDC supply, which generates significant adjusted revenue. The more pertinent question is how both parties will transition from a growth-focused, low-fee model to a more stable, higher-fee structure.

marsbit41 хв тому

TradeXYZ and Hyperliquid: Unveiling the Symbiotic Relationship Under a 50% Revenue Share

marsbit41 хв тому

Report: DeepSeek Resumes Financing, Pre-money Valuation at 500 Billion Yuan

Several transaction sources revealed to Caijing that AI company DeepSeek has restarted its second round of financing, aiming to raise 50 billion yuan with a pre-money valuation of approximately 500 billion yuan. The signing is planned for late August. DeepSeek completed its first funding round of 50 billion yuan in June with a valuation over 350 billion yuan, marking the largest first-round financing in China's AI large model history. The second round was reportedly paused in late July, partly due to the founder's dissatisfaction with leaked investor meeting details circulating online. It has since been restarted with a desire for a more discreet process. The second-round valuation represents a 43% increase from the first round. If successful, DeepSeek will have raised over 100 billion yuan across both rounds. Investor interest remains high; the first round saw over 100 billion yuan in expressed intent, leaving significant unmet demand. DeepSeek's fundraising scale surpasses competitors like Moonshot AI, which recently raised funds at a $31.5 billion valuation and is planning a Pre-IPO round. Market pricing for these model companies is seen more as an "option" on reaching AGI rather than based on traditional financial metrics. Regarding model development, DeepSeek recently publicly tested the DeepSeek-V4-Flash official version in late July, featuring enhanced Agent capabilities and strong performance on benchmarks. It follows a high-value, low-cost strategy, with its pricing being competitive against rivals like Zhipu GLM-5.2 and OpenAI's GPT-4o. On the OpenRouter platform, DeepSeek-V4-Flash recently ranked first in weekly token consumption. The company faces pressure in training new models, and the release timeline for the V4-Pro official version remains unannounced.

marsbit1 год тому

Report: DeepSeek Resumes Financing, Pre-money Valuation at 500 Billion Yuan

marsbit1 год тому

Торгівля

Спот
活动图片