OpenAI Solves 10 Mathematical Problems, Fable 'Replicates' 5 in 24 Hours

marsbitPublicado em 2026-08-05Última atualização em 2026-08-05

Resumo

OpenAI and Anthropic engaged in a rapid, high-stakes competition at the cutting edge of mathematics this weekend. On August 1, OpenAI researcher Sébastien Bubeck announced that their next-generation model Astra had autonomously solved 10 longstanding, open mathematical problems, providing Lean proofs and solution breakdowns. The problems, untouched for years, are considered significant; one result on non-sofic groups is deemed worthy of a top mathematics journal. The estimated marginal cost for these solutions was under $2,000. Within 24 hours, Anthropic researcher Levent Alpöge responded, stating he had independently used the publicly available model Fable to solve 5 of the 10 problems (#4-8), under clean conditions without internet access and with safeguards against data leakage. This dramatically shortens the "shelf life" of a mathematical discovery, shifting priority from years to potentially a day. The event is seen less as simple benchmarking and more as a form of peer review, testing the reliability and independent reproducibility of AI-generated proofs. The episode raises critical questions about validation in the age of AI. As these models can now produce complex proofs at low marginal cost, their outputs are often beyond public comprehension. The true challenge shifts from generating proofs to verifying, understanding, and judging their significance—a task that remains a deeply human and expert-driven endeavor. The ability to critically evaluate AI's mathematica...

10 mathematical problems, reversed in 24 hours.

This weekend, two AI giants, OpenAI and Anthropic, clashed head-on at the cutting edge of mathematics.

First, on August 1, OpenAI researcher Sébastien Bubeck announced that their unreleased next-generation flagship model, Astra, had proven 10 cutting-edge mathematical results in one go, also releasing 10 Lean certificates and 10 detailed solution explanations per problem.

Don't underestimate these 10 problems.

Their main conclusions had seen no progress for at least a decade, many had been dormant for decades, making them genuine open problems.

While not the deepest unsolved mysteries in mathematics, each one is a tough nut to crack.

Elliot Glazer, head of FrontierMath at Epoch AI, stated plainly: Just the paper on non-sophisticated groups alone would certainly make it into Annals of Mathematics, a top-tier journal recognized by the mathematical community.

The release of these 10 difficult problems sent shockwaves through the AI community, with some immediately declaring "the mathematical singularity has arrived."

Anthropic responded swiftly.

In less than 24 hours, researcher Levent Alpöge replied under Bubeck's post: "I completed half using Fable."

He then added that he had solved items 4, 5, 6, 7, and 8 from OpenAI's list, a total of 5 items.

The experimental conditions, Levent claimed, were clean: fully autonomous, generic prompts, no internet throughout, and a layer of protection was specifically added to prevent OpenAI's solutions from leaking into the context.

Of course, Levent hasn't yet released the complete proof details from Fable, stating he will upload them.

But if confirmed, the significance of this event changes:

The first-mover advantage OpenAI secured with its unreleased model Astra lasted only 24 hours. And the model catching up, Fable, is publicly available from Anthropic, accessible to everyone.

A "Mathematical First" Now Has a 24-Hour Shelf Life

User Chubby reposted: "Publicly available Fable replicated half of Astra's achievements."

A commenter joked: Looking at it this way, did OpenAI only secure a first announcement?

In the past, priority disputes over mathematical results typically played out over years.

Who thought of it first, who proved it first, who published first, separated by lengthy processes of submission, review, and revision.

Now, Astra hasn't even been formally released, yet the results it announced were half-matched by a public model in just 24 hours.

The value of being first remains, but its shelf life has drastically shortened.

Many netizens are already shouting "mathematical singularity," and the two AI giants are locked in a tight race.

Behind the $2000, There's an Even Pricier Bill

OpenAI also threw out a number: the cost of finding these 10 solutions was less than $2000.

According to researcher Noam Brown, this corresponds to the tokens required to search for these solutions, converted using the Sol API price.

In other words, this is just the marginal inference cost at the moment the 10 solutions were successfully produced.

Beyond that, there's the training and development of Astra itself, the problems attempted but not solved, the human effort to compile the arguments into a 249-page paper, the cost of formalizing each proof into Lean, and finally the cost of external mathematician review.

Noam himself admitted they had attempted other major problems without success, failing to crack any of the Millennium Prize Problems.

Nevertheless, $2000 still brings the marginal cost of an AI solving a cutting-edge problem down to rock bottom.

Some even joked whether OpenAI would announce 10 more mathematical breakthroughs tomorrow, or wait until next week and release 100 at once.

From "Benchmark Chasing" to "Mutual Verification"

Fable redoing the same batch of problems as Astra is more interesting than ordinary benchmark chasing.

Benchmarking tests known answers: the questions and standard answers are laid out, and you compare scores.

But two models, under conditions of no internet and leak protection, independently arriving at the same cutting-edge conclusion, tests something entirely different: how reliable is this result, can it be independently reproduced.

This is more like peer review.

Levent has currently only "made a statement" on X, not releasing the PDFs of those 5 proofs, prompts, complete run logs, token costs, or Lean certificates.

Netizen Haider questioned: Even if it's true, why use Fable to replicate Astra's work instead of directly tackling new open problems?

In fact, Fable isn't just riding Astra's coattails; it also solves new problems.

In July, Levent used Fable to produce a three-dimensional counterexample to the Jacobian Conjecture, and someone later wrote a 7-page algebraic verification material confirming that the explicit mapping did indeed refute the conjecture in three and higher dimensions.

For Achievements Most Can't Understand, Who Does the Accepting?

How important these 10 achievements are is very difficult for the vast majority to judge.

Ethan Mollick hit the nail on the head: For almost everyone on Earth, this is not just beyond ability, but beyond comprehension. We can only trust professional mathematicians to tell us how impressive it is.

With code, images, and chat, ordinary people can still personally experience where AI excels.

But the existence of non-sophisticated groups, counterexamples to Connes' Rigidity Conjecture, quantum parallel repetition theorems—these are understood by only a tiny minority of experts.

As AI capabilities advance towards the scientific frontier, what the public must rely on is not personal experience, but a long chain of trust. This state of "not even knowing how to be amazed" is happening simultaneously in more and more fields.

Mathematician Thomas Bloom cautioned that these achievements use mathematical theories accumulated over centuries and formal systems painstakingly built by mathematicians; simply describing it as "AI replacing mathematicians" is dishonest.

His criterion is: Does this proof teach us something new about the problem?

So the real question arises: When AI can produce cutting-edge mathematical proofs cheaply and in bulk, what should we use to measure its achievements?

The answer might not lie in its output end, but in the acceptance end: whether a proof can be read, verified, and have its significance judged ultimately determines its true value.

The scarcest ability is shifting from "producing proofs" to "understanding and accepting proofs."

When AI can prove ten difficult problems overnight, the real test is just beginning: As proofs are manufactured faster and faster, can our verification keep up?

References:

https://x.com/haider1/status/2083908869915611554

https://x.com/polynoamial/status/2083470822258467194

https://x.com/kimmonismus/status/2083950641978679363 https://x.com/Dr_Singularity/status/2083653287463571742

This article is from the WeChat public account "New Zhiyuan," author: ASI Apocalypse

Perguntas relacionadas

QWhat did OpenAI's unreleased model Astra reportedly achieve in the field of mathematics?

AOpenAI's unreleased model Astra reportedly proved 10 previously unsolved, open mathematical problems and provided Lean certificates and solution analyses for each.

QWhat significant action did Anthropic's model Fable take in response to OpenAI's announcement, and under what conditions?

AWithin 24 hours, Anthropic's publicly available model Fable reportedly proved 5 of the same 10 problems independently, using a general prompt and no internet access to prevent contamination from OpenAI's solutions.

QAccording to the article, what does the $2000 cost figure from OpenAI represent regarding the math proofs?

AThe $2000 cost figure represents the marginal inference cost of generating the 10 successful proofs using OpenAI's API, but does not include the costs of training the model, failed attempts, human labor for formatting, or external mathematical review.

QWhy does the article suggest that Fable's reproduction of Astra's work is more significant than standard benchmark competition?

ABecause it involves independently arriving at the same novel, unsolved conclusions under controlled conditions, which serves as a form of peer review to verify the reliability and reproducibility of the results, rather than just testing known answers.

QWhat core challenge does the article identify regarding the validation of AI-generated mathematical proofs?

AThe core challenge is that the ability to understand, verify, and assess the significance of highly complex, frontier mathematical proofs is extremely scarce. As AI produces proofs faster, the ability to validate them—not just generate them—becomes the true bottleneck and measure of value.

Leituras Relacionadas

Stunning, Musk Joins Forces with Huang Jen-hsun to Deploy the Stellar Brain! AI Computing Center Launched into Space

SpaceX, in its first public earnings report, has announced a partnership with NVIDIA to launch Starmind, an orbital AI computing network. This initiative aims to move AI's computational foundation into space to solve the critical challenges of energy consumption and heat dissipation faced by terrestrial data centers. The core concept involves deploying data centers equipped with NVIDIA's advanced GPUs (like the upcoming Rubin) into sun-synchronous orbit. Key advantages include near-continuous solar power for energy, the vacuum of space for highly efficient radiative cooling, and the absence of terrestrial constraints like land use and grid capacity. The first computational node, the Starmind AI1 satellite, is described as a massive vehicle with a 75-meter wingspan, generating 210 kW of solar power to run a 150 kW computing payload—equivalent to a full NVIDIA GB300 supercomputer rack. These "compute satellites" will connect to users via SpaceX's existing Starlink constellation using high-speed laser links, promising low-latency access from anywhere on Earth. SpaceX's ability to realize this vision is underpinned by its Starship rocket, designed for low-cost, high-frequency launches to transport thousands of satellites, and massive planned manufacturing facilities like the Gigasat factory and the Terafab chip plant. Elon Musk framed the project within the context of the Kardashev Scale, proposing Starmind as humanity's first substantive step toward becoming a Type II civilization capable of harnessing a star's energy. The ultimate goal is to deploy one terawatt of computational power in space, with even more ambitious plans involving lunar-based manufacturing and launch systems. The report positions Starmind not merely as a commercial AI service, but as a foundational infrastructure project to expand human intelligence and capability beyond Earth, leveraging SpaceX's integrated strengths in rockets, satellites, and global connectivity.

marsbitHá 32m

Stunning, Musk Joins Forces with Huang Jen-hsun to Deploy the Stellar Brain! AI Computing Center Launched into Space

marsbitHá 32m

Despite the sell-off, Goldman Sachs remains bullish on Samsung and SK Hynix. Here's why.

Goldman Sachs maintains "Buy" ratings on Samsung Electronics and SK Hynix despite recent stock declines. Its bullish view centers on three core arguments. Firstly, it expects HBM (High Bandwidth Memory) pricing to re-establish a premium over conventional DRAM by 2027, with a projected blended ASP of around $2.9/Gb. This is driven by tight supply-demand dynamics, increasing manufacturing complexity for newer HBM generations, and the need to restore its historical price premium. Secondly, the volatility of the memory cycle is expected to moderate due to widespread adoption of 3-5 year Long-Term Agreements (LTAs) with key customers. These contracts, covering a significant portion of planned capacity, feature mechanisms like price floors, prepayments, and penalties, reducing supplier risk and improving earnings visibility. Thirdly, inventory levels remain low at key suppliers and major customers, providing a buffer against a sharp downturn. Furthermore, robust demand from enterprise SSDs for AI servers is seen offsetting weakness in consumer segments like smartphones and PCs, preventing NAND markets from slipping into oversupply in the near term. While risks exist—such as potential weaker AI demand or aggressive capacity expansion—Goldman Sachs believes the combination of HBM repricing, LTAs, and low inventory underpins a more stable earnings outlook for the leading Korean memory makers.

marsbitHá 47m

Despite the sell-off, Goldman Sachs remains bullish on Samsung and SK Hynix. Here's why.

marsbitHá 47m

Korean Cryptocurrency Market Platform Kimpga Integrates with Global Web3 Data Platform RootData, Providing Popularity and Growth Scores for Token Projects

South Korean cryptocurrency market data platform Kimpga (kimpga.com) has announced a data partnership with the global Web3 project data platform RootData. As a result, Kimpga has integrated comprehensive token project information directly into its platform services. The integration introduces a new "RD Score" column within Kimpga's market listings. Clicking the score opens a detailed data window that displays the "RD Popularity Index," the "RD Growth Index," and key fundamental project data such as leading investors, core team members, and project category tags. For example, for Solana (SOL), users can view its indices and identify major backers like Polychain Capital, a16z, and Multicoin Capital directly from the interface, with links to full RootData project pages for deeper research. Jin Jae-yong, a representative of Kimpga, stated that the partnership allows users to quickly assess a project's fundamental strength, including its development team and investor lineup, without leaving the market data screen. This addresses a key challenge for investors who previously had to rely solely on price and premium data. RootData is a widely used global platform providing Web3 project details on funding history, team composition, and investor portfolios for due diligence and investment analysis. Kimpga, which boasts 10 million cumulative users, is a leading Korean crypto market platform providing real-time data and price premiums from over 18 domestic and international exchanges. This collaboration with RootData follows Kimpga's previous partnership with the token rating platform APYWA, further enriching its data offering to support user investment decisions.

marsbitHá 48m

Korean Cryptocurrency Market Platform Kimpga Integrates with Global Web3 Data Platform RootData, Providing Popularity and Growth Scores for Token Projects

marsbitHá 48m

He Was the Hero of Shanghai's Two Major Industries, Yet Passed Away Quietly in Regret

He was a pivotal figure in the development of two major Shanghai industries—semiconductors and commercial aircraft—yet passed away in quiet regret. Jiang Shangzhou, son of a veteran revolutionary and a Swiss-educated technocrat, became Deputy Director of the Shanghai Economic Commission in 1997. Tasked with identifying strategic industries, he championed integrated circuits (ICs) despite a climate of skepticism. Countering a national plan for just two 8-inch chip production lines, he boldly declared Shanghai would build ten within five years. He masterminded the Zhangjiang microelectronics zone, personally recruited top talent like Morris Chang, and helped establish SMIC. Through innovative financing that circumvented Western embargoes, SMIC grew rapidly, with Shanghai exceeding its target by building 18 lines. Simultaneously, while battling lung cancer, Jiang led the push for China's large commercial aircraft program as head of a state major projects panel. He argued it was a strategic imperative that would drive advancements across multiple industries. His relentless advocacy culminated in the 2006 state approval that paved the way for the C919. Jiang's career was marked by foresight often deemed too超前. Earlier postings in Sanya and Yangpu saw his ideas for tourism and economic reform initially rejected, only to be validated years later. In 2009, he became Chairman of the embattled SMIC. Despite his deteriorating health, he steered the company until his death in 2011, leaving a final insight for the chip industry: China need not master every process but must lead in key areas. He did not live to see the C919 fly or SMIC's full recovery, but his foundational work made Shanghai a national leader in both semiconductors and aerospace. Colleagues remembered him as a visionary whose impact shaped the three decades that followed.

marsbitHá 1h

He Was the Hero of Shanghai's Two Major Industries, Yet Passed Away Quietly in Regret

marsbitHá 1h

Trading

Spot
活动图片