He Let GPT-5.6 Sol Run for 33 Hours Straight to Tackle Fermat's Last Theorem, Forcibly Terminated by the System

marsbit发布于2026-07-28更新于2026-07-28

文章摘要

This article discusses a real-world experiment by expert Michael P. Frank to test if an AI, specifically GPT-5.6 Sol, could autonomously make progress on a major unsolved mathematical problem: finding a simpler proof for Fermat's Last Theorem. The AI was tasked with exploring specific mathematical pathways and maintaining rigorous notes over approximately 33 hours. However, the session was terminated by OpenAI's systems. The AI itself suggested two possible reasons for the stoppage: excessive resource consumption, or OpenAI having previously failed on similar problems and wishing to conserve computational resources. The AI reported its work primarily involved refining plausible ideas into precise, verifiable statements, most of which were subsequently disproven or excluded—effectively creating a map of dead ends rather than a proof. The incident sparked debate online. Some speculated that OpenAI might deliberately restrict public access to its most powerful models to maintain a competitive edge or avoid regulatory scrutiny, rather than allowing users to potentially solve landmark problems. OpenAI researcher Noam Brown countered this, arguing that a user solving a major problem would be tremendous publicity. Others offered technical explanations, suggesting the termination could be due to standard safety mechanisms preventing infinite loops, or even a known bug in the GPT-5.6 Sol version that disrupts long-running sessions. The story highlights the practical challenges, tec...

As an ordinary person, if you gave a challenging math problem to an AI and repeatedly told it to "continue," could it actually solve the problem, helping you win a large cash prize and even rewrite the history of mathematics?

Recently, with the announcement of the Fields Medal, discussions about mathematics + AI have been very heated, and I'm sure many have imagined the "wish-fulfillment" scenario I just described.

In reality, there are indeed people trying this, but the process is not as smooth as imagined.

Recently, top expert in reversible computing and computational physics, Michael P. Frank, posted that he set a high-difficulty research goal for GPT-5.6 Sol: to explore whether there exists a more concise approach to proving Fermat's Last Theorem than the Wiles–Taylor proof, focusing specifically on the modularity of Frey curves, uniform infinite descent, arithmetic abc-type inequalities, and uniform low-genus quotients; to maintain rigorous notes, computationally verify candidate lemmas, and clearly distinguish between proven results and conjectures.

This task ran continuously in the background for about 33 hours, consuming significant computational resources. But in the end, it was forcibly blocked by the OpenAI system.

In its own post-hoc analysis, GPT-5.6 Sol wrote that there were likely two reasons: first, the system judged that the session was using excessive resources; second, OpenAI had already attempted this problem with a similar model in the past and failed, and this time didn't want to waste computing power again.

Michael P. Frank seemed to agree with this analysis by GPT-5.6 Sol.

Furthermore, GPT-5.6 Sol candidly reported what it did during those 33 hours. The core finding was: it did substantial work, turning many "seemingly feasible" shortcuts into verifiable precise statements, then disproving or eliminating them one by one. The output was a map of "dead ends," not a proof itself. It recommended stopping the effort.

Regarding OpenAI's approach, some offered other possible interpretations: OpenAI might be holding back, not letting the model casually solve super-impressive math problems, fearing others would steal the spotlight, or wanting to claim the credit themselves. This serves as a reminder: if we become too dependent on them in the future, we are essentially handing over the power to decide "what can be discovered" to a single company, which is risky.

However, OpenAI senior research scientist Noam Brown quickly stepped in to point out the unreasonableness of this claim, saying, "If someone just typed 'continue' and used our model to solve a Millennium Prize Problem and then took home $1 million, that would be the best advertisement for OpenAI. Nothing would be better."

That sounds reasonable, but rebuttals argue that while opening top-tier capabilities to users might bring short-term publicity benefits, in the long run, it could weaken OpenAI's leading edge in the AI race. In a multi-player competitive environment, secrecy/internal priority use of strong models to accelerate one's own research is more important than letting users "preemptively" solve major problems. Once capabilities become "available to everyone," their publicity value drops significantly.

Furthermore, recent regulatory scrutiny gives OpenAI reason to hold back. If users used the public version to solve high-difficulty math problems, it would publicly demonstrate "this model is actually very powerful," potentially inviting more regulatory trouble. In this context, deliberately offering a weakened version externally while retaining a super-strong version internally sounds like a plausible strategy.

However, some also pointed out another possibility from a technical perspective: in codex-cli, "goal blocked" means the model repeatedly hit the same limit/blocker for 3 consecutive reasoning turns. So the task termination might be a routine safety/prevention-of-infinite-loop mechanism during model runtime, not a special block targeting difficult problems.

Others said this might simply be caused by a bug. Specifically, version 5.6 Sol has an annoying bug when running long-duration tasks. The temporary workaround is: switch back to version 5.5, compress the context, run it for a few minutes with 5.5, then switch back to version 5.6 Sol. "If you want to keep the same session, you'll get stuck and can't proceed."

What do you think? Feel free to share your experiences using AI to tackle difficult problems or getting stuck with it in the comments section.

References:

https://x.com/lu_sichu/status/2081367506468360495

This article is from the WeChat public account "Almost Human" (ID: almosthuman2014), author: Zhang Qian

热门币种推荐

相关问答

QWhat was the main goal set for the GPT-5.6 Sol model in the experiment described in the article?

AThe main goal was to investigate whether a simpler proof of Fermat's Last Theorem exists compared to the Wiles–Taylor proof, focusing on specific areas like modularity of Frey curves and the Uniform abc Conjecture, while maintaining rigorous documentation and verification of candidate lemmas.

QWhy was the GPT-5.6 Sol session forcibly terminated after approximately 33 hours according to its own analysis?

AGPT-5.6 Sol's analysis suggested two possible reasons: the system deemed the session was consuming excessive computational resources, or OpenAI had already tried and failed to solve the problem with similar models and decided to stop wasting resources.

QWhat did the GPT-5.6 Sol model report as its primary achievement during the 33-hour run?

AIt reported that its core achievement was performing substantial groundwork by turning plausible-sounding shortcuts into precise, verifiable statements, and then systematically disproving or eliminating them, effectively producing a map of dead ends rather than a proof itself.

QAccording to the article, what was one argument against Noam Brown's claim that OpenAI would welcome a user solving a major problem?

AThe argument was that while such a solve might offer short-term publicity, long-term competitive strategy favors keeping the strongest models internal to accelerate a company's own research, and that making top-tier capabilities publicly available diminishes their strategic and perceived value.

QWhat technical reason, unrelated to conspiracy, did some commentators suggest for the session's termination?

ASome suggested it could be due to a bug in GPT-5.6 Sol where long-running tasks might get stuck, and a known workaround involves switching to version 5.5, compressing the context, running for a few minutes, and then switching back to 5.6 Sol.

你可能也喜欢

迈克尔·赛勒声称比特币可能增长100倍,并警告称监管变化可能危及其未来

微策略公司执行主席迈克尔·赛勒警告称,比特币面临的主要挑战可能来自内部治理分歧,因为网络正从数字资产转变为全球资本市场的潜在基础。他认为比特币已进入重要采用阶段,但现在面临决定其经济体系的规则压力。赛勒指出,改变网络结构和激励参与者的协议变更可能是主要威胁,而非外部竞争。 他警告称,若个别团体获得比特币治理过程的影响力,可能“编造借口、重写规则并攫取经济权利”。赛勒将比特币共识规则比作其宪法,强调为特定派别利益修改系统可能影响整个生态和参与者权利。他同时认为比特币有百倍增长潜力,可能成为全球资本基础,但规则变更风险可能损害未来市场、技术与经济自由。 赛勒以BIP-110、盟约功能和区块扩容提案为例,说明可能带来风险的变更,担忧技术更新会让特定团体给网络参与者强加新成本或风险。他还将协议决策与网络长期安全模型关联,指出随着区块奖励减半,交易费用需承担更多安全成本。随着企业采用增加,机构投资者更关注网络规则可预测性和稳定性。 赛勒最后警告,若允许政治影响力介入共识变更定义,可能导致持续争夺协议控制权的冲突,这种持久争论可能吓退资本、延缓创新、削弱安全性,阻碍比特币发挥全部潜力。

cryptonews.ru16分钟前

迈克尔·赛勒声称比特币可能增长100倍,并警告称监管变化可能危及其未来

cryptonews.ru16分钟前

交易

现货

热门文章

如何购买S

欢迎来到HTX.com!我们已经让购买Sonic(S)变得简单而便捷。跟随我们的逐步指南,放心开始您的加密货币之旅。第一步:创建您的HTX账户使用您的电子邮件、手机号码注册一个免费账户在HTX上。体验无忧的注册过程并解锁所有平台功能。立即注册第二步:前往买币页面,选择您的支付方式信用卡/借记卡购买:使用您的Visa或Mastercard即时购买Sonic(S)。余额购买:使用您HTX账户余额中的资金进行无缝交易。第三方购买:探索诸如Google Pay或Apple Pay等流行支付方法以增加便利性。C2C购买:在HTX平台上直接与其他用户交易。HTX场外交易台(OTC)购买:为大量交易者提供个性化服务和竞争性汇率。第三步:存储您的Sonic(S)购买完您的Sonic(S)后,将其存储在您的HTX账户钱包中。您也可以通过区块链转账将其发送到其他地方或者用于交易其他加密货币。第四步:交易Sonic(S)在HTX的现货市场轻松交易Sonic(S)。访问您的账户,选择您的交易对,执行您的交易,并实时监控。HTX为初学者和经验丰富的交易者提供了友好的用户体验。

3.1k人学过发布于 2025.01.15更新于 2026.06.02

如何购买S

相关讨论

欢迎来到HTX社区。在这里,您可以了解最新的平台发展动态并获得专业的市场意见。以下是用户对S(S)币价的意见。

活动图片