他让GPT-5.6 Sol连跑33小时攻关费马大定理,被系统强制终止

marsbitPublished on 2026-07-28Last updated on 2026-07-28

Abstract

近日,菲尔兹奖公布引发数学与AI结合的讨论。可逆计算专家Michael P. Frank尝试让GPT-5.6 Sol探究费马大定理是否存在比现有证明更简洁的途径,任务运行约33小时后被OpenAI系统强制终止。 GPT-5.6 Sol分析可能原因包括会话占用资源过多,或OpenAI此前已尝试类似问题但失败。模型在运行期间进行了扎实工作,将许多“听起来可行”的思路转化为精确陈述并逐一验证或排除,最终产出了一张“此路不通”的地图,而非证明本身。 对于终止原因,存在多种解读:有人认为OpenAI可能刻意限制模型解决重大难题,以保持自身竞争优势或避免监管风险;OpenAI研究员Noam Brown则反驳称,若用户能用模型解决千禧年难题并获得奖金,对OpenAI将是最好的宣传。也有观点指出,终止可能是技术性原因,如模型运行时触发了防卡死机制,或是版本bug导致。 这一事件引发关于AI能力边界、资源分配与科技公司策略的思考。

作为一个普通人,如果你把一道数学难题给到 AI,并一直让它「继续」,它有没有可能真的把这道题解出来,从而帮你赢得一大笔奖金,甚至改写数学发展进程?

最近,菲尔兹奖公布,大家对于数学 + AI 的讨论非常热烈,相信也有不少人想象过我刚刚描述的「爽文」剧情。

其实,现实里也真的有人在尝试,但过程并没有想象中顺利。

最近,可逆计算、计算物理学顶尖专家 Michael P. Frank 发文表示,他给 GPT-5.6 Sol 设定了一个高难度研究目标:探究费马大定理是否存在比怀尔斯 - 泰勒证明更简洁的途径,重点放在专门的 Frey 曲线模性、一致无限下降法、算术 abc 型不等式以及一致低亏格商上;保持严谨的笔记,通过计算验证候选引理,并清楚区分已证明的结果与推测。

这个任务在后台持续运行了大约 33 个小时,消耗了大量计算资源。但最终,它被 OpenAI 系统强制阻止了。

GPT-5.6 Sol 在自己给出的事后分析中写到,可能的原因有两个:一是系统判定该会话占用了过多资源;二是 OpenAI 之前已经用类似模型尝试过这个问题但失败了,这次不想再浪费算力。

Michael P. Frank 似乎赞同 GPT-5.6 Sol 的这一分析。

此外,GPT-5.6 Sol 还坦诚报告了这 33 小时所做的事情,核心是:做了不少扎实工作,把很多「听起来可行」的捷径变成了可验证的精确陈述,然后一一证伪或排除。产出是一张「此路不通」的地图,而不是证明本身。建议就此打住。

对于 OpenAI 的这种做法,还有人给出了其他可能的解读:OpenAI 可能在藏着掖着,不让模型随便解决超级牛的数学问题,怕被别人抢了风头,或者想自己拿去吹牛。这就给大家提了个醒:如果以后太依赖他们,就等于把「什么能被发现」的权力全交给他们一家了,这很危险。

但很快,OpenAI 高级研究科学家 Noam Brown 就站出来指出了这一说法的不合理性,他说:「要是有人只是输入一个 continue(继续),就用我们的模型解决了千禧年大奖难题,然后拿走 100 万美元奖金,这对 OpenAI 来说才是最好的广告,没有比这更好的了。」

听起来很有道理,但反驳者认为,开放顶尖能力给用户确实可能带来短期宣传收益,但长期看会削弱 OpenAI 在 AI 竞赛中的领先优势。在多玩家竞赛环境下,保密 / 内部优先使用强模型来加速自家研究,比让用户「抢先」解决大问题更重要。而一旦能力变得「人人可用」,宣传价值就会大幅下降。

此外,最近的监管风波也让 OpenAI 有理由保存实力。如果用户用公开版解决了高难度数学问题,就等于公开证明「这模型其实很强」,反而会引来监管麻烦。这种情况下,刻意对外提供弱化版模型,而内部保留超强版本听起来是个合理的做法。

不过,也有人从技术层面指出了另一种可能性:在 codex-cli 里,「goal blocked」的意思是模型在连续 3 个推理轮次(turns)里反复撞上同一个限制 / 阻挡点(blocker)。所以任务被组织可能是模型运行时的常规安全 / 防止卡死机制,而不是针对难题的特殊阻挡。

还有人说,这可能就是一个 bug 导致的。具体来说,5.6 版本的 Sol 在长时间运行任务时有个烦人的 bug,临时解决办法是:切换回 5.5 版本,压缩一下上下文,用 5.5 继续跑几分钟,然后再切回 5.6 版本的 Sol。「如果你想保留同一个会话(session),就会卡住出不来。」

你怎么看?欢迎在评论区分享用 AI 攻克难题或被它卡住的经历。

参考链接:

https://x.com/lu_sichu/status/2081367506468360495

本文来自微信公众号“机器之心”(ID:almosthuman2014),作者:张倩

Trending Cryptos

Related Questions

QMichael P. Frank 让 GPT-5.6 Sol 运行了多长时间来攻关费马大定理?

A大约33小时。

QGPT-5.6 Sol 在33小时的运行后,得出了什么核心成果?

A它做了不少扎实工作,将许多听起来可行的捷径转化为可验证的精确陈述,然后一一证伪或排除,产出是一张‘此路不通’的地图,而非证明本身,并建议就此打住。

QOpenAI 高级研究科学家 Noam Brown 针对‘OpenAI 刻意阻止用户解决难题’的说法是如何回应的?

A他认为,如果用户仅用‘continue’指令就解决了千禧年大奖难题并赢得奖金,这对 OpenAI 来说是最好的广告,没有比这更好的了。

Q文章中提到,关于任务被强制终止,有哪些可能的解读?

A主要有四种解读:一是系统判定会话占用了过多资源;二是OpenAI可能不想在已尝试失败的问题上再浪费算力;三是监管风波导致公司可能对外提供弱化版模型;四是可能是版本bug或模型运行时的常规安全机制导致的。

Q文章中提到的一个临时解决 GPT-5.6 Sol 长时间运行任务卡住的建议方法是什么?

A切换回5.5版本,压缩一下上下文,用5.5继续跑几分钟,然后再切回5.6版本的Sol。

Related Reads

UNI Doubles in Two Months Against the Trend: A 5-Year-Overdue Value Realization

Amidst a generally stagnant crypto market in June and July, UNI, the governance token of Uniswap, saw a significant surge, nearly doubling in price from around $2.3 to $4.6. This rally represents a delayed but significant value reassessment, triggered by the practical implementation of its long-debated "fee switch" mechanism. The key turning point was the on-chain execution of the UNIfication proposal in December 2025. It activated a protocol fee on select pools, directed Unichain sequencer revenue (net of costs) to a communal treasury, executed a one-time burn of 100 million UNI, and established a system where all protocol revenue flows into a "TokenJar" contract. This treasury has a single exit: purchasing and permanently burning UNI via a "Firepit" contract. Initially, the market reacted tepidly as the generated revenue and corresponding burn rate were modest. The narrative shifted dramatically in July 2025 with two major developments. First, the launch of Robinhood Chain, tailored for tokenized stocks, rapidly became a primary source of volume and fees for Uniswap, at one point contributing nearly half of its weekly fees. Second, governance votes successfully expanded the fee mechanism to v4 pools and initiated a temperature check for fees on Robinhood Chain. The activation of v4 fees caused the protocol's daily revenue earmarked for UNI burns to nearly triple. The core of UNI's recent price action is the transition from a pure governance token to a cash-flow asset with a permanent, protocol-funded buyer. Its effectiveness is amplified by UNI's mature and widely distributed supply, with no major impending unlocks to dilute the impact of the buybacks. The sustainability of this rally now hinges on whether the transaction volume, particularly on Robinhood Chain, persists after its initial gas subsidies expire, determining if this is a genuine value realization or a subsidy-fueled spike.

marsbit44m ago

UNI Doubles in Two Months Against the Trend: A 5-Year-Overdue Value Realization

marsbit44m ago

Breaking: Google Earth Urgently Pulls Back Nano Banana 2 Image Generation Feature!

Google Earth's newly launched "Create image" feature, powered by the Nano Banana 2 AI image generation model, was abruptly withdrawn shortly after its release due to being "played" by users. The feature allowed users to generate and overlay AI-created visuals directly onto real-world satellite and 3D maps in Google Earth. The tool enabled creative applications like historical recreations (e.g., visualizing ancient Pompeii), generating informational graphics for landmarks, and envisioning architectural projects or futuristic cityscapes on real terrain. It operated under "geospatial grounding," meaning the AI respected the underlying geography, topography, and perspective of the chosen map view. The model also integrated with Gemini to retrieve relevant factual information. However, upon release, users quickly tested its limits. A prominent example involved reimagining Philadelphia's historic Independence Hall as a post-apocalyptic ruin overrun by "happy" zombies, evil clowns, and giant alien mechs. This highlighted both the feature's playful potential and its risks regarding the generation of inappropriate or misleading content on realistic maps, leading to its swift temporary removal. Google stated it would re-release the feature after implementing "enhanced guardrails." Analysts note this move strategically leverages Google's vast proprietary geospatial data, positioning its AI not just for artistic generation but for spatially accurate world visualization—a unique advantage in the competitive AI image generation landscape.

marsbit2h ago

Breaking: Google Earth Urgently Pulls Back Nano Banana 2 Image Generation Feature!

marsbit2h ago

Altman Admits: Overestimated AI Snatching Jobs! Huang Renxun: The Unemployment Narrative Is Completely Backwards

Sam Altman has revised his earlier predictions about AI rapidly replacing jobs, admitting he overestimated the speed at which AI would eliminate entry-level white-collar roles. Speaking on the "Invest Like the Best" podcast, he stated that people do not truly want an AI CEO, as accountability and human connection remain critical. He found that individuals prefer interacting with people who can be held responsible for decisions. Similarly, NVIDIA's Jensen Huang argued that the narrative of AI destroying jobs is misguided. He distinguishes between tasks and jobs, noting that while AI can automate specific tasks, entire jobs—encompassing communication, judgment, coordination, and accountability—are not eliminated. He cited examples like radiologists and software engineers, where demand for these roles has increased as AI handles repetitive tasks, allowing for business expansion and the creation of more positions. Data from a University of Maryland and LinkUp study supports this, showing that U.S. job postings for new graduates have actually risen, countering the fear of vanishing entry-level roles. However, a significant shift is occurring: the traditional entry-level tasks that help newcomers gain experience are being automated, making initial career access more challenging. The key insight is that as AI takes over standardized tasks, the enduring value of human work shifts toward areas of responsibility, trust-building, and final decision-making—aspects that AI cannot replicate. The real "moat" for professionals lies in these irreplaceable human elements.

marsbit2h ago

Altman Admits: Overestimated AI Snatching Jobs! Huang Renxun: The Unemployment Narrative Is Completely Backwards

marsbit2h ago

Trading

Spot

Hot Articles

Discussions

Welcome to the HTX Community. Here, you can stay informed about the latest platform developments and gain access to professional market insights. Users' opinions on the price of S (S) are presented below.

活动图片