世界杯才踢几天，AI预测已经有模型封神，有模型翻车

Odaily星球日报Опубліковано о 2026-06-15Востаннє оновлено о 2026-06-15

Анотація

世界杯期间，AI预测模型成为预测市场的新兴参考工具。首日比赛，阿里千问成功预测墨西哥2:0胜南非，并提示南非红牌风险，随后又命中韩国2:1逆转捷克，引发关注。微软Copilot对完整赛程进行了预测，成功押中墨西哥、韩国及巴西被摩洛哥逼平的具体比分，但也出现多次误判，尤其在冷门比赛如澳大利亚胜土耳其、日本平荷兰等场次中表现不佳。 ChatGPT在单场比赛分析中展现出完整逻辑，如准确预测揭幕战比分并给出合理理由，但其完整赛程预测更偏向纸面强队，对爆冷赛事敏感度不足。其他模型如Gemini、Grok和Claude在测试中表现各异，预测结果存在差异。总体而言，目前AI模型在世界杯预测中已有亮眼表现，可作为辅助参考，但远非绝对准确。其稳定性、对冷门的识别能力仍有待更多比赛检验。后续将持续追踪各模型预测与实际赛果的对比。

原创 | Odaily 星球日报(@OdailyChina)

作者 | Asher(@Asher_ 0210)

本届世界杯,最热闹的地方不只在球场上。

随着世界杯相关预测事件热度升温,越来越多用户开始用真金白银参与交易。谁能赢、几比几、会不会爆冷、有没有红牌、哪名球员能进球,这些原本属于球迷赛前闲聊的话题,如今被拆成了一个个可以交易的预测事件。

而当预测变成交易,用户需要的就不只是情绪和直觉:赔率变化、球队状态、伤病信息、历史交锋、市场情绪,都会成为交易前的参考。在这一过程中,AI 模型开始被频繁拉进世界杯预测场景里。

千问、ChatGPT、Gemini、Claude、DeepSeek、Qwen 以及 Copilot 等大模型,不仅能回答“哪支球队更可能赢”,还能给出比分判断、爆冷可能、红牌风险、关键球员表现和比赛走势分析。对于预测市场参与者来说,AI 的赛前推演,正在成为赔率、新闻、球队数据和市场情绪之外的另一层参考。

不过,预测最终仍要回到比赛本身。

随着世界杯正式开赛,前几场比赛结果已经陆续出炉。那些赛前被用户拿来辅助判断的 AI 分析,也终于有了可以对照的答案:比分有没有押中,爆冷有没有提前看到,红牌、绝杀、比赛走势这些细节,又有多少真正被模型捕捉到了。

最先出圈的,竟是千问

世界杯首日最有节目效果的,无疑是千问。

揭幕战墨西哥对南非,千问赛前给出的预测是墨西哥 2:0 南非。比赛结束后,比分真的定格在 2:0。更有看点的是,全场一共出现三张红牌,也和千问赛前提到的“南非防守动作过大、可能早早陷入少打一人”的风险判断基本吻合。

如果只是判断墨西哥取胜,这并不算太意外。作为东道主之一,墨西哥本身就更被看好。但千问这次踩中的是更具体的比赛细节:2:0 的比分、南非的红牌风险,以及比赛中后段被逐渐拉开的节奏。

紧接着,韩国对捷克这场,千问又给出了韩国 2:1 的判断。

这场比赛赛前并不算好猜。捷克有身体对抗,有定位球威胁,也有欧洲球队一贯的大赛经验。比赛过程也确实没有一边倒,捷克先取得领先,韩国随后扳平,比赛一度长时间僵在 1:1。直到最后阶段,韩国打进制胜球,比分最终变成 2:1。

这一下,千问的预测就有了更强的“剧本感”。胜负判断可以靠纸面实力,比分预测可以有运气成分,但红牌、逆转、最后阶段制胜这些过程细节,才真正让人觉得“有点东西”。首日两场之后,千问先把 AI 预测世界杯的关注度拉了起来。

Copilot:有神来一笔,也有明显翻车

赛前,USA Today 曾让 Copilot 预测了本届世界杯全部 104 场比赛。从目前已经结束的比赛来看,这份预测既有高光,也有明显失手。

其中,有三场比赛的预测最亮眼。

揭幕战墨西哥对南非,Copilot 给出的预测是墨西哥 2:0,最终比分正好命中。韩国对捷克,它预测韩国 2:1,同样与赛果一致。到了巴西对摩洛哥,Copilot 又给出 1:1 的判断,结果巴西真的被摩洛哥逼平。

尤其是巴西 1:1 摩洛哥这场,含金量不低。巴西毕竟是传统豪门,阵容和关注度都在第一梯队。摩洛哥虽然上届世界杯打进四强,但面对巴西,赛前直接预测双方打平,并不是一个特别安全的选择。结果比赛踢完,巴西没有拿下开门红,摩洛哥也延续了自己在大赛中的韧性,Copilot 这场预测确实是“神来一笔”。

但 Copilot 的问题也很快暴露出来。

它预测加拿大 2:1 战胜波黑,结果双方踢成 1:1;预测瑞士 1:0 小胜卡塔尔,结果瑞士同样被逼平;预测美国 2:0 巴拉圭,方向虽然对了,但实际比分是 4:1,进攻强度被明显低估。

更明显的翻车,出现在几场爆冷和强队受阻的比赛里。

土耳其对澳大利亚,Copilot 预测土耳其 2:1 取胜,结果澳大利亚 2:0 爆冷赢球。厄瓜多尔对科特迪瓦,它预测厄瓜多尔 2:1,结果科特迪瓦 1:0 拿下。荷兰对日本,它预测荷兰 2:1,结果日本两度追平,最终双方 2:2 战平。瑞典对突尼斯,它预测 1:1,结果瑞典直接踢出 5:1。

Copilot 能押中墨西哥、韩国、巴西这几场具体比分,说明并不是只会顺着热门队给答案。但澳大利亚击败土耳其、卡塔尔逼平瑞士、日本逼平荷兰这些比赛,也暴露出它对冷门和平局的判断仍然偏保守。

ChatGPT:分析很完整,但冷门抓得不够准

相比 Copilot 的完整赛程预测,ChatGPT 更像是一个“赛前分析型选手”。

在揭幕战预测中,ChatGPT 预测墨西哥 2:0 南非,最终比分命中。它给出的理由也比较完整,包括墨西哥的主场优势、近期状态、南非进攻乏力,以及墨西哥城高海拔和主场氛围等因素。这次预测中,ChatGPT 不只是给了结果,背后的判断逻辑也和比赛结果对上了。

但到了对世界杯完整赛程预测里,ChatGPT 的稳定性就没那么强。虽然它命中了墨西哥 2:0 南非和巴西 1:1 摩洛哥,也看对了苏格兰、德国、瑞典等几场比赛的胜负方向。但在韩国 2:1 捷克、卡塔尔 1:1 瑞士、澳大利亚 2:0 土耳其、日本 2:2 荷兰这些比赛上,ChatGPT 的判断都预测了纸面实力更强的队伍。比如瑞士应该赢卡塔尔,土耳其应该赢澳大利亚,荷兰应该小胜日本。

ChatGPT 不是没有预测能力,它能把球队实力、主场环境、近期状态拆得很清楚,也能在部分比赛里命中比分。但从目前结果看,它更擅长解释“为什么热门队更合理”,而不是提前识别哪些比赛可能偏离热门剧本。

Gemini、Grok、Claude:同一场比赛,不同模型写出不同剧本

除了千问、Copilot 和 ChatGPT,还有一些社媒用户把同一场比赛喂给多个模型做赛前预测。

以揭幕战墨西哥对南非为例,有博主同时测试了 ChatGPT、Gemini、Grok 和 Claude 四款 AI 模型进行赛前预测。结果显示,ChatGPT 和 Gemini 都给出了墨西哥 2:0 南非的预测,最终比分正好命中;Grok 预测墨西哥 2:1,Claude 预测墨西哥 3:1,虽然都看对了墨西哥取胜,但没有押中具体比分。

这次揭幕战的预测,不同模型给出了三种不同的“剧本”。ChatGPT Go 和 Gemini Pro 更接近实际比赛:墨西哥占优,南非进攻乏力,最终被零封。Grok 更像是给了一个相对开放的比分,认为南非会有反击收获。Claude Sonnet 则把墨西哥的进攻预期拉得更高,给出了 3:1 这种更大开大合的结果。

小结

由于目前可回溯的 AI 预测样本仍然有限,现阶段还不能直接判断哪个模型最“懂球”。

但只看已经结束的几场比赛,差异已经开始显现。千问目前最有记忆点,首日连续命中墨西哥 2:0 南非、韩国 2:1 捷克,还踩中了红牌风险和比赛走势,属于小样本里的高光表现。不过,后续能否持续命中,还需要更多比赛验证。

Copilot 和 ChatGPT,两者都有命中具体比分的高光,但也都暴露出一个共同问题——面对澳大利亚击败土耳其、卡塔尔逼平瑞士、日本战平荷兰这类偏离纸面实力的比赛,判断仍然不够敏感。

至于 Gemini、Grok、Claude 等模型,目前公开样本更多集中在单场或社媒对照,参考价值有,但还不适合直接下排名。

AI 已经可以成为世界杯预测市场用户的一层参考,但还远不是标准答案。接下来,Odaily星球日报也会继续收集各模型赛前预测,并随着比赛推进持续回看:哪些模型只是开局运气好,哪些模型真的能在更多场次里经得起赛果检验。

Пов'язані питання

Q在文章提到的AI模型中，哪个模型在世界杯首日的预测中表现最为突出，并具体说明了哪些细节？

A在世界杯首日的预测中，千问的表现最为突出。它成功预测了墨西哥2:0战胜南非的比分，并提到了南非可能因防守动作过大而吃到红牌的风险，这都与实际比赛情况吻合。此外，它还准确预测了韩国2:1战胜捷克的比分和比赛过程。

QCopilot在哪些比赛的预测中表现亮眼，又在哪些比赛中出现了明显的翻车？

ACopilot在墨西哥2:0南非、韩国2:1捷克和巴西1:1摩洛哥这几场比赛的预测中表现亮眼，准确命中了比分。然而，它在加拿大对波黑（预测2:1，实际1:1）、瑞士对卡塔尔（预测1:0，实际1:1）以及土耳其对澳大利亚（预测土耳其2:1胜，实际澳大利亚2:0胜）等比赛中出现了明显的预测失误。

Q根据文章描述，ChatGPT在世界杯预测中表现出了什么特点？

AChatGPT在世界杯预测中表现出了“赛前分析型选手”的特点。它不仅能给出预测结果，还能提供相对完整的分析逻辑，例如在预测墨西哥2:0南非时，提到了主场优势、近期状态和高海拔等因素。但文章指出，它在判断可能偏离纸面实力的比赛（如冷门或平局）时，表现不够敏感，更倾向于支持热门队伍。

Q文章中提到有博主测试了多个AI模型对同一场比赛（墨西哥对南非）的预测，结果如何？

A有博主同时测试了ChatGPT、Gemini、Grok和Claude四款AI模型对墨西哥对南非揭幕战的预测。结果是：ChatGPT和Gemini都准确预测了墨西哥2:0获胜；Grok预测墨西哥2:1获胜；Claude预测墨西哥3:1获胜。后两个模型虽然判断对了胜负，但没有命中具体比分。

Q文章作者对目前AI模型在世界杯预测中的总体表现做出了怎样的评价和展望？

A文章作者认为，由于目前可回溯的预测样本仍然有限，尚不能直接判断哪个模型最“懂球”。AI可以作为预测市场用户的一层参考，但远非标准答案。作者指出，不同的模型在部分场次有高光表现，但也暴露出对冷门比赛判断不够敏感等问题。文章最后表示，将继续收集各模型的赛前预测，并随着比赛推进检验其长期表现。

Пов'язані матеріали

Microsoft Identifies New Crypto Malware Targeting Wallet Addresses and Private Keys

In February 2026, Microsoft identified a new crypto clipper malware, dubbed Trojan/CryptoBandits.A, targeting Windows systems. The malware spreads via malicious shortcut files on USB drives and operates without a traditional installer or control servers by leveraging Windows Script Host and ActiveX to deploy a Tor proxy. Once active, it runs two modules: one for spreading and another for stealing information. The malware continuously monitors the clipboard for 12 or 24-word recovery phrases, Bitcoin/Ethereum private keys, and wallet addresses. When a user copies a wallet address, the malware silently swaps it with one controlled by attackers to divert funds. It also captures screenshots to gather information on wallet balances and user activity, sending data through Tor connections. Additional capabilities include remote code execution and persistence via scheduled tasks. Microsoft advises disabling auto-run features, restricting script interpreters and executable shortcuts from USB drives, and monitoring for suspicious activities like JavaScript execution, localhost:9050 proxy use, PowerShell screenshot capture, and clipboard monitoring.

TheNewsCrypto16 хв тому

Microsoft Identifies New Crypto Malware Targeting Wallet Addresses and Private Keys

TheNewsCrypto16 хв тому

No Sales Team, $20 Million in Revenue: How Did AI Employee Viktor Win Over 30,000 Companies?

The AI employee Viktor, developed by a team with DeepMind background, has achieved $20 million in annual revenue without a traditional sales team, serving over 30,000 companies. Its core innovation lies in positioning itself as a "Tier 3 AI Coworker" capable of "end-to-end execution and delivery of results," moving beyond the "draft and wait for human completion" model of typical AI assistants. Users can simply mention Viktor in Slack or Microsoft Teams using natural language commands, and it autonomously performs tasks like pulling sales data from a CRM, generating reports, or even cross-tool operations like creating board meeting PPTs by aggregating data from six different sources. Key to its growth is a pure Product-Led Growth (PLG) model, eliminating complex implementation cycles and per-seat licensing. Instead, it charges based on task credits or consumption, lowering the trial barrier with a $100 free credit offer and no credit card required. This enabled viral, bottom-up adoption within organizations. Viktor's interaction paradigm removes the barrier of prompt engineering, allowing non-technical employees to delegate complex workflows seamlessly. It also features proactive, automated task execution (e.g., overnight bookkeeping, scheduled reports) based on triggers, effectively embedding AI as an automated "process layer" within business operations. However, its expansion into Microsoft Teams—a platform with 320 million users—highlights challenges. Large enterprises require stringent IT compliance, security reviews (e.g., SOC 2), and governance, potentially hindering the frictionless, user-driven adoption that succeeded in Slack. Additionally, the "black box" nature of its autonomous decision-making raises concerns about operational risks, data integrity, and the need for robust audit logs and permission controls. Balancing efficiency gains with security and trust remains a critical hurdle for Viktor and similar AI agents aiming to become core enterprise infrastructure.

marsbit54 хв тому

No Sales Team, $20 Million in Revenue: How Did AI Employee Viktor Win Over 30,000 Companies?

marsbit54 хв тому

Interview with CoreWeave Co-founders: AI Demand Seems to 'Intensify' Every Day

An Interview with CoreWeave Executives: AI Demand Seems to 'Intensify' Every Day In an interview, CoreWeave executives highlight a structural shift in AI infrastructure demand. While GPU availability remains crucial, the primary bottlenecks are evolving to include powered data center shells, skilled labor (like electricians), and complex supply chain execution. They note that AI demand, particularly for agentic AI and reasoning models, continues to intensify daily, accelerating since Q1 2024. This demand is driving a need for more balanced infrastructure. CoreWeave is redesigning data centers to allocate more space for storage and CPUs alongside GPUs, with significant interest in Nvidia's upcoming Vera CPUs. The company, serving top AI labs and hyperscalers, emphasizes its client-driven model, building precisely to customer specifications. CoreWeave attributes its competitive edge to proven execution, performance, and a mature platform for AI deployment. Pricing is structured to pass component cost increases (e.g., for HBM memory) to customers, protecting margins. Looking ahead, they anticipate Vera Rubin platform deployments to begin meaningfully in late 2025, with a major ramp throughout 2027, mirroring the Blackwell (GB) series rollout pattern. The competition is shifting from merely acquiring chips to holistic engineering and delivery capability.

marsbit1 год тому

Interview with CoreWeave Co-founders: AI Demand Seems to 'Intensify' Every Day

marsbit1 год тому

Manus Buyback Plan Emerges: Chinese Investors Plan to Repurchase Equity with $2 Billion, Path to Hong Kong IPO Becomes Clearer

According to a report by The Information, early Chinese investors of Manus, including Tencent, Sequoia Capital China, and ZhenFund, are planning to repurchase the company from Meta for $2 billion—the same price Meta paid in its acquisition last December. This move is a direct response to the Chinese government's prohibition of the foreign acquisition in April. As part of the repurchase plan, Manus is considering establishing a Sino-foreign joint venture within China. This structure is seen as a way to ensure regulatory compliance for its Chinese investors and to pave the way for a future IPO in Hong Kong. Notably, U.S. investor Benchmark will not participate in the buyback, which will concentrate ownership even more among Chinese capital. Since its acquisition by Meta, Manus's business has grown rapidly, with its annualized revenue run rate reportedly increasing four-to-fivefold to $400-$500 million in roughly six months. This strong growth underpins the investors' willingness to repurchase at the original price. Financially, the forced unwinding of the deal may benefit the early investors, allowing them to regain equity at a cost far below the company's current implied valuation, with the added prospect of an independent future listing. However, specific terms of the repurchase, including funding proportions and the joint venture's equity structure, are still under negotiation. This "repurchase-joint venture-Hong Kong IPO" approach could serve as a reference model for other Chinese AI startups navigating cross-border M&A regulations.

marsbit1 год тому

Manus Buyback Plan Emerges: Chinese Investors Plan to Repurchase Equity with $2 Billion, Path to Hong Kong IPO Becomes Clearer

marsbit1 год тому

STRC Loses Peg by 11%, Can Strategy's Perpetual Motion Machine Keep Running?

The article discusses the significant and concerning depegging of MicroStrategy's (MSTR) preferred stock, STRC. Designed to trade near its $100 target par value, STRC has recently fallen sharply, reaching a low of $83.26 and closing at $88.59, representing an over 11% discount. STRC is a core component of MicroStrategy's financial strategy. As a perpetual preferred stock, it allows the company to raise capital through an "at-the-market" (ATM) issuance program without diluting common shareholders (MSTR). This capital is primarily used to purchase Bitcoin, creating a "capital flywheel": issuing STRC → raising cash → buying BTC → increasing net assets → supporting STRC's value. The flywheel's operation depends on STRC maintaining its $100 price. To enforce this, MicroStrategy employs a dynamic dividend mechanism, recently raising the rate to 11.5% and increasing payout frequency. However, this has failed to halt the depegging, indicating market concerns extend beyond yield. Analysts cite two main reasons. First, technical factors like forced liquidations from leveraged arbitrage trades may have exacerbated the sell-off. Second, and more fundamentally, is waning confidence in MicroStrategy's financial resilience. A JPMorgan report highlighted the company's limited cash relative to its ~$1.7 billion annual dividend obligation, raising liquidity concerns. While MicroStrategy counters that its massive Bitcoin holdings provide decades of coverage, this argument relies on the potential need to sell BTC—a departure from its long-standing "never sell" narrative. The company's recent sale of a small amount of Bitcoin for "testing," despite being framed as minor, has intensified these fears. The persistent depegging threatens to cripple MicroStrategy's primary funding channel. If STRC remains discounted, the company's ability to fund further Bitcoin purchases weakens. Should cash reserves dwindle while financing is constrained, the market may increasingly price in the risk of MicroStrategy becoming a forced seller of Bitcoin to meet obligations. This shift from a major marginal buyer to a potential seller could pose significant downside risk to the broader Bitcoin market.

链捕手1 год тому

STRC Loses Peg by 11%, Can Strategy's Perpetual Motion Machine Keep Running?

链捕手1 год тому

Торгівля

Спот

Ф'ючерси