The World Cup Has Only Just Begun, But AI Predictions Already Have Models Hailed as 'Godly' and Others Flipping Over

Odaily星球日报Publicado em 2026-06-15Última atualização em 2026-06-15

Resumo

After only a few days of the World Cup, AI models are being widely used for match predictions, with mixed early results. These models analyze details like scores, upsets, red cards, and key players, offering users in prediction markets an extra layer of analysis beyond odds and news. Qwen gained early attention for its remarkably accurate calls on the opening day, correctly predicting Mexico's 2-0 win over South Africa and Korea's 2-1 victory over the Czech Republic, while also highlighting red card risks and match flow. Copilot had its own highlights, accurately forecasting the Mexico 2-0 result, the Korea 2-1 win, and a surprising 1-1 draw between Brazil and Morocco. However, it also misjudged several matches, like predicting a Swiss win that ended in a draw with Qatar and missing Australia's upset over Turkey. ChatGPT provided detailed pre-match analysis and correctly called the Mexico 2-0 score, explaining factors like home-field advantage. Yet, it struggled to anticipate upsets, often siding with the stronger team on paper, as seen in its missed calls for the Australia-Turkey and Japan-Netherlands matches. Social media tests pitted models like Gemini, Grok, and Claude against each other for the same games, revealing different predictive "scripts" even for the same fixture. Overall, while AI models like Qwen and Copilot have shown promising, high-profile successes in early matches, their consistency and ability to predict genuine upsets remain in question. As the tour...

Original | Odaily Planet Daily (@OdailyChina)

Author | Asher(@Asher_ 0210)

The most lively place in this World Cup isn't just on the pitch.

As the heat around World Cup prediction events rises, more and more users are starting to participate in trading with real money. Who will win, what will the score be, will there be an upset, will there be a red card, which player will score—topics that originally belonged to pre-match fan chatter have now been broken down into tradable prediction events.

And when predictions become trades, what users need is more than just emotion and intuition: odds movements, team form, injury news, historical matchups, market sentiment, all become references before a trade. In this process, AI models are being frequently pulled into World Cup prediction scenarios.

Large models like Qwen, ChatGPT, Gemini, Claude, DeepSeek, Qwen, and Copilot can not only answer "which team is more likely to win," but also give score predictions, upset possibilities, red card risks, key player performance, and match flow analysis. For prediction market participants, AI's pre-match deduction is becoming another layer of reference beyond odds, news, team data, and market sentiment.

However, predictions ultimately have to return to the matches themselves.

As the World Cup officially kicks off and the results of the first few matches come out one after another, those AI analyses that users used to aid their judgment before the matches finally have answers to compare against: Was the score predicted, were upsets spotted in advance, how many details like red cards, last-minute winners, and match flow were truly captured by the models.

The First to Go Viral Was, Unexpectedly, Qwen

The most dramatic performance on the first day of the World Cup was undoubtedly Qwen's.

For the opening match between Mexico and South Africa, Qwen's pre-match prediction was Mexico 2:0 South Africa. After the match ended, the score was indeed 2:0. What's more interesting is that the total of three red cards in the match also basically matched Qwen's pre-match risk assessment of "South Africa's overly physical defending, possibly getting into an early one-man-down situation."

If it were just predicting a Mexico win, that wouldn't be too surprising. As one of the hosts, Mexico was already favored. But what Qwen hit this time were more specific match details: the 2:0 scoreline, South Africa's red card risk, and the gradually widening gap in the latter stages of the match.

Immediately after, for the South Korea vs. Czech Republic match, Qwen gave another prediction of South Korea 2:1.

This match wasn't easy to call pre-match. The Czechs have physicality, set-piece threats, and the usual big-tournament experience of European teams. The match process was indeed not one-sided—the Czechs took the lead first, South Korea equalized later, and the match was deadlocked at 1:1 for a long time. Until the final stage, South Korea scored the winning goal, and the final score became 2:1.

At this point, Qwen's prediction gained a stronger "scripted" feel. Judging the winner can rely on paper strength, score prediction can have luck involved, but it's process details like red cards, comebacks, and last-stage winners that truly make people feel "there's something to it." After two matches on the first day, Qwen first pulled up the attention on AI World Cup predictions.

Copilot: Moments of Genius, and Clear Flops

Before the tournament, USA Today had Copilot predict all 104 matches of this World Cup. Looking at the matches that have concluded so far, this prediction has both highlights and clear misses.

Three match predictions stood out the most.

For the opening match Mexico vs. South Africa, Copilot predicted Mexico 2:0, which hit the final score exactly. For South Korea vs. Czech Republic, it predicted South Korea 2:1, again matching the result. For Brazil vs. Morocco, Copilot gave a 1:1 prediction, and Brazil was indeed held to a draw by Morocco.

Especially the Brazil 1:1 Morocco match, the prediction had considerable merit. Brazil, after all, is a traditional powerhouse, with squad strength and attention in the top tier. Although Morocco reached the semi-finals in the last World Cup, directly predicting a draw against Brazil pre-match wasn't a particularly safe choice. After the match, Brazil didn't get a winning start, and Morocco continued its tournament resilience. Copilot's prediction for this match was indeed a "stroke of genius."

But Copilot's issues also quickly surfaced.

It predicted Canada would beat Bosnia and Herzegovina 2:1, but they drew 1:1; predicted Switzerland would narrowly beat Qatar 1:0, but Switzerland was also held to a draw; predicted the USA would beat Paraguay 2:0—the direction was correct, but the actual score was 4:1, significantly underestimating the offensive intensity.

More obvious flops appeared in several upset matches and matches where strong teams were held back.

For Turkey vs. Australia, Copilot predicted Turkey to win 2:1, but Australia pulled off a 2:0 upset win. For Ecuador vs. Ivory Coast, it predicted Ecuador 2:1, but Ivory Coast won 1:0. For Netherlands vs. Japan, it predicted Netherlands 2:1, but Japan equalized twice, resulting in a 2:2 draw. For Sweden vs. Tunisia, it predicted 1:1, but Sweden directly won 5:1.

That Copilot could nail the exact scores for Mexico, South Korea, and Brazil matches shows it's not just giving answers favoring the favorites. But matches like Australia beating Turkey, Qatar drawing with Switzerland, and Japan drawing with Netherlands also expose that its judgment on upsets and draws is still relatively conservative.

ChatGPT: Analysis Is Comprehensive, But Not Sharp Enough on Upsets

Compared to Copilot's full schedule prediction, ChatGPT is more of a "pre-match analysis player."

In the opening match prediction, ChatGPT predicted Mexico 2:0 South Africa, hitting the final score. The reasoning it gave was also quite comprehensive, including Mexico's home advantage, recent form, South Africa's lack of attacking power, and factors like Mexico City's high altitude and home atmosphere. In this prediction, ChatGPT didn't just give a result; the underlying judgment logic also aligned with the match outcome.

But when it comes to full World Cup schedule predictions, ChatGPT's stability isn't as strong. While it hit Mexico 2:0 South Africa and Brazil 1:1 Morocco, and correctly called the winner in several matches like Scotland, Germany, and Sweden, in matches like South Korea 2:1 Czech Republic, Qatar 1:1 Switzerland, Australia 2:0 Turkey, and Japan 2:2 Netherlands, ChatGPT's judgments all predicted the team with stronger paper strength. For example, Switzerland should beat Qatar, Turkey should beat Australia, the Netherlands should narrowly beat Japan.

ChatGPT isn't without predictive ability; it can break down team strength, home environment, and recent form clearly, and can hit the score in some matches. But based on current results, it's better at explaining "why the favorite is more reasonable" than identifying in advance which matches might deviate from the favorite's script.

Gemini, Grok, Claude: Different Models Write Different Scripts for the Same Match

Besides Qwen, Copilot, and ChatGPT, some social media users fed the same match to multiple models for pre-match predictions.

Taking the opening match Mexico vs. South Africa as an example, a blogger tested four AI models—ChatGPT, Gemini, Grok, and Claude—simultaneously for pre-match predictions. The results showed that ChatGPT and Gemini both predicted Mexico 2:0 South Africa, hitting the final score exactly; Grok predicted Mexico 2:1, Claude predicted Mexico 3:1; while both correctly saw Mexico winning, they didn't nail the exact score.

For this opening match prediction, different models gave three different "scripts." ChatGPT and Gemini Pro were closer to the actual match: Mexico dominant, South Africa lacking in attack, eventually kept scoreless. Grok seemed to give a more open scoreline, thinking South Africa would get a goal back from a counter. Claude Sonnet raised expectations for Mexico's attack higher, giving a more open 3:1 result.

Summary

Since the currently reviewable sample of AI predictions is still limited, we can't directly judge which model is the most "football-savvy" at this stage.

But just looking at the few matches that have concluded, differences are starting to show. Qwen currently has the most memorable performance, hitting Mexico 2:0 South Africa and South Korea 2:1 Czech Republic on the first day, and also stepping on red card risks and match flow, a highlight in a small sample. However, whether it can sustain accuracy needs more matches to verify.

Copilot and ChatGPT both have highlights of hitting exact scores, but also expose a common problem—their judgment still isn't sensitive enough when facing matches that deviate from paper strength, like Australia beating Turkey, Qatar drawing Switzerland, or Japan drawing the Netherlands.

As for models like Gemini, Grok, and Claude, currently public samples are more concentrated on single matches or social media comparisons, having reference value but not yet suitable for direct ranking.

AI can already serve as a layer of reference for World Cup prediction market users, but it's far from a standard answer. Next, Odaily Planet Daily will continue to collect pre-match predictions from various models and keep reviewing them as the tournament progresses: which models just had good luck at the start, and which models can truly withstand the test of results across more matches.

Perguntas relacionadas

QWhich AI model achieved a notable success in predicting the first two matches of the World Cup, including specific scores and details like red cards?

AThe AI model Qwen (千问) successfully predicted the exact scores of Mexico 2:0 South Africa and South Korea 2:1 Czech Republic in the first two matches. It also correctly identified the risk of red cards for South Africa.

QWhat were the major strengths and weaknesses of Copilot's World Cup match predictions according to the article?

ACopilot's strengths included accurately predicting the exact scores of Mexico 2:0 South Africa, South Korea 2:1 Czech Republic, and Brazil 1:1 Morocco. Its weaknesses were significant misses on several matches, such as incorrectly predicting wins for Turkey against Australia and for Switzerland against Qatar, showing a conservative bias against upsets and draws.

QHow did ChatGPT's approach to World Cup prediction differ from a model like Copilot?

AChatGPT functioned more as a 'pre-match analytical player,' providing detailed reasoning behind its predictions (e.g., considering Mexico's home advantage and altitude). While it accurately predicted some scores like Mexico 2:0, its full-tournament predictions showed less stability and a tendency to favor the stronger team on paper, missing several upsets and draws.

QWhat were the different predictions made by ChatGPT, Gemini, Grok, and Claude for the opening match between Mexico and South Africa?

AFor the opening match Mexico vs. South Africa, ChatGPT and Gemini predicted a 2:0 win for Mexico. Grok predicted a 2:1 win for Mexico, and Claude predicted a 3:1 win for Mexico. Only ChatGPT and Gemini predicted the exact final score of 2:0.

QWhat is the article's overall conclusion about the current reliability of AI models for World Cup predictions?

AThe article concludes that while AI models can provide an additional layer of reference for prediction market participants, they are far from being a standard answer. Their reliability varies, with some showing promising initial results but needing more matches for validation, and most models currently showing a lack of sensitivity in predicting upsets and unexpected draws.

Leituras Relacionadas

Xpeng and NIO Compete on Computing Power, Li Auto Shifts Architecture

On June 15, 2026, Li Auto unveiled details of its self-developed chip, Mahe M100, for its new L9 Livis model. CTO Xie Yan stated the goal was not just a faster chip, but a fundamentally different one, targeting the chip architecture itself. While competitors like NIO, Xpeng, and Huawei highlight TOPS (computing power) figures for their self-developed chips, Li Auto’s Mahe M100 focuses on redesigning the underlying architecture. It employs a "dynamic data flow architecture" to address memory bandwidth bottlenecks in large model inference, claiming up to 3x the effective computing power of Nvidia's Thor U for its specific workloads and a 40% reduction in latency. The chip's design was peer-reviewed and accepted at ISCA 2026. However, this performance is highly optimized for Li Auto's own VLA2.1 algorithm, meaning it may not generalize as well to other tasks. Li Auto aims to achieve full-stack in-house development with Mahe M100, covering chip, compiler, OS, AI algorithms, and domain controller—a level of vertical integration few competitors match. Beyond the chip, CEO Li Xiang introduced a new strategic narrative: the "embodied intelligent vehicle," defined as an integration of an EV, a professional driver, an AI computer, and a life assistant. This shifts competition from features like large screens to systemic AI capabilities. A key commitment was that Li Auto's Mahe VLA autonomous driving model will match Tesla's FSD V14 by Q4 2026, with specific OTA milestones set for July, September, and December. Financially, Li Auto faces pressure with declining revenue and vehicle gross margins since Q4 2025, while maintaining high R&D investment (approx. ¥12B in 2026, 50% AI-related). Its 2026 sales target is 550,000 vehicles, up from 406,000 in 2025. The new L9 Livis garnered over 10,000 pre-orders in two weeks. The effectiveness of these strategic moves—new products, OTAs, and the novel chip architecture—will begin to show in Q3 2026 financial results, with the year-end FSD V14 benchmark being the ultimate test.

marsbitHá 38m

Xpeng and NIO Compete on Computing Power, Li Auto Shifts Architecture

marsbitHá 38m

The Year of AI Applications: Saying 'Yes' While Ignoring Risks? A Comprehensive Open Source Log of Software Development's Journey

The Year of AI Applications: Blindly Saying "Yes" While Ignoring Risks? A Software Development Log Goes Fully Open Source. AI-generated code harbors risks hidden within seemingly correct programs, potentially leading to data leaks or asset loss. The open-source project "Narwhal AI Code Risks," from Peking University's Narwhal-Lab, compiles real-world cases, early warning signs, and typical risk pathways. Its goal is to help developers identify potential hazards early and avoid repeating past mistakes. In 2026, code is generated faster than ever but deployed with less scrutiny. The danger often lies not in glaring errors, but in code that appears normal—syntactically correct, passing all checks—yet introduces subtle but critical flaws like non-existent dependencies, excessive permissions, or exposed databases. A stark example is the Moonwell cbETH oracle incident. A configuration file error, where a cryptocurrency price was set to ~$1.12 instead of ~$2,200, slipped through 28 checks and a pull request signed by both AI (Claude, Copilot) and human developers. This "semantic deviation" resulted in a loss of $1.78 million. The risk is that AI can produce functionally valid code that is semantically wrong for the business context. As AI moves beyond simple code completion to modifying configurations, installing dependencies, and operating via autonomous agents, it traverses longer, less traceable paths within software engineering, blurring traditional boundaries and oversight points. The Narwhal AI Code Risks project structures information into three layers: `/cases` for documented real-world incidents, `/inferred` for early warning signals, and `/scenarios` for clear, generalized risk patterns not yet tied to specific events. This aims to create a lasting, public record to prevent collective amnesia about past AI-coding pitfalls. Risks are categorized into seven areas: Software Supply Chain (e.g., recommending fake packages), Code-Level Vulnerabilities (e.g., reintroducing path traversal bugs), Cloud & Infrastructure Misconfiguration (e.g., overly permissive settings), Agent Risks (from autonomous tool execution), Vertical Domain Risks (e.g., in finance, healthcare), Intellectual Property & Compliance issues, and Human Factors (like over-reliance on AI output). The project's core value is transforming isolated incidents into reusable knowledge—a foundational resource for developers to spot similar issues, for security researchers to build upon, for toolmakers to create detection rules, and for the community to contribute new findings. As AI integration accelerates, this open-source "logbook" serves as a crucial navigational aid, charting past errors to help future projects steer clear of the same traps.

marsbitHá 38m

The Year of AI Applications: Saying 'Yes' While Ignoring Risks? A Comprehensive Open Source Log of Software Development's Journey

marsbitHá 38m

The Foundation of SpaceX's Trillion-Dollar Valuation: Who is Dividing Up Musk's Annual Tens of Billions in Capital Expenditure?

SpaceX's trillion-dollar valuation is built on its three core businesses: Starlink (profitable, 60% of revenue), rockets (driving down launch costs), and AI (a major investment area). This creates a financial cycle: Starlink funds rocket development, which enables low-cost launches for AI hardware, generating future revenue. This cycle fuels annual capital expenditures of tens of billions, flowing to a vast supply chain. Suppliers are categorized by their replaceability. The first group includes irreplaceable players like NVIDIA (GPU/CUDA ecosystem), Eutelsat (critical radio spectrum), Filtronic (specialized amplifiers), Materion (strategic beryllium), and STMicroelectronics (antenna chips). The second group consists of hard-to-replace suppliers due to high switching costs, such as Honeywell (flight control), Carpenter Technology (specialty alloys), Hexcel (carbon fiber), Broadcom (data exchange), and Linde (industrial gases). The third group comprises high-volume, cost-critical suppliers for mass-produced items like Starlink terminals. Key names include Wistron NeWeb (primary manufacturer) and several A-share companies like Shenzhen Sunway (connectors), Pies New Materials (forgings), Western Superconducting (alloys), and Yingliu (castings). Other niche players include Trimble (timing), Astronics (power distribution), and CTS (thermal management). The article argues that investing in these suppliers, rather than SpaceX stock directly, offers an alternative opportunity. The rationale is threefold: procurement is just beginning to scale, SpaceX's IPO brings new transparency to its supply chain, and the situation mirrors early stages of past "super terminal" ecosystems like Apple or Tesla. While risks exist (commodity cycles, geopolitical factors, technology shifts), the core thesis is that SpaceX's massive, ongoing procurement will translate into reliable revenue for its key suppliers, regardless of its own stock price volatility.

marsbitHá 1h

The Foundation of SpaceX's Trillion-Dollar Valuation: Who is Dividing Up Musk's Annual Tens of Billions in Capital Expenditure?

marsbitHá 1h

SpaceX's Trillion-Dollar Valuation Base: Who's Sharing in Musk's Annual Tens of Billions in Capital Expenditure?

**Title: The Foundation of SpaceX's Trillion-Dollar Valuation: Who Benefits from Musk's Annual $100 Billion Capital Expenditure?** This article argues that investors seeking to benefit from SpaceX's growth might find greater opportunities in its supply chain rather than directly investing in the company itself, drawing parallels to historical successes with Apple, Tesla, and NVIDIA suppliers. **SpaceX's Business Model & Cash Flow:** SpaceX generates revenue from three main areas: 1. **Starlink:** Its profitable core, earning $11.3B in 2023 (60% of revenue), funding other ventures. 2. **Rockets (Falcon/Starship):** Requires $3B+ in annual R&D but achieves the world's lowest launch costs. 3. **AI:** Currently unprofitable (-$6B+ in 2023), investing heavily in ground-based supercomputers (220,000 GPUs) and future orbital data centers. The cycle is: Starlink profits → fund cheaper rockets → low-cost launches deploy AI hardware → AI compute rentals generate future revenue. This cycle drives annual procurement spending of tens of billions of dollars. **The Supply Chain Beneficiaries:** Suppliers are categorized by their replaceability: **1. Nearly Irreplaceable (High Barriers to Entry):** * **NVIDIA:** Powers the Colossus supercomputer; its CUDA ecosystem creates immense switching costs. * **Eutelsat (SATS):** Controls critical radio spectrum for satellite communications; holds a ~3% stake in SpaceX. * **Filtronic (FTC):** Supplies millimeter-wave signal amplifiers for Starlink satellites; SpaceX constitutes 83% of its revenue. * **Materion (MTRN):** Global leader in beryllium production, a strategic material used in Starship structures. * **STMicroelectronics (STM):** Supplies phased-array antenna chips for Starlink satellites. **2. Replaceable, but Switching Cost is Prohibitively High:** * **Honeywell (HON):** Provides flight control and inertial navigation systems with decades of certification. * **Carpenter Technology (CRS):** Manufactures ultra-pure specialty steel alloys for Raptor engines. * **Hexcel (HXL):** Supplies custom carbon fiber composites developed over a decade with SpaceX. * **Broadcom (AVGO):** Manages high-speed data switching. * **Linde Group:** Supplies industrial gases (liquid oxygen/nitrogen) from facilities built near SpaceX launch sites. **3. High-Volume, Cost-Critical Manufacturing:** Focuses on mass-producing components like Starlink user terminals (target: 30 million units). * **Key Players:** Wistron NeWeb (6285, primary terminal manufacturer), several Chinese A-share companies (e.g., Sunway Communication, PAX New Materials, Western Metal Materials, Yingliu Co.), and smaller US firms like Trimble (TRMB, timing systems). **Why Now?** Three factors make the supply chain opportunity timely: 1. **Volume Ramp-Up:** SpaceX plans 100 launches in 2026, aims for 30 million Starlink terminals, and will deploy AI data centers, meaning procurement will accelerate. 2. **Increased Transparency:** The IPO provides public financial data, allowing investors to track supplier order growth. 3. **Historical Precedent:** The current phase is likened to Tesla's early mass-production stage (circa 2018), suggesting a long growth runway for suppliers. **Conclusion:** The article posits that while investing in SpaceX stock is betting on Elon Musk's ambitious vision at a high valuation, investing in its established suppliers is a bet on the tangible, recurring revenue from its massive procurement budget, which is largely decoupled from day-to-day stock price volatility.

链捕手Há 1h

SpaceX's Trillion-Dollar Valuation Base: Who's Sharing in Musk's Annual Tens of Billions in Capital Expenditure?

链捕手Há 1h

Trading

Spot
Futuros
活动图片