# Accuracy İlgili Makaleler

HTX Haber Merkezi, kripto endüstrisindeki piyasa trendleri, proje güncellemeleri, teknoloji gelişmeleri ve düzenleyici politikaları kapsayan "Accuracy" hakkında en son makaleleri ve derinlemesine analizleri sunmaktadır.

Predicting World Cup Knockout Matches: Why Are Different AI Models So Far Apart?

AI performance in predicting the 2026 FIFA World Cup knockout matches varied significantly, according to an analysis of models including ChatGPT, Grok, DeepSeek, Gemini, and Claude. The standout predictions came from DeepSeek and Gemini for the Netherlands vs. Morocco match. Gemini precisely forecasted a 1-1 draw and a penalty shootout win for Morocco, while DeepSeek correctly identified the high probability of a draw and Morocco's potential to advance via a defensive and counter-attacking strategy. Grok and Tongyi Qianwen (千问) demonstrated strength in predicting accurate scores for matches with clearer favorites. They correctly called the narrow 1-0 win for Canada over South Africa and Brazil's 2-1 victory over Japan, as well as Norway's 2-1 win over Ivory Coast. ChatGPT and Claude excelled more in match process analysis than in predicting exact scores or upsets. They frequently identified potential challenges for favorites, such as Japan's pressing against Brazil or DR Congo's defensive tactics against England, even when predicting the favorite's ultimate victory. A notable failure was the unanimous misjudgment of Germany vs. Paraguay. All models incorrectly favored Germany, underestimating Paraguay's ability to force a penalty shootout and cause an upset. In summary, Gemini and DeepSeek showed the most insight for high-stakes, unpredictable matches. Grok and Qianwen were reliable "score predictors" for less volatile games. ChatGPT and Claude were strong "analytical models," adept at outlining match dynamics but often hesitant to predict upsets.

Odaily星球日报07/02 01:44

Predicting World Cup Knockout Matches: Why Are Different AI Models So Far Apart?

Odaily星球日报07/02 01:44

Someone Predicts South Korean Stock Market with Hyperliquid, Achieving 74% Accuracy?

A study analyzed whether weekend price movements of four Korean stock perpetual futures contracts (Samsung Electronics, SK Hynix, Hyundai Motor, and EWY) on Hyperliquid could predict their Monday opening directions on their respective primary exchanges (KRX, NYSE). Over 62 weekend observations across the four assets, Hyperliquid correctly predicted the Monday opening direction 45 times (73.8% accuracy). However, performance varied significantly. Samsung Electronics showed the strongest and statistically significant signal, with Hyperliquid's weekend close correctly predicting its KRX Monday open in 15 out of 16 cases (94% accuracy, p-value < 0.001). This signal remained strong (75% accuracy) even when using Saturday's close instead of Sunday's, suggesting genuine price discovery beyond last-minute convergence. Hyundai Motor also showed high accuracy (81%, 13/16 correct), but this was not statistically significant after accounting for a baseline downward bias in its Monday opens. SK Hynix performed marginally better than a coin flip (63%, 10/16). EWY performed the worst (54%, 7/13), underperforming a simple strategy of always predicting a Monday rise. The stark difference between Samsung and EWY is largely attributed to market timing. KRX opens shortly after Hyperliquid's Sunday close, while NYSE opens ~14 hours later, allowing new information to flow in. The results suggest that for assets like Samsung Electronics, where weekend trading on Hyperliquid precedes the primary market open by only minutes, the platform can provide a valuable predictive signal worth monitoring before the Monday auction, despite the currently small weekend trading volumes.

Foresight News06/10 10:08

Someone Predicts South Korean Stock Market with Hyperliquid, Achieving 74% Accuracy?

Foresight News06/10 10:08

Probability in the Price: How World Cup Odds Are Calculated

**The Probability in the Price: How World Cup Odds Are Calculated** Two major systems released their "championship probabilities" before the 2026 World Cup, and they disagreed on the favorite. Prediction market aggregators listed France at around **17%**, while the Opta supercomputer gave European champion Spain **16.1%**. These numbers look similar, but their production methods are fundamentally different. The market's **17%** is the **price** that clears after hundreds of millions of dollars in trading across platforms like Polymarket and Kalshi, where contracts trade between 0 and 100 cents, directly representing implied probability. This liquidity is provided by crypto-native market makers like Wintermute, though the market still has "the liquidity profile of an early-stage" asset class. In contrast, Opta's **16.1%** is a **simulated frequency**. Its model uses team data (including betting market odds as an input) to estimate match probabilities, then runs **10,000 full tournament simulations**, counting how often each team wins. Which is more accurate? There is **no rigorous, cross-tournament academic study** directly comparing their track records. However, a persistent **longshot bias**—where low-probability outcomes are systematically overvalued—observed in traditional betting for nearly a century, has also been found in modern crypto prediction markets. Research shows low-price contracts on Kalshi/Polymer less likely to pay out than their implied odds suggest. Unlike traditional bookmakers, prediction markets operate on **public blockchain ledgers**, making every transaction auditable and enabling such research. However, price formation is also influenced by **regulatory uncertainty**, as seen in recent US state-level bans and legal battles over jurisdiction. In summary, the "probability" you see is either a **market-clearing price** subject to behavioral biases and liquidity constraints, or a **model-simulated frequency** that partially incorporates market data. The question of which method is more reliable remains open, highlighting the importance of asking: **How was this number produced?**

marsbit06/05 00:26

Probability in the Price: How World Cup Odds Are Calculated

marsbit06/05 00:26

Topping GitHub's Trending, the Essential Guide for Claude Code Users

The CLAUDE.md file, trending on GitHub, is a project-level guide for Claude Code designed to dramatically improve its accuracy and efficiency. It addresses key issues like repetitive context explanations, unauthorized code changes, and forgotten decisions across sessions. By placing this plain-text file in a project root, Claude Code reads it automatically at the start of each session. The guide includes rules to eliminate redundant explanations, enforce strict behavioral constraints (e.g., no modifications outside the requested scope without confirmation), and establish a "memory" system using companion files like MEMORY.md and ERRORS.md to log past decisions and failures. It also locks in the project's specific tech stack to prevent inappropriate tool recommendations. Highlighted are four foundational rules from Andrej Karpathy that reportedly increased coding accuracy from 65% to 94%: always ask for clarity first, implement the simplest solution, never touch unrelated code, and explicitly flag uncertainties. The article quantifies significant weekly cost savings for developers and teams by eliminating wasted time on re-explaining context, rolling back unauthorized edits, and re-evaluating previously rejected solutions. The core message is that a small, upfront investment in creating a CLAUDE.md file leads to a more predictable, controlled, and cost-effective AI programming assistant.

marsbit05/18 09:38

Topping GitHub's Trending, the Essential Guide for Claude Code Users

marsbit05/18 09:38

Polymarket Is Not an All-Powerful "Truth Machine"

Polymarket, a crypto-based betting platform, is often hailed as a "truth machine" for its ability to aggregate crowd wisdom through financial stakes. While it has demonstrated remarkable accuracy in predicting major events like the 2024 U.S. presidential election—outperforming traditional polls—its overall reliability is highly inconsistent. Analysis using the Brier score reveals that its predictive power excels in high-liquidity domains like politics and economics but falls to near-random or worse in categories like sports, culture, and tech. The platform’s growing influence is concerning as its odds are increasingly cited by major media outlets like The Wall Street Journal and CNN, lending them an air of authority. This visibility creates a feedback loop where the odds themselves can influence the outcomes they are meant to predict—a phenomenon known as endogeneity. Moreover, the market is vulnerable to manipulation by well-resourced "whales" with access to exclusive information, such as private polls or even military intelligence, as seen in cases involving bets on geopolitical events. While useful for short-term, high-stakes events, Polymarket’s predictions are often unreliable for the vast majority of its contracts due to low liquidity and wide bid-ask spreads. The danger lies not in its occasional failures, but in the unchecked trust it receives—risking a future where a handful of traders can shape perceived reality through a platform masquerading as an oracle of truth.

marsbit04/15 11:40

Polymarket Is Not an All-Powerful "Truth Machine"

marsbit04/15 11:40

Wikipedia Implements New Editing Rules: Vote Passes, Strictly Prohibits Using AI to Generate or Rewrite Article Content

On March 26, Wikipedia officially passed a new policy through a community vote that explicitly prohibits users from directly using AI to generate or rewrite article content. This decision reinforces the platform's commitment to content accuracy and human editorial control. The updated policy strengthens previous guidelines by moving from a recommendation against generating articles from scratch to a strict ban on using large language models (LLMs) for content creation or rewriting. The policy was approved overwhelmingly by volunteer editors, with a vote of 40 to 2, reflecting deep concerns within the community about AI-generated misinformation and inaccuracies. While AI tools are still permitted for suggesting basic edits, they must not introduce any unverified content. All AI-assisted contributions must undergo human review to prevent factual errors or hallucinations. This move highlights Wikipedia’s effort to balance technological efficiency with content integrity amid the growing use of generative AI across digital platforms. By clearly distinguishing between AI-assisted editing and AI-generated content, Wikipedia aims to preserve human-driven knowledge curation and prevent trust issues caused by automated content production. The decision sets a significant precedent for ethical knowledge management in the age of artificial intelligence.

marsbit03/27 01:08

Wikipedia Implements New Editing Rules: Vote Passes, Strictly Prohibits Using AI to Generate or Rewrite Article Content

marsbit03/27 01:08

Only 60% Real Win Rate: Data Reveals the Truth Behind ICO Predictions on Polymarket

Polymarket's TokenSale markets have processed nearly $250 million in volume, boasting impressive accuracy rates—100% for fundraising amounts and over 90% for fully diluted valuations (FDV). However, an analysis of 231 prediction markets across 29 token sales reveals these figures are misleading. The platform functions more as a sentiment indicator, often acting as a contrarian signal. Key findings show that the true prediction accuracy one week before market close is only 66.7%, meaning the crowd is wrong one-third of the time, with errors consistently skewing toward over-optimism. FDV predictions averaged a 35% overestimation. Analysis of 24-hour post-launch volatility showed an average price swing of ±23%, with 75% of tokens facing sell-offs. Only 62.5% of 24-hour FDV predictions were accurate. The 100% accuracy claim is meaningless because markets close after results are known. High trading volume on Polymarket often serves as a reverse indicator—more optimism typically leads to greater inaccuracy. Tokens with conservative predictions (e.g., Monad, Football.fun) saw smaller declines. Actionable signals: High volume (>$50M) and high optimism (>50% FDV overestimation) are bearish. Low volume (<$5M) and accurate predictions (within 20% of actual FDV) are relatively bullish. In a market where most tokens fall below ICO price, "less bad" is the best outcome. Polymarket’s token sales market is essentially a hype meter—extreme confidence often signals maximum investor pain.

marsbit01/31 03:19

Only 60% Real Win Rate: Data Reveals the Truth Behind ICO Predictions on Polymarket

marsbit01/31 03:19

活动图片