Tens of Millions of Errors Per Hour: Investigation Reveals the 'Accuracy Illusion' of Google AI Search

marsbitPublished on 2026-04-13Last updated on 2026-04-13

Abstract

A New York Times investigation, in collaboration with AI startup Oumi, reveals significant accuracy and reliability issues with Google's AI Overviews search feature. Testing over 4,300 queries showed the accuracy rate improved from 85% (powered by Gemini 2) to 91% (Gemini 3). However, given Google's scale of ~5 trillion annual searches, this 9% error rate translates to nearly 57 million incorrect answers generated hourly. A critical finding is the prevalence of "unsubstantiated citations." For correct answers, the rate of citations that do not support the AI's summary surged from 37% to 56% with the Gemini 3 upgrade, making it difficult for users to verify information. The AI heavily relies on low-quality sources, with Facebook and Reddit being among its top-cited websites. Furthermore, the system is highly manipulable. A BBC journalist successfully "poisoned" it by publishing a fabricated article; Google's AI began presenting the false information as fact within 24 hours. Google disputed the study's methodology, criticizing its use of the SimpleQA benchmark and an AI model (Oumi's own) to evaluate another AI. The company maintains its AI Overviews, combined with its search ranking systems, perform better than the underlying model alone. Critics note this defense does little to bolster user confidence in the feature's reliability.

Author: Claude, Deep Tide TechFlow

Deep Tide Guide: A recent test conducted by The New York Times in collaboration with AI startup Oumi shows that the accuracy rate of Google Search's AI Overviews feature is approximately 91%. However, given Google's scale of processing 5 trillion searches annually, this translates to tens of millions of incorrect answers generated every hour. More troublingly, even when the answers are correct, over half of the cited links fail to support their conclusions.

Google is disseminating misinformation on an unprecedented scale, and most people are completely unaware.

According to The New York Times, AI startup Oumi, commissioned by the publication, used the industry-standard test SimpleQA, developed by OpenAI, to evaluate the accuracy of Google's AI Overviews feature. The test covered 4,326 search queries, conducted in two rounds: one in October last year (powered by Gemini 2) and another in February this year (upgraded to Gemini 3). The results showed that Gemini 2's accuracy was about 85%, which improved to 91% with Gemini 3.

91% sounds good, but it's a different story when considering Google's massive scale. Google processes approximately 5 trillion search queries annually. With a 9% error rate, AI Overviews generates over 57 million inaccurate answers per hour, nearly 1 million per minute.

Correct Answers, Wrong Sources

More alarming than the accuracy rate is the issue of "unsubstantiated citations."

Oumi's data shows that in the Gemini 2 era, 37% of correct answers had the problem of "unsubstantiated citations," meaning the links attached to the AI summary did not support the information provided. After upgrading to Gemini 3, this proportion increased instead of decreasing, jumping to 56%. In other words, while the model gives correct answers, it is increasingly failing to "show its work."

Oumi CEO Manos Koukoumidis pointedly questioned: "Even if the answer is correct, how do you know it's correct? How do you verify it?"

The heavy reliance on low-quality sources by AI Overviews exacerbates this problem. Oumi found that Facebook and Reddit are the second and fourth most cited sources for AI Overviews, respectively. In inaccurate answers, Facebook was cited 7% of the time, higher than the 5% rate in accurate answers.

BBC Journalist's Fake Article "Poisons" Results Within 24 Hours

Another serious flaw of AI Overviews is its susceptibility to manipulation.

A BBC journalist tested the system with a deliberately fabricated false article. In less than 24 hours, Google's AI Overview presented the false information from the article as fact to users.

This means anyone who understands how the system works could potentially "poison" AI search results by publishing false content and boosting its traffic. Google spokesperson Ned Adriance responded by stating that the search AI feature is built on the same ranking and security mechanisms used to block spam, and claimed that "most examples in the test are unrealistic queries that people wouldn't actually search for."

Google's Rebuttal: The Test Itself Is Flawed

Google raised several concerns about Oumi's study. A Google spokesperson called the research "seriously flawed," citing reasons including: the SimpleQA benchmark itself contains inaccurate information; Oumi used its own AI model, HallOumi, to judge another AI's performance, potentially introducing additional errors; and the test content does not reflect real user search behavior.

Google's internal tests also showed that when Gemini 3 operates independently outside the Google Search framework, it produces false outputs at a rate as high as 28%. However, Google emphasized that AI Overviews, leveraging the search ranking system, performs better in accuracy than the model alone.

Nevertheless, as PCMag pointed out in a logical paradox: If your defense is that "the report pointing out our AI's inaccuracies itself uses potentially inaccurate AI," this likely does not enhance user confidence in your product's accuracy.

Related Questions

QWhat was the accuracy rate of Google's AI Overviews feature as tested by Oumi, and how many errors does this translate to per hour given Google's search volume?

AThe accuracy rate of Google's AI Overviews was found to be 91% in the test. Given Google's annual volume of 5 trillion searches, this 9% error rate translates to over 57 million inaccurate answers generated every hour.

QAccording to the Oumi study, what was the trend in 'unsubstantiated citations' between the Gemini 2 and Gemini 3 versions of the AI Overviews?

AThe problem of 'unsubstantiated citations' (where the provided links did not support the AI's answer) increased from 37% with Gemini 2 to 56% with the upgraded Gemini 3.

QWhich low-quality websites were identified as major sources frequently cited by Google's AI Overviews?

AFacebook and Reddit were identified as the second and fourth most frequently cited sources by the AI Overviews feature.

QHow did a BBC journalist demonstrate the vulnerability of Google's AI Overviews to manipulation?

AA BBC journalist tested the system by publishing a deliberately fabricated article. Within 24 hours, Google's AI Overviews began presenting the false information from that article as a factual answer to user queries.

QWhat were Google's main criticisms of the Oumi study's methodology?

AGoogle criticized the study for having 'serious flaws,' stating that the SimpleQA benchmark itself contains inaccuracies, that using Oumi's own AI model to judge another AI could introduce errors, and that the test queries did not reflect real user search behavior.

Related Reads

Home Furnishing Listed Companies Are Trying to Turn Around with AI and Semiconductors

Chinese home furnishing listed companies are increasingly turning to AI and semiconductor cross-border ventures to revive their fortunes amid declining core businesses. A prominent example is PVC flooring giant **Elegant Home Furnishing**, whose stock price skyrocketed with 10 consecutive trading limits after announcing a plan to acquire a majority stake in **Ou Kang Nuo**, a semiconductor storage testing equipment and services company. Despite Elegant Home's own financial struggles—including recent losses and tight cash flow—the acquisition, funded through asset sales and loans, has dramatically boosted its market value. The deal involves a cross-shareholding structure with Ou Kang Nuo's controlling shareholder. This trend is widespread. Over the past year, numerous home furnishing firms with stagnant main operations have seen their stock prices surge after announcing moves into hot sectors like AI, semiconductors, or computing power. Key cases include: * **Markor Home Furnishings** ("high-end home furnishing first share"), after years of heavy losses, acquired an AI server high-speed copper cable company and later introduced AI computing power investors. * **Zhenai Meijia** ("blanket king") became the first A-listed manufacturer controlled by an AI large model company following a takeover. * **Fasilon** (integrated ceiling) saw its stock soar after establishing an AI subsidiary. * **Jinlong Decoration** experienced multiple trading limits after being linked to "commercial aerospace" and "AI computing," despite clarifying these segments contribute less than 1% of its business. Statistics show over 20 A-share companies announced semiconductor crossovers in 2025, with home furnishing and building materials firms being the most active. This shift is largely driven by pressures from the real estate downturn, shrinking overseas demand, and trade tariffs, pushing companies to seek growth through trendy concepts rather than core business improvement. However, the article warns that such speculative frenzies often end when the hype fades, leaving stock prices to eventually reflect the companies' weak fundamentals, as seen in Jinlong Decoration's subsequent sharp price correction. The repeated attempts to "ride the trend" highlight a desperate struggle for survival in a challenging traditional industry.

marsbit50m ago

Home Furnishing Listed Companies Are Trying to Turn Around with AI and Semiconductors

marsbit50m ago

"Sell America" Trade Resurfaces: Global Funds Reprice Washington Policy Risks, Dollar and Treasuries Bear the Brunt

"Sale of America" Trade Resurges as Global Funds Reprice Washington Policy Risks, Hitting Dollar and Treasuries First Recent policy signals from Washington are prompting global bond and foreign exchange investors to reignite the "Sale of America" debate. Uncertainty stems from Fed Chair Wash's shift towards less policy communication, Treasury Secretary Bessant's approval of joint intervention with Japan to support the Yen (the first such coordinated move in nearly 30 years), expanding fiscal deficits, and trade war fears. These factors are undermining confidence in U.S. assets. The 30-year Treasury yield recently broke above 5%, hitting a high not seen since 2007, while the Bloomberg Dollar Spot Index has fallen about 2% from its June peak. This dollar weakness is unusual given still-high U.S. interest rates. Investors cite policy uncertainty as a key driver, with one manager calling the Fed Chair and Treasury Secretary a "double whammy" creating a "Trump administration premium." While the S&P 500 continues hitting record highs and foreign holdings of U.S. Treasuries remain high, parts of the bond and FX markets are adjusting. A core concern is whether the Fed under Wash can effectively anchor inflation expectations; analysts warn that if the Fed lags the hiking cycle, long-end yields could face further upward pressure. The 30-year term premium has risen sharply. The U.S.-backed Yen intervention has also sparked debate on the dollar's structural outlook. While framed by Bessant as a "reallocation of reserves," market participants caution that if Japan—America's largest foreign creditor—must sell Treasuries to fund interventions, ripple effects could hit the bond market. Strategists predict moderate dollar depreciation ahead. Most experts are not predicting an end to dollar hegemony or the safe-haven status of Treasuries. However, a key underlying risk is highlighted: the pace of foreign buying of U.S. debt may not keep up with the speed of American borrowing, gradually eroding the structural advantages that have long supported "American exceptionalism."

marsbit51m ago

"Sell America" Trade Resurfaces: Global Funds Reprice Washington Policy Risks, Dollar and Treasuries Bear the Brunt

marsbit51m ago

Trading

Spot
活动图片