Tens of Millions of Errors Per Hour: Investigation Reveals the 'Accuracy Illusion' of Google AI Search

marsbitXuất bản vào 2026-04-13Cập nhật gần nhất vào 2026-04-13

Tóm tắt

A New York Times investigation, in collaboration with AI startup Oumi, reveals significant accuracy and reliability issues with Google's AI Overviews search feature. Testing over 4,300 queries showed the accuracy rate improved from 85% (powered by Gemini 2) to 91% (Gemini 3). However, given Google's scale of ~5 trillion annual searches, this 9% error rate translates to nearly 57 million incorrect answers generated hourly. A critical finding is the prevalence of "unsubstantiated citations." For correct answers, the rate of citations that do not support the AI's summary surged from 37% to 56% with the Gemini 3 upgrade, making it difficult for users to verify information. The AI heavily relies on low-quality sources, with Facebook and Reddit being among its top-cited websites. Furthermore, the system is highly manipulable. A BBC journalist successfully "poisoned" it by publishing a fabricated article; Google's AI began presenting the false information as fact within 24 hours. Google disputed the study's methodology, criticizing its use of the SimpleQA benchmark and an AI model (Oumi's own) to evaluate another AI. The company maintains its AI Overviews, combined with its search ranking systems, perform better than the underlying model alone. Critics note this defense does little to bolster user confidence in the feature's reliability.

Author: Claude, Deep Tide TechFlow

Deep Tide Guide: A recent test conducted by The New York Times in collaboration with AI startup Oumi shows that the accuracy rate of Google Search's AI Overviews feature is approximately 91%. However, given Google's scale of processing 5 trillion searches annually, this translates to tens of millions of incorrect answers generated every hour. More troublingly, even when the answers are correct, over half of the cited links fail to support their conclusions.

Google is disseminating misinformation on an unprecedented scale, and most people are completely unaware.

According to The New York Times, AI startup Oumi, commissioned by the publication, used the industry-standard test SimpleQA, developed by OpenAI, to evaluate the accuracy of Google's AI Overviews feature. The test covered 4,326 search queries, conducted in two rounds: one in October last year (powered by Gemini 2) and another in February this year (upgraded to Gemini 3). The results showed that Gemini 2's accuracy was about 85%, which improved to 91% with Gemini 3.

91% sounds good, but it's a different story when considering Google's massive scale. Google processes approximately 5 trillion search queries annually. With a 9% error rate, AI Overviews generates over 57 million inaccurate answers per hour, nearly 1 million per minute.

Correct Answers, Wrong Sources

More alarming than the accuracy rate is the issue of "unsubstantiated citations."

Oumi's data shows that in the Gemini 2 era, 37% of correct answers had the problem of "unsubstantiated citations," meaning the links attached to the AI summary did not support the information provided. After upgrading to Gemini 3, this proportion increased instead of decreasing, jumping to 56%. In other words, while the model gives correct answers, it is increasingly failing to "show its work."

Oumi CEO Manos Koukoumidis pointedly questioned: "Even if the answer is correct, how do you know it's correct? How do you verify it?"

The heavy reliance on low-quality sources by AI Overviews exacerbates this problem. Oumi found that Facebook and Reddit are the second and fourth most cited sources for AI Overviews, respectively. In inaccurate answers, Facebook was cited 7% of the time, higher than the 5% rate in accurate answers.

BBC Journalist's Fake Article "Poisons" Results Within 24 Hours

Another serious flaw of AI Overviews is its susceptibility to manipulation.

A BBC journalist tested the system with a deliberately fabricated false article. In less than 24 hours, Google's AI Overview presented the false information from the article as fact to users.

This means anyone who understands how the system works could potentially "poison" AI search results by publishing false content and boosting its traffic. Google spokesperson Ned Adriance responded by stating that the search AI feature is built on the same ranking and security mechanisms used to block spam, and claimed that "most examples in the test are unrealistic queries that people wouldn't actually search for."

Google's Rebuttal: The Test Itself Is Flawed

Google raised several concerns about Oumi's study. A Google spokesperson called the research "seriously flawed," citing reasons including: the SimpleQA benchmark itself contains inaccurate information; Oumi used its own AI model, HallOumi, to judge another AI's performance, potentially introducing additional errors; and the test content does not reflect real user search behavior.

Google's internal tests also showed that when Gemini 3 operates independently outside the Google Search framework, it produces false outputs at a rate as high as 28%. However, Google emphasized that AI Overviews, leveraging the search ranking system, performs better in accuracy than the model alone.

Nevertheless, as PCMag pointed out in a logical paradox: If your defense is that "the report pointing out our AI's inaccuracies itself uses potentially inaccurate AI," this likely does not enhance user confidence in your product's accuracy.

Câu hỏi Liên quan

QWhat was the accuracy rate of Google's AI Overviews feature as tested by Oumi, and how many errors does this translate to per hour given Google's search volume?

AThe accuracy rate of Google's AI Overviews was found to be 91% in the test. Given Google's annual volume of 5 trillion searches, this 9% error rate translates to over 57 million inaccurate answers generated every hour.

QAccording to the Oumi study, what was the trend in 'unsubstantiated citations' between the Gemini 2 and Gemini 3 versions of the AI Overviews?

AThe problem of 'unsubstantiated citations' (where the provided links did not support the AI's answer) increased from 37% with Gemini 2 to 56% with the upgraded Gemini 3.

QWhich low-quality websites were identified as major sources frequently cited by Google's AI Overviews?

AFacebook and Reddit were identified as the second and fourth most frequently cited sources by the AI Overviews feature.

QHow did a BBC journalist demonstrate the vulnerability of Google's AI Overviews to manipulation?

AA BBC journalist tested the system by publishing a deliberately fabricated article. Within 24 hours, Google's AI Overviews began presenting the false information from that article as a factual answer to user queries.

QWhat were Google's main criticisms of the Oumi study's methodology?

AGoogle criticized the study for having 'serious flaws,' stating that the SimpleQA benchmark itself contains inaccuracies, that using Oumi's own AI model to judge another AI could introduce errors, and that the test queries did not reflect real user search behavior.

Nội dung Liên quan

Chú ý: Tuần này sẽ diễn ra đợt mở khóa token quy mô lớn của 10 altcoin! Đây là danh sách theo ngày và giờ

Tuần qua, thị trường tiền mã hóa trải qua đợt suy giảm do sự cố hack ví ColdCard và ảnh hưởng từ các sự kiện địa chính trị. Tuy nhiên, tuần này sẽ chứng kiến đợt mở khóa token quy mô lớn của 10 altcoin, có thể tác động đến nguồn cung và giá cả. Lịch trình mở khóa chi tiết (giờ UTC+3): - **Lagrange (LGR)**: 4/8, 03:00 - Giá trị: 1.38 triệu USD (15.04% vốn hóa). - **Proof (PROOF)**: 5/8, 03:00 - 39.11 triệu USD (119.59% vốn hóa). - **Power Protocol (POWER)**: 5/8, 03:00 - 1.62 triệu USD (8.93% vốn hóa). - **Verona (VERONA)**: 5/8, 03:00 - 1.37 triệu USD (12.61% vốn hóa). - **Ethena (ENA)**: 5/8, 11:00 - 15.28 triệu USD (1.80% vốn hóa). - **Goldfinger (GF)**: 6/8, 03:00 - 11.52 triệu USD (5.05% vốn hóa). - **Infinity (INF)**: 7/8, 03:00 - 2.31 triệu USD (20.30% vốn hóa). - **Stable (STBL)**: 8/8, 03:00 - 28.75 triệu USD (3.55% vốn hóa). - **Name (NAME)**: 9/8, 03:00 - 48.47 triệu USD (74.54% vốn hóa). - **Move (MOVE)**: 9/8, 03:00 - 1.22 triệu USD (3.90% vốn hóa). Đặc biệt, cần lưu ý đến PROOF với khối lượng mở khóa vượt 119% vốn hóa thị trường và NAME với gần 75%. Các sự kiện này có thể làm gia tăng áp lực bán trên thị trường. Đây không phải là lời khuyên đầu tư.

cryptonews.ru1 giờ trước

Chú ý: Tuần này sẽ diễn ra đợt mở khóa token quy mô lớn của 10 altcoin! Đây là danh sách theo ngày và giờ

cryptonews.ru1 giờ trước

Với giá 100.000 đô la mỗi tháng: Truth Social bán quyền truy cập bài đăng của Trump cho các công ty đầu tư

Trump Media and Technology Group (TMTG) đã ra mắt dịch vụ Truth API từ ngày 1/8/2026. Đây là kênh dữ liệu có phí cung cấp cho các khách hàng tổ chức, chủ yếu là các công ty đầu tư và giao dịch tần suất cao, quyền truy cập thời gian thực đến các bài đăng từ những tài khoản có ảnh hưởng nhất trên nền tảng Truth Social, bao gồm cả cựu Tổng thống Donald Trump. Theo các nguồn tin, gói dịch vụ này có giá lên tới 100.000 USD một tháng, với mức giảm giá xuống 60.000 USD/tháng cho hợp đồng ba năm. TMTG tuyên bố đây là một phần trong chiến lược tạo ra nguồn thu ổn định và lợi nhuận cao từ tài sản của công ty. Tuy nhiên, sáng kiến này đã vấp phải chỉ trích từ các nhà lập pháp cả hai đảng. Các Thượng nghị sĩ Dân chủ Elizabeth Warren và Adam Schiff đã yêu cầu Ủy ban Chứng khoán Mỹ (SEC) điều tra xem liệu việc bán quyền truy cập ưu tiên đến các bài đăng của tổng thống có vi phạm luật hay không. Thượng nghị sĩ Cộng hòa Bill Cassidy cũng chỉ trích đây là hành vi bán quyền truy cập đặc quyền không thể chấp nhận được. Phân tích AI trong bài báo cảnh báo về rủi ro tiềm ẩn, so sánh với sự kiện năm 2013 khi thị trường chứng khoán sụt giảm nhanh chóng do tin tức giả mạo. Việc biến tài khoản tổng thống thành một nút tín hiệu thị trường với độ trễ mili giây có thể tạo ra mục tiêu cho tin tặc hoặc thao túng, và đặt ra câu hỏi về trách nhiệm nếu thông tin sai lệch được phát tán qua kênh này.

cryptonews.ru3 giờ trước

Với giá 100.000 đô la mỗi tháng: Truth Social bán quyền truy cập bài đăng của Trump cho các công ty đầu tư

cryptonews.ru3 giờ trước

Các giao dịch rút Bitcoin tiếp tục: 8 năm lưu trữ trong ví lạnh Coldcard kết thúc bằng số không

Ví phần cứng Coldcard bị xâm phạm, dẫn đến làn sóng rút tiền mới từ các thiết bị dễ bị tấn công. Theo Galaxy Research, tổng số tiền bị đánh cắp đã lên tới 1.367,05 BTC (khoảng 88,6 triệu USD). Vấn đề không nằm ở phần mềm cập nhật, mà ở seed phrase (cụm từ khôi phục) được tạo từ tháng 3/2021 do lỗi lập trình, khiến chúng dễ bị dò tìm. Lỗi này xảy ra khi thiết bị chuyển từ bộ tạo số ngẫu nhiên phần cứng sang bộ tạo phần mềm Yasmarang, được khởi tạo bằng dữ liệu có thể dự đoán được. Người dùng các model Mk2-Mk5 và Q với phiên bản phần mềm nhất định cần tạo seed phrase mới trên bản cập nhật đã sửa và chuyển tài sản sang đó để bảo vệ. Câu chuyện đau lòng của một nhà đầu tư 39 tuổi đã mất 2 BTC (130.000 USD) tích góp suốt 8 năm trong vài phút, dù áp dụng chiến lược "mua và giữ trong ví lạnh" thận trọng. Anh mua Bitcoin như một lá chắn chống siêu lạm phát và kế hoạch nghỉ hưu sớm, nhưng lỗ hổng đã phá hỏng mọi thứ. Sự việc nhấn mạnh rằng lưu trữ offline không tự động đảm bảo an toàn, và cộng đồng hy vọng nhà sản xuất có thể tìm cách khắc phục, hoàn trả tài sản cho người dùng.

cryptonews.ru5 giờ trước

Các giao dịch rút Bitcoin tiếp tục: 8 năm lưu trữ trong ví lạnh Coldcard kết thúc bằng số không

cryptonews.ru5 giờ trước

Giao dịch

Giao ngay
活动图片