Tens of Millions of Errors Per Hour: Investigation Reveals the 'Accuracy Illusion' of Google AI Search

marsbitXuất bản vào 2026-04-13Cập nhật gần nhất vào 2026-04-13

Tóm tắt

A New York Times investigation, in collaboration with AI startup Oumi, reveals significant accuracy and reliability issues with Google's AI Overviews search feature. Testing over 4,300 queries showed the accuracy rate improved from 85% (powered by Gemini 2) to 91% (Gemini 3). However, given Google's scale of ~5 trillion annual searches, this 9% error rate translates to nearly 57 million incorrect answers generated hourly. A critical finding is the prevalence of "unsubstantiated citations." For correct answers, the rate of citations that do not support the AI's summary surged from 37% to 56% with the Gemini 3 upgrade, making it difficult for users to verify information. The AI heavily relies on low-quality sources, with Facebook and Reddit being among its top-cited websites. Furthermore, the system is highly manipulable. A BBC journalist successfully "poisoned" it by publishing a fabricated article; Google's AI began presenting the false information as fact within 24 hours. Google disputed the study's methodology, criticizing its use of the SimpleQA benchmark and an AI model (Oumi's own) to evaluate another AI. The company maintains its AI Overviews, combined with its search ranking systems, perform better than the underlying model alone. Critics note this defense does little to bolster user confidence in the feature's reliability.

Author: Claude, Deep Tide TechFlow

Deep Tide Guide: A recent test conducted by The New York Times in collaboration with AI startup Oumi shows that the accuracy rate of Google Search's AI Overviews feature is approximately 91%. However, given Google's scale of processing 5 trillion searches annually, this translates to tens of millions of incorrect answers generated every hour. More troublingly, even when the answers are correct, over half of the cited links fail to support their conclusions.

Google is disseminating misinformation on an unprecedented scale, and most people are completely unaware.

According to The New York Times, AI startup Oumi, commissioned by the publication, used the industry-standard test SimpleQA, developed by OpenAI, to evaluate the accuracy of Google's AI Overviews feature. The test covered 4,326 search queries, conducted in two rounds: one in October last year (powered by Gemini 2) and another in February this year (upgraded to Gemini 3). The results showed that Gemini 2's accuracy was about 85%, which improved to 91% with Gemini 3.

91% sounds good, but it's a different story when considering Google's massive scale. Google processes approximately 5 trillion search queries annually. With a 9% error rate, AI Overviews generates over 57 million inaccurate answers per hour, nearly 1 million per minute.

Correct Answers, Wrong Sources

More alarming than the accuracy rate is the issue of "unsubstantiated citations."

Oumi's data shows that in the Gemini 2 era, 37% of correct answers had the problem of "unsubstantiated citations," meaning the links attached to the AI summary did not support the information provided. After upgrading to Gemini 3, this proportion increased instead of decreasing, jumping to 56%. In other words, while the model gives correct answers, it is increasingly failing to "show its work."

Oumi CEO Manos Koukoumidis pointedly questioned: "Even if the answer is correct, how do you know it's correct? How do you verify it?"

The heavy reliance on low-quality sources by AI Overviews exacerbates this problem. Oumi found that Facebook and Reddit are the second and fourth most cited sources for AI Overviews, respectively. In inaccurate answers, Facebook was cited 7% of the time, higher than the 5% rate in accurate answers.

BBC Journalist's Fake Article "Poisons" Results Within 24 Hours

Another serious flaw of AI Overviews is its susceptibility to manipulation.

A BBC journalist tested the system with a deliberately fabricated false article. In less than 24 hours, Google's AI Overview presented the false information from the article as fact to users.

This means anyone who understands how the system works could potentially "poison" AI search results by publishing false content and boosting its traffic. Google spokesperson Ned Adriance responded by stating that the search AI feature is built on the same ranking and security mechanisms used to block spam, and claimed that "most examples in the test are unrealistic queries that people wouldn't actually search for."

Google's Rebuttal: The Test Itself Is Flawed

Google raised several concerns about Oumi's study. A Google spokesperson called the research "seriously flawed," citing reasons including: the SimpleQA benchmark itself contains inaccurate information; Oumi used its own AI model, HallOumi, to judge another AI's performance, potentially introducing additional errors; and the test content does not reflect real user search behavior.

Google's internal tests also showed that when Gemini 3 operates independently outside the Google Search framework, it produces false outputs at a rate as high as 28%. However, Google emphasized that AI Overviews, leveraging the search ranking system, performs better in accuracy than the model alone.

Nevertheless, as PCMag pointed out in a logical paradox: If your defense is that "the report pointing out our AI's inaccuracies itself uses potentially inaccurate AI," this likely does not enhance user confidence in your product's accuracy.

Câu hỏi Liên quan

QWhat was the accuracy rate of Google's AI Overviews feature as tested by Oumi, and how many errors does this translate to per hour given Google's search volume?

AThe accuracy rate of Google's AI Overviews was found to be 91% in the test. Given Google's annual volume of 5 trillion searches, this 9% error rate translates to over 57 million inaccurate answers generated every hour.

QAccording to the Oumi study, what was the trend in 'unsubstantiated citations' between the Gemini 2 and Gemini 3 versions of the AI Overviews?

AThe problem of 'unsubstantiated citations' (where the provided links did not support the AI's answer) increased from 37% with Gemini 2 to 56% with the upgraded Gemini 3.

QWhich low-quality websites were identified as major sources frequently cited by Google's AI Overviews?

AFacebook and Reddit were identified as the second and fourth most frequently cited sources by the AI Overviews feature.

QHow did a BBC journalist demonstrate the vulnerability of Google's AI Overviews to manipulation?

AA BBC journalist tested the system by publishing a deliberately fabricated article. Within 24 hours, Google's AI Overviews began presenting the false information from that article as a factual answer to user queries.

QWhat were Google's main criticisms of the Oumi study's methodology?

AGoogle criticized the study for having 'serious flaws,' stating that the SimpleQA benchmark itself contains inaccuracies, that using Oumi's own AI model to judge another AI could introduce errors, and that the test queries did not reflect real user search behavior.

Nội dung Liên quan

Các đối thủ của Bitcoin không vượt qua được thử thách: Báo cáo mới nhất tiết lộ sự thật về altcoin! 'Chỉ một số ít altcoin là người chiến thắng!'

Trong khi Bitcoin (BTC) thiết lập các kỷ lục mới vào năm 2025, nhiều altcoin, bao gồm cả Ethereum, đã cho thấy đà tăng trưởng hạn chế hơn. Một báo cáo toàn diện từ Blockworks Research đã phân tích hiệu suất của thị trường altcoin trong những năm gần đây. Báo cáo tiết lộ rằng altcoin thường hoạt động kém hơn Bitcoin. Chỉ 1,7% số altcoin đạt vốn hóa thị trường tối thiểu 50 triệu USD và có lịch sử 24 tháng (từ đầu 2020 đến tháng 6/2026) mới có thể vượt trội hơn Bitcoin. Về dài hạn, Bitcoin đã vượt mặt đại đa số các altcoin. Trong đợt thị trường tăng giá 2020-2021, khoảng một phần ba số altcoin đã mang lại lợi nhuận cao hơn Bitcoin. Tuy nhiên, thành công này không bền vững: 86% trong số những altcoin đó đã mất hơn 90% giá trị kể từ đỉnh. Đáng chú ý, trong số 187 altcoin vượt Bitcoin năm 2021, chỉ có OKB duy trì được hiệu suất vượt trội sau đó. Phân tích xác định rằng chỉ có 22 altcoin vượt Bitcoin về lợi nhuận trong giai đoạn 2020-2026. Danh sách này bao gồm các cái tên như BNB, OKB, GT, LEO, BGB, WBT, MX, CAKE, Ethereum, Solana và Tron. Cần lưu ý rằng nghiên cứu chỉ tập trung vào các dự án đạt vốn hóa tối thiểu 50 triệu USD, không phải tất cả tiền điện tử.

cryptonews.ru20 phút trước

Các đối thủ của Bitcoin không vượt qua được thử thách: Báo cáo mới nhất tiết lộ sự thật về altcoin! 'Chỉ một số ít altcoin là người chiến thắng!'

cryptonews.ru20 phút trước

Doanh nghiệp gia đình niêm yết đều muốn dựa vào AI, bán dẫn để lật ngược tình thế

Công ty sàn giao dịch chứng khoán sàn nhà cửa, Ai muốn dựa vào AI và bán dẫn để lật ngược tình thế. Một công ty sản xuất sàn nhựa PVC là Aili Home Furnishing, thông qua việc mua lại ngành bán dẫn, đã đạt được 10 phiên tăng giá liên tiếp, với tổng mức tăng lên đến 159.31% từ ngày 21/7 đến ngày 6/8. Điều này cho thấy việc mua lại chéo đang trở thành một kịch bản phổ biến để các doanh nghiệp sàn nhà cửa tìm cách thoát khỏi tình trạng khó khăn. Aili Home Furnishing dự định mua ít nhất 77.08% cổ phần của Oukangnuo, một công ty chuyên về thiết bị kiểm tra bộ nhớ và dịch vụ kiểm tra liên quan. Chi tiết giao dịch bao gồm việc trao đổi cổ phần hai chiều và cam kết lợi nhuận. Tuy nhiên, tình hình tài chính của Aili không mấy khả quan: lợi nhuận ròng năm 2025 giảm mạnh 87.6% và dự kiến thua lỗ trong nửa đầu năm 2026 là từ 34.5 đến 40.5 triệu nhân dân tệ. Việc mua lại chủ yếu dựa vào việc xử lý tài sản và vay mượn. Aili không phải là trường hợp duy nhất. Nhiều công ty sàn nhà cửa khác như Markor Home Furnishings, Zhenai Meijia, Faslong và Jin Tailong cũng đã công bố các động thái liên quan đến AI, bán dẫn hoặc điện toán hiệu suất cao, dẫn đến việc tăng giá cổ phiếu đáng kể. Thống kê ban đầu cho năm 2025 cho thấy hơn 20 công ty niêm yết A-share đã công bố việc mở rộng sang lĩnh vực bán dẫn, với tỷ lệ cao nhất là các doanh nghiệp sàn nhà cửa và vật liệu xây dựng. Động lực đằng sau làn sóng này là sự suy giảm của ngành bất động sản, nhu cầu trong nước và quốc tế yếu, cùng với áp lực thuế quan, khiến cho các doanh nghiệp truyền thống tìm kiếm các khái niệm đầu tư mới để kích thích thị trường và tạo ra khả năng sinh lời. Tuy nhiên, nhiều trường hợp trong quá khứ cho thấy sau khi cơn sốt đầu cơ qua đi, giá cổ phiếu thường quay trở lại với thực tế kinh doanh cơ bản. Bản thân Aili cũng đã cảnh báo về sự không chắc chắn của thỏa thuận mua lại chính thức. Cuối cùng, việc các doanh nghiệp sàn nhà cửa đổ xô vào các lĩnh vực thời thượng phần lớn là một nỗ lực tuyệt vọng để tồn tại khi hoạt động chính gặp khó khăn. Nhưng sau màn kịch sôi động, luôn có người phải trả giá.

marsbit49 phút trước

Doanh nghiệp gia đình niêm yết đều muốn dựa vào AI, bán dẫn để lật ngược tình thế

marsbit49 phút trước

Giao Dịch 'Bán Mỹ' Tái Xuất Hiện: Vốn Toàn Cầu Định Giá Lại Rủi Ro Chính Sách Washington, Đô La Mỹ, Trái Phiếu Mỹ Chịu Tác Động Trước Tiên

Tín hiệu chính sách liên tục từ Washington đang khiến nhà đầu tư toàn cầu tái kích hoạt thảo luận về "bán tháo Mỹ", tập trung vào đồng USD và trái phiếu kho bạc Mỹ. Chủ tịch Fed Waller thay đổi phong cách giao tiếp, Bộ trưởng Tài chính Bousquet phê duyệt hỗ trợ Nhật can thiệp tỷ giá (lần đầu tiên sau 30 năm), cùng với thâm hụt ngân sách mở rộng và mây mưa chiến tranh thương mại, đang làm lung lay niềm tin vào tài sản Mỹ. Lợi suất trái phiếu kho bạc kỳ hạn 30 năm vượt 5%, cao nhất kể từ 2007. Chỉ số Bloomberg Dollar Spot giảm khoảng 2% so với đỉnh tháng 6. Các nhà quản lý quỹ như Gama Asset Management và Brandywine Global đang bán hoặc giữ quan điểm yếu đối với USD và trái phiếu Mỹ, do lo ngại rủi ro chính sách dưới thời chính quyền Trump. Mối lo ngại chính là uy tín của Fed trong việc kiểm soát lạm phát. Phần bù rủi cho trái phiếu 30 năm ở mức cao nhất kể từ 2013. Áp lực nguồn cung từ việc Bộ Tài chính Mỹ điều chỉnh tăng dự báo vay nợ quý này lên 7390 tỷ USD cũng đè nặng lên thị trường. Việc can thiệp ngoại hối hỗ trợ đồng Yên đặt ra câu hỏi về triển vọng cấu trúc của USD. Chuyên gia cảnh báo nếu Nhật Bản - chủ nợ nước ngoài lớn nhất của Mỹ - phải bán trái phiếu kho bạc để tài trợ can thiệp, hiệu ứng lan tỏa có thể xảy ra. Mặc dù vậy, "chủ nghĩa ngoại lệ Mỹ" chưa chấm dứt. Tài sản Mỹ vẫn hấp dẫn, thiếu bằng chứng về một đợt bán tháo đồng loạt quy mô lớn. Tuy nhiên, điểm yếu tiềm ẩn là dòng vốn nước ngoài có thể không theo kịp tốc độ mở rộng nợ công của Mỹ, đe dọa vị thế tài sản an toàn và có thể dẫn đến sự suy yếu của đồng USD trong những năm tới.

marsbit50 phút trước

Giao Dịch 'Bán Mỹ' Tái Xuất Hiện: Vốn Toàn Cầu Định Giá Lại Rủi Ro Chính Sách Washington, Đô La Mỹ, Trái Phiếu Mỹ Chịu Tác Động Trước Tiên

marsbit50 phút trước

Giao dịch

Giao ngay
活动图片