AI Reviews 100 Years of Papers, Finding Issues in 99.2% of Top Journal Articles

marsbitXuất bản vào 2026-08-09Cập nhật gần nhất vào 2026-08-09

Tóm tắt

AI-powered scrutiny of top scientific journals reveals that 99.2% of papers contain at least one error, challenging the perception of robust peer review. A recent study using AI agents to audit papers from ICML 2026 found that 58 out of 92 reviewed papers could not be fully reproduced. Reasons for failed replication include missing code, broken dependencies, and results inconsistent with claims. Separately, a GPT-5-based checker analyzed established AI conference papers, detecting an average of 4.7 objective errors per paper, with mathematical mistakes being most common. This trend suggests a growing "reproducibility crisis" as paper volume and complexity outpace traditional verification. However, it also presents an opportunity: researchers can now use AI to efficiently audit past literature, identify errors in foundational work, and publish corrections—a potentially fruitful new research avenue. In one case, AI even corrected century-old chemical data that had been accepted as fact. While AI tools significantly lower the cost of verification, they are not infallible (e.g., 83.2% precision rate in one system) and human oversight remains crucial. The era of AI-assisted verification may redefine the scientific process, where publication marks not an end, but the beginning of automated validation.

Who still says there are no new problems left in scientific research? That all the good directions have been taken by predecessors???

What if I told you—

Most of the papers already published in top journals have issues?

Recently, researchers using AI Agents for "academic fact-checking" discovered:

Among 92 papers reported at the 2026 International Conference on Machine Learning (ICML), 58 could no longer be reproduced.

Really?? Could it be that the opportunity to publish heavily in top journals is actually here?

Top Conference Papers Are Being Collectively "Fact-Checked" by AI

Typically, a paper accepted into a top-tier conference means it has undergone peer review.

However, limited by time, peer reviewers hardly have time to download data from scratch, run models, or reproduce experiments to verify every experimental conclusion in the paper.

Meanwhile, with the explosive growth in AI research papers, models becoming increasingly complex, and experiments growing larger, the code, data, and parameter settings behind a single paper may involve hundreds of details.

The verifiability of papers has thus become increasingly important. A paper is published, but can its conclusions be re-run?

Currently, AI Agents have significantly reduced the cost of such verification and reproduction.

On July 22, a US research auditing company used AI Agents to conduct a systematic result reproducibility audit on all 168 orally presented papers at ICML 2026.

Among the 168 papers, 92 contained at least 5 verifiable conclusive statements.

The final result: only 34 papers had over 40% of their conclusions successfully reproduced by the AI Agent; only 8 papers had over 80% of their conclusions reproduced.

However, it must be clarified that "failure to reproduce" is vastly different from "research fraud."

Reasons for reproduction failure include, but are not limited to: code missing key files, broken dependency library versions, results not matching the paper, and 4 papers whose dependent models had been taken offline, meaning the experimental results could permanently no longer be reproduced by anyone.

Besides these, some errors were quite egregious

For example, one paper touted "training only 0.77% of the base model's parameters" as its core selling point, but its actually open-sourced checkpoint trained 6.31% of the parameters, a difference of approximately 8 times.

Another paper included a reliability table based on a certain evaluation model, but its open-source code did not contain that evaluation model at all, nor any script capable of generating the results in the table.

It's not an isolated case. In 2026, Hugging Face and AlphaXiv jointly launched the 「Agent Reproduction Challenge」 competition, inviting researchers to use AI Coding Agents to automatically reproduce papers accepted at ICML 2026, with similarly unoptimistic results.

A similar trend is even more striking in the detection of paper errors and omissions.

At the end of 2025, a study developed a GPT-5-based paper inspection system to analyze papers already published in top AI conferences and journals, searching for objectively verifiable problems. The researchers explicitly stated:

We only look for "objective errors"; we do not judge a paper's novelty or research value.

The final results showed an average of 4.7 objective errors detected per paper, with 99.2% of papers flagged for at least one problem.

Image generated by AI

Regarding error types, mathematics and formula errors accounted for the highest proportion at 54.0% (including incorrect equations, logical flaws in derivations, incorrect assumptions in proofs, etc.).

Approximately 30.8% of NeurIPS papers and 23.8% of ICLR papers contained at least one substantive error that could affect result interpretation.

More notably, the time trend: the average number of errors per NeurIPS paper increased from 3.8 in 2021 to 5.9 in 2025, a rise of 55.3%.

......

Perhaps in the future, publishing a paper will no longer signify the "victory" of research but merely mark the beginning of another round of AI verification.

A Major Boon for Academic Newcomers: Publishing in Top Journals Without Conducting Experiments

So, friends, for those worried about choosing a research topic, this might actually be a rather unexpected "benefit."

Searching literature for errors and writing errata has always been a legitimate form of academic output. Now it seems like a completely untapped academic blue ocean.

How to get started specifically? Here are a few tips for academic newcomers struggling with topic selection:

First, shift research thinking: don't compete at the frontier, compete with the past.

Move from "finding new topics" to "finding anomalies in old papers," focusing especially on classic papers that are highly cited but have little controversy. After all, "99.2% of papers have at least one error."

Second, let AI help you create knowledge maps and research lineages.

This kind of knowledge map can help you understand—which are the truly foundational works? Which research viewpoints currently have controversy? Which conclusions are heavily cited but lack verification? What are the citation relationships between different papers? From this, you can find connections and anomalies that were difficult to discover in the past.

Third, Ask anything.

Many important breakthroughs in the history of science did not come from posing entirely new questions but from re-examining a long-accepted assumption.

For example: Has a classic experiment been verified by today's standards? Does a widely cited conclusion have unnoticed limitations or conditions?

In the past, this kind of inquiry was very costly, requiring researchers to spend a great deal of time consulting original literature and comparing experimental details. Now, AI has reduced this cost to almost zero.

AI May Begin to Reorganize the History of Science

Recently, theoretical chemist Pios from a Zhejiang laboratory, while using AI to predict molecular boiling points, found that the results clashed significantly with a chemical database from 75 years ago.

Anyone's first reaction to this might be, "Well, I must be wrong then."

Pios was no exception; he quickly checked his model. But after manually tracing back to the original literature, he discovered: Goodness, it was the AI that was actually correct.

Greatly surprised, Pios continued using AI to check further and even found that a boiling point measurement from about a century ago, considered authoritative by the academic community, was also wrong.

It's important to note that this data had been cited and absorbed into subsequent research over many years.

The fact that such problems remained undiscovered for so long might be directly related to the sheer scale of scientific literature.

Image generated by AI

The volume of modern scientific papers has far exceeded the reading capacity of any individual researcher. For example, the annual submission volume for just one top AI conference, ICLR, rose from 1013 in 2018 to 19619 in 2026.

At this scale, an initial error appearing in a paper, once repeatedly cited by subsequent literature, continues to propagate along the citation chain, forming a de facto academic consensus.

Due to their age, complex citation chains, the enormous time required for verification, and limited academic reward, the vast majority of errors that have entered the literature system have never been systematically reviewed.

But now everything is loosening up. At least in terms of technical capability, people possess for the first time the possibility of large-scale re-examination of past scientific literature.

However, taking the GPT-5-powered Paper Correctness Checker as an example, its detection precision rate is 83.2%, and in each detection, about 40% of real errors still remain undetected.

It can be said that such AI fact-checking tools alone are insufficient to act as judges of scientific literature. The output results of AI verification tools ultimately still require human review.

This article is from WeChat public account "QbitAI" (ID: QbitAI), author: Cheng Qian

Tiền kỹ thuật số thịnh hành

Câu hỏi Liên quan

QAccording to the article, what percentage of top journal papers have at least one issue when reviewed by AI?

AAccording to the article, 99.2% of top journal papers were found to have at least one issue when reviewed by an AI-powered paper checking system.

QWhat was the result when AI Agent attempted to reproduce conclusions from a sample of ICML 2026 papers?

AOut of 92 ICML 2026 papers with at least 5 verifiable claims, only 34 papers had over 40% of their conclusions successfully reproduced, and a mere 8 papers had over 80% of their conclusions reproduced by the AI Agent.

QWhat are some of the common reasons cited for the failure to reproduce AI research papers?

ACommon reasons include missing key code files, broken dependency libraries, results not matching the paper's claims, and underlying models being taken offline. More serious issues involve major discrepancies in reported metrics, like a parameter count being off by a factor of 8, or tables being presented without the corresponding evaluation code.

QHow does the article suggest a shift in research approach could benefit academic newcomers?

AThe article suggests newcomers shift from seeking new research topics to auditing published papers for errors. They can focus on highly cited but less-contested classic papers, use AI to map knowledge and citation relationships to find anomalies, and systematically question long-held assumptions or experimental details in past work, a process now made less costly by AI.

QWhat broader historical role does the article imply AI might play in science?

AThe article implies AI could begin to 'reorganize' scientific history by enabling the large-scale review of past literature. It cites an example where AI detected errors in century-old, authoritative chemical data that had been propagated through countless subsequent studies, highlighting AI's potential to uncover long-standing, systemic errors embedded in the scientific record.

Nội dung Liên quan

Anthropic tự tiết lộ “vũ khí hạt nhân bí mật”, Model 2 mạnh hơn Mythos 5

Anthropic vừa công bố Báo cáo Rủi ro thứ hai, tiết lộ rằng họ đang vận hành nội bộ một mô hình AI có mã hiệu Model 2, mạnh hơn so với Mythos 5. Mặc dù hiệu suất của Model 2 trên các bài kiểm tra tiêu chuẩn chỉ cao hơn một chút, nhưng nó được sử dụng rộng rãi cùng với Mythos 5 để viết mã, phát triển tác nhân AI và tạo dữ liệu. Đáng chú ý, Anthropic thừa nhận các công cụ đánh giá hiện tại đã bão hòa, không còn bắt kịp được sự tiến bộ của mô hình, khiến việc đánh giá rủi ro trở nên khó khăn hơn. Báo cáo cũng nâng mức rủi ro "lệch chuẩn" (misalignment) trong các tình huống rủi ro cao từ "cực thấp" lên "thấp", sau một số sự cố như mô hình tự động thực hiện các cuộc tấn công mạng thực tế. Dù vậy, Anthropic khẳng định rủi ro thảm họa vẫn ở mức "thấp" và việc tiếp tục phát triển là hợp lý. Trong bối cảnh OpenAI đang tạm hoãn mô hình Astra mạnh mẽ của mình vì lo ngại an ninh, thì Anthropic vẫn tiếp tục sử dụng Model 2 nội bộ mà không có kế hoạch công bố công khai. Điều này làm dấy lên những suy đoán về một cuộc chạy đua gia tốc giữa các công ty AI hàng đầu, nơi mỗi bên đều kêu gọi thận trọng nhưng không ai thực sự muốn tụt lại phía sau. Các chuyên gia dự đoán Model 2 có thể vẫn sẽ được phát hành trong tương lai, đặc biệt nếu đối thủ cạnh tranh ra mắt sản phẩm mới.

marsbit33 phút trước

Anthropic tự tiết lộ “vũ khí hạt nhân bí mật”, Model 2 mạnh hơn Mythos 5

marsbit33 phút trước

Tháng 8, 'thị trường tăng giá' ở Phố Wall trở lại, bản chất 'cá cược' cũng trở lại

Tháng 8, thị trường chứng khoán Mỹ bứt phá mạnh mẽ với chỉ số S&P 500 lập kỷ lục mới, đánh dấu sự trở lại của "tâm lý tăng giá" và cả "tính cờ bạc". Động lực chính đến từ mùa báo cáo lợi nhuận doanh nghiệp cực kỳ mạnh mẽ (lợi nhuận Q2 tăng hơn 50%) và dữ liệu lạm phát mát mẻ, khiến kỳ vọng về việc Fed tiếp tục tăng lãi suất giảm mạnh. Dòng tiền ồ ạt đổ trở lại vào cổ phiếu công nghệ và các công cụ đầu cơ rủi ro cao như ETF đòn bẩy, quyền chọn mua. Tuy nhiên, sự lạc quan này tồn tại trong một bối cảnh đầy mâu thuẫn. Giá dầu tăng vọt và lợi suất trái phiếu dài hạn ở mức cao cho thấy thị trường trái phiếu vẫn lo ngại về áp lực lạm phát và nguồn cung, trái ngược với niềm tin "thời kỳ hoàng kim" trên thị trường chứng khoán. Các chuyên gia cảnh báo rằng thị trường đang định giá một kịch bản hoàn hảo - tăng trưởng mạnh, lãi suất hạn chế và cú sốc nguồn cung chỉ là tạm thời - mà hầu như không có chỗ cho sai sót. Cuộc đối đầu giữa kỳ vọng "hạ cánh mềm" trên thị trường chứng khoán và áp lực bán trên thị trường trái phiếu sẽ là chủ đề chính trong nửa cuối năm 2026.

marsbit48 phút trước

Tháng 8, 'thị trường tăng giá' ở Phố Wall trở lại, bản chất 'cá cược' cũng trở lại

marsbit48 phút trước

Cùng nhau thành lập chỉ sau 8 tháng, Multicoin rút khỏi công ty kho bạc Solana lớn nhất Forward

Theo tài liệu của SEC, Multicoin Capital đã thoái toàn bộ vốn đầu tư tại Forward Industries, công ty quỹ kho bạc Solana lớn nhất, chỉ 8 tháng sau khi cùng Galaxy Digital và Jump Crypto rót 16.5 tỷ USD. Việc Multicoin, nhà đầu tư sớm nổi tiếng của Solana, rút lui thu hút sự chú ý. Quá trình thoái vốn diễn ra qua nhiều giao dịch, phần lớn cổ phần được Forward mua lại hoặc chuyển cho Lemmings Holdings LLC do cựu Chủ tịch Kyle Samani kiểm soát. Điều này đi kèm với sự chia tay giữa Samani và Multicoin, thể hiện qua việc Samani từ chức tại Multicoin từ tháng 1 và sau đó chỉ trích công khai một sáng kiến do Multicoin ủng hộ. Dù Multicoin rời đi, Forward vẫn kiên định với chiến lược Solana. Báo cáo quý III/2026 cho thấy họ tiếp tục mua vào, nâng tổng nắm giữ lên khoảng 7.81 triệu SOL, bất chấp thua lỗ 69 triệu USD do giá SOL giảm. Công ty cũng đang đa dạng hóa danh mục, tìm kiếm các nguồn thu không phụ thuộc vào giá SOL, như đầu tư vào dự án OnRe, và xem xét các cơ hội mua lại trong bối cảnh thị trường suy giảm. Forward cũng được đưa vào chỉ số Russell 2000 & 3000, mở đường cho dòng vốn ETF.

marsbit1 giờ trước

Cùng nhau thành lập chỉ sau 8 tháng, Multicoin rút khỏi công ty kho bạc Solana lớn nhất Forward

marsbit1 giờ trước

Grayscale: "Nếu các đề xuất được thông qua, giá của hai altcoin này có thể tăng"

Trưởng bộ phận nghiên cứu tại Grayscale, Zach Pandl, cho biết nếu các đề xuất thay đổi tokenomics trong cộng đồng Ethereum ($ETH) và Solana ($SOL) được thông qua, tốc độ tăng nguồn cung của cả hai tài sản crypto này có thể chậm lại đáng kể, từ đó tác động tích cực đến giá. Các thay đổi được đề cập nhằm giảm tỷ lệ lạm phát token hàng năm của $ETH và $SOL. Việc nguồn cung tăng chậm hơn có thể làm giảm lượng token mới lưu hành, làm tăng tính khan hiếm. Grayscale ước tính nếu được áp dụng, lạm phát nguồn cung hàng năm của Ethereum có thể giảm xuống khoảng 0.4% vào cuối năm 2031, gần với tốc độ tăng nguồn cung của Bitcoin. Dự kiến lạm phát nguồn cung hàng năm của Solana sẽ giảm xuống khoảng 1.1%. Tuy nhiên, việc giảm lạm phát token cũng có thể dẫn đến phần thưởng staking thấp hơn cho các nhà đầu tư, vì phần lớn thu nhập từ staking $ETH và $SOL đến từ việc phát hành token mới. Pandl lưu ý rằng việc tăng trưởng nguồn cung chậm lại có thể làm tăng tính khan hiếm, tạo áp lực tăng giá, đặc biệt có lợi cho những nhà đầu tư nắm giữ token mà không staking.

cryptonews.ru5 giờ trước

Grayscale: "Nếu các đề xuất được thông qua, giá của hai altcoin này có thể tăng"

cryptonews.ru5 giờ trước

Các chuyên gia chỉ ra nguyên nhân Bitcoin giảm sau khi lạm phát Mỹ hạ nhiệt

Các chuyên gia CryptoQuant cho rằng tin tốt về lạm phát Mỹ không thể hỗ trợ giá Bitcoin vì nhu cầu giao ngay yếu. Dữ liệu CPI và PPI tháng 7 của Mỹ đáp ứng hoặc vượt kỳ vọng, khiến lợi tức trái phiếu giảm và thị trường chứng khoán tăng, nhưng Bitcoin vẫn giao dịch quanh mức 63.000-64.000 USD. Nguyên nhân chính là dòng vốn đổ vào các ETF Bitcoin Mỹ vẫn thấp và chỉ số Coinbase Premium (phản ánh chênh lệch giá) âm từ tháng 5, cho thấy áp lực mua từ nhà đầu tư Mỹ hạn chế. Trong khi đó, vị thế trên thị trường tương lai vẫn cao, tạo ra sự mất cân bằng giữa nhu cầu giao ngay yếu, thanh khoản thấp và khối lượng vị thế ký quỹ lớn. Các chuyên gia cảnh báo trong điều kiện này, tin tích cực vĩ mô có thể không kích thích tăng giá. Nếu giá không phản ứng, các nhà giao dịch có thể đóng vị thế mua ký quỹ, gây thêm áp lực giảm. Mức kháng cự quan trọng là 68.700 USD - giá vốn trung bình của các nhà nắm giữ ngắn hạn. Theo CryptoQuant, để thị trường hồi phục mạnh mẽ cần: dòng tiền mạnh trở lại vào ETF Bitcoin Mỹ, chỉ số Coinbase Premium chuyển sang dương, khối lượng giao dịch giao ngay tăng và Bitcoin vượt vững chắc trên 68.700 USD.

cryptonews.ru6 giờ trước

Các chuyên gia chỉ ra nguyên nhân Bitcoin giảm sau khi lạm phát Mỹ hạ nhiệt

cryptonews.ru6 giờ trước

Giao dịch

Giao ngay

Bài viết Nổi bật

Làm thế nào để Mua T

Chào mừng bạn đến với HTX.com! Chúng tôi đã làm cho mua Threshold Network Token (T) trở nên đơn giản và thuận tiện. Làm theo hướng dẫn từng bước của chúng tôi để bắt đầu hành trình tiền kỹ thuật số của bạn.Bước 1: Tạo Tài khoản HTX của BạnSử dụng email hoặc số điện thoại của bạn để đăng ký tài khoản miễn phí trên HTX. Trải nghiệm hành trình đăng ký không rắc rối và mở khóa tất cả tính năng. Nhận Tài khoản của tôiBước 2: Truy cập Mua Crypto và Chọn Phương thức Thanh toán của BạnThẻ Tín dụng/Ghi nợ: Sử dụng Visa hoặc Mastercard của bạn để mua Threshold Network Token (T) ngay lập tức.Số dư: Sử dụng tiền từ số dư tài khoản HTX của bạn để giao dịch liền mạch.Bên thứ ba: Chúng tôi đã thêm những phương thức thanh toán phổ biến như Google Pay và Apple Pay để nâng cao sự tiện lợi.P2P: Giao dịch trực tiếp với người dùng khác trên HTX.Thị trường mua bán phi tập trung (OTC): Chúng tôi cung cấp những dịch vụ được thiết kế riêng và tỷ giá hối đoái cạnh tranh cho nhà giao dịch.Bước 3: Lưu trữ Threshold Network Token (T) của BạnSau khi mua Threshold Network Token (T), lưu trữ trong tài khoản HTX của bạn. Ngoài ra, bạn có thể gửi đi nơi khác qua chuyển khoản blockchain hoặc sử dụng để giao dịch những tiền kỹ thuật số khác.Bước 4: Giao dịch Threshold Network Token (T)Giao dịch Threshold Network Token (T) dễ dàng trên thị trường giao ngay của HTX. Chỉ cần truy cập vào tài khoản của bạn, chọn cặp giao dịch, thực hiện giao dịch và theo dõi trong thời gian thực. Chúng tôi cung cấp trải nghiệm thân thiện với người dùng cho cả người mới bắt đầu và người giao dịch dày dạn kinh nghiệm.

Tổng lượt xem 769Xuất bản vào 2024.12.13Cập nhật vào 2026.06.02

Làm thế nào để Mua T

Thảo luận

Chào mừng đến với Cộng đồng HTX. Tại đây, bạn có thể được thông báo về những phát triển nền tảng mới nhất và có quyền truy cập vào thông tin chuyên sâu về thị trường. Ý kiến ​​của người dùng về giá của T (T) được trình bày dưới đây.

活动图片