Do people in frontier labs hardly read papers anymore?
An OpenAI researcher takes a broad swipe at three top-tier conferences with one sentence: Too much exaggeration and fraud!

The origin of the matter was an ICLR paper with results that were suspiciously good. After some investigation, netizens discovered the "secret" behind it.

Has publishing in top conferences "evolved" to this point?

First, check if the code is open source; second, see if the experimental methods hide any "underhanded tricks"; and third, examine whether the results actually support the core claims — after this multi-layered interrogation, how many top conference papers can truly withstand scrutiny?
Someone actually did it.
Only 8 Out of 105 Papers Were "Qualified"
In July this year, SAI, co-founded by Chenghao Tan, an associate professor of computer science and data science at the University of Chicago, published the results of a large-scale "experimental review."
They selected all 168 Oral papers from ICML 2026.
This year, ICML received 23,918 submissions, and only 168 were selected for Oral presentations, accounting for about 0.7%.
In other words, SAI was reviewing the very top layer filtered from over twenty thousand submissions.
In addition to reading the paper's methodology, experimental design, and results like traditional reviewers, SAI Review also downloads the code, models, and data, sets up the environment, runs the experiments, and compares the results item by item with the original paper.
SAI examined all 168 Oral papers. Only 104 of those papers had open-sourced code, and they completed full replication for 105 papers.
Among them, only 34 papers replicated more than 40% of the claimed results; only 8 papers had a replication rate exceeding 80%.

Whether each paper contained at least 1, 3, or 5 verifiable claims, the median replication score remained at 28%–30%, with little change in the overall distribution.

Even after excluding experiments that couldn't run, terminated early, or exceeded hardware capabilities, the median replication score was only 42%–50%; when each paper contained at least 3 verifiable claims, the median stabilized at 42%.
They encountered the most common issues in replication: code that wouldn't run, missing files, incomplete instructions, broken dependencies, or code results that didn't match the paper.
There were also 4 papers that depended on models which were already taken offline. Later researchers, even if willing to spend money and time, couldn't possibly obtain the same results.
SAI also listed two more specific examples.
One paper's main selling point was "training only 0.77% of the base model's parameters," but the checkpoints they released actually trained 6.31% of the parameters, approximately 8 times the claimed figure.
Another paper presented a reliability table scored by a judge model, but the open-sourced code did not include this judge model, nor were there any scripts capable of calculating the numbers in the table.
Poor Papers: High Reward, Low Risk
The replication results from SAI are quite dismal.
What's worse, the problems discovered here are only those that were "found under specific conditions."
Papers without code are directly "flawless."

And even with completely open code, the time, effort, and money required to verify a single paper are enormous. When the final outcome is unsatisfactory, the frustration is immeasurable.
According to SAI's estimates based on Google Cloud's public on-demand pricing, the median cost to fully rerun one ICML Oral paper is about $8,900. Among the 105 papers, 17 exceeded $100,000, and the most expensive one approached $2.2 million.
The more a paper relies on large-scale computing power, the more difficult independent replication becomes.
Without code, others can't check; with code, few can afford the cost.
Thus, a problematic paper can smoothly pass review, gain citations, be added to a resume, and secure admissions, faculty positions, or jobs in major labs.
Occasionally, someone might find mismatched results, but conferences rarely re-examine them, and the papers are seldom retracted.
This is what the comments refer to as "there's a reward for producing bad research, but almost no cost."

Not Reading Papers, But Hiring Based on Papers
In the field of large language models, many truly significant advances no longer appear in the form of papers.
As OpenAI engineers standing at the industry forefront, it's understandable to feel this way. With more abundant computing power, faster experimental feedback, and a wealth of unpublished internal results, they have the confidence to "not read papers."
But saying it out loud carries a certain nuance.
One comment mocked people in tech companies for "pulling up the ladder after climbing ashore": recruiting researchers from academia, building on publicly accumulated academic achievements, snatching talent and GPUs with higher prices, themselves increasingly avoiding peer review, and then turning around to declare academic research "mostly a scam."

It's slightly harsh, but there's truth to it.
People in these frontier labs can "not read papers," but those wanting to enter still have to publish papers first.
For students and young researchers lacking industry experience, top conference papers remain the most direct proof of research capability.
Applying for PhD programs, seeking faculty positions, or entering frontier labs like OpenAI all rely on papers to gain attention.
People in the industry, after entering major labs, may look down on papers, yet they still use paper counts, conference tiers, and citation metrics to filter those standing outside the gate.
Thus, papers occupy an awkward position: their credibility in knowledge dissemination is questioned, yet their value in talent competition has not diminished in the slightest.
Therefore, "major labs don't read papers anymore" sounds somewhat lofty and arrogant. Because those truly qualified not to read papers have often already crossed that threshold by virtue of their own papers.
Of course, not all students are meticulously designing fraudulent academic antagonists.
An academic researcher joked that he wished his own students were even capable of writing a paper full of exaggeration or fraud.

Some are "striving to be academic fraudsters," while others are still struggling with LaTeX.
Reference Links:
https://x.com/MathewShen42/status/2084465434506768867
https://x.com/kellerjordan0/status/2084721463089902074
https://sai.science/blog/how-much-science-is-verifiable
https://x.com/mengyer/status/2085134204786921886
This article is from the WeChat public account "Machine Heart" (ID: almosthuman2014), author: Machine Heart





