One-Third of arXiv 'Contaminated', 65% of CS Papers Smell of AI, Only 0.7% in Math

marsbitPublished on 2026-07-27Last updated on 2026-07-27

Abstract

Approximately one-third of recently posted arXiv papers show signs of significant AI-generated text, according to a new study. An analysis of 12,750 papers from January 2023 to July 2026 across ten disciplines found a sharp increase in AI text markers following ChatGPT's release, with the overall detection rate reaching 32% in the latest quarter and peaking near 39% in early 2026. The rate varies drastically by field. Computer Science papers lead at 65%, followed by Quantitative Biology (56.3%) and Electrical Engineering (51.3%). Mathematics, however, has the lowest detection rate at just 0.7%. The study's authors note this could be due to mathematicians using AI less or because the detector struggles with the high volume of formulas and symbolic notation in math papers, leaving the true cause unclear. The research highlights that the detector identifies a statistical "AI style" in the text rather than proving full AI authorship. It cannot distinguish between light AI-assisted editing and fully AI-generated content. Furthermore, the detector can produce false positives, as some pre-ChatGPT academic writing also exhibits patterns now flagged as "AI-like." The growing use of AI, particularly in highly competitive fields, is creating a cycle where researchers may feel pressured to adopt AI tools to keep pace. The findings raise questions about the changing nature of academic writing and the emergence of a new "AI style" that is increasingly difficult to distinguish from human...

A third of arXiv papers are actually written by AI?

A recent unslop study has caused a stir overseas.

Over the past year, 65% of new papers in computer science have been flagged by its detector.

Netizens have even coined a term for it – 'The Great Slurry Era'.

Yet, for the same period, only 0.7% of math papers were flagged.

ChatGPT Appears, Curve Takes Off

In this study, the team scanned 12,750 arXiv papers spanning from January 2023 to July 2026.

The target was ten disciplines, with about 25 full-text papers sampled per field per month.

The control group was set from 2021 to 2022, the pure human era before ChatGPT's birth.

The results show that from 2021-2022, the flagging rate was stable at 0.4%; but with ChatGPT's release, the curve shot up within months.

In the most recent complete quarter, the flagging rate reached 32%, and in early 2026, the peak approached 39%.

By discipline, computer science fared the worst, with a flagging rate of 65%.

Before ChatGPT, its baseline was only 0.2%. In just three years, the number soared over 300-fold.

Following closely were quantitative biology at 56.3%, electrical engineering at 51.3%, economics & finance at 47%; then applied physics at 34%, statistics at 31.3%, condensed matter physics at 24%, high-energy physics at 14%, and astrophysics at 10.7%.

The bottom of the list was mathematics: 0.7%.

And this is just the lower bound. If AI text is hidden well enough, the detector can't identify it.

The same detector, the same paper database, a gap nearing 100-fold.

Are mathematicians collectively sticking to pure manual writing? Or is this machine simply blind in the face of mathematical formulas?

The Bottomed-Out 0.7%: Machine Blindness

In fact, the answer has been written into the unslop report itself.

Think, what does a standard mathematics paper look like? Screens full of symbols, formulas, theorems, and proofs. The genuine prose sentences written in English are pitifully few.

This detector only recognizes text. It focuses on sentence structure, word usage habits, paragraph organization.

And the few remaining lines of text in math papers are all formulaic expressions like 'Let X be such and such' or 'By Lemma 3, we can derive', which basically aren't in the same realm as the scientific English the detector was trained on.

So the team itself admits the study currently has several limitations:

The 0.7% for math doesn't prove much yet. It might be that mathematicians really don't use AI much, or the detector simply can't read math papers, but this data can't distinguish which.

The control group is also relatively small. Only 200 pre-ChatGPT papers per discipline, with a 0.4% false positive rate, means only about 8 flagged papers out of 2000 total. This level can only be considered a rough estimate.

Then there's the aforementioned issue of incomplete detection. So, these numbers are just lower bounds; the real proportions are certainly higher.

The More Competitive the Field, the More Reliance on AI

So, who is using AI to write papers?

The answer lies in another two-year Stanford study.

In March 2024, Stanford's Weixin Liang team first measured peer reviews: in ICLR 2024 review comments, about 10.6% of sentences were heavily modified by LLMs. The reviewers used it themselves before even finishing evaluating the papers.

In August 2025, the same team expanded the sample to 1.12 million papers and directly published in Nature Human Behaviour.

The research showed that as of September 2024, the proportion of LLM-modified text in CS paper abstracts had reached up to 22.5%. And the heaviest users were researchers publishing preprints most frequently, operating in the most competitive fields.

The more competitive the field, the heavier the use.

Here, AI writing has become an arms race: peers are using AI to speed up; those writing by hand alone fall a step behind.

No wonder a widely circulated related repost on X had the first line: 'Prompt: Write a paper publishable on arXiv.'

It Sniffs the 'Smell', Not the AI

But don't rush to convict based on that 65%.

The team also admits in the report that the detector cannot distinguish between 'AI polished the grammar' and 'the whole thing was machine-made'. It only detects the concentration of machine-like scent in the text.

So the correct understanding of 65% is '65% of papers smell like AI', not '65% of papers are written by AI'.

The trouble is, humans themselves can acquire this machine scent.

A skeptical researcher fed their own old papers from years ago into a detection tool, and 27% to 74% of the content was flagged red.

Remember, in that era, ChatGPT wasn't even a glimmer.

The reason isn't hard to find.

Open any academic journal today, and pages are filled with large language models, benchmark tests, state-of-the-art. The more standardized the phrasing, the more structured the composition, the more likely it is to hit the machine-like features recognized by the detector.

Besides, LLMs were trained on such papers. It's not that scholars write like AI; it's that AI was born writing like scholars.

When All Text Becomes Suspect

Actually, in recent years, we've all been doing the same thing as the detector.

Reading a passage that's perfectly balanced, neatly phrased, occasionally popping with 'not only... but also...', our hearts skip a beat: 'This must be AI-written, right?'

Unslop's detector essentially turned this mystical sense of smell into an instrument that mass-produces numbers.

After all, once suspicion enters the mind, it's hard to remove.

In the past, 'writing it down' was itself an endorsement. Someone willing to spend hours organizing text at least showed seriousness.

Now, no matter how beautifully written, it might only earn a 'AI, right?'.

Some have even started deliberately avoiding words they've used all their lives, simply because they 'sound too AI'.

Now, 'AI scent' is becoming a new type of original sin sweeping through the world of writing.

The ironic part is, to this day, we still haven't precisely measured what exactly constitutes this 'AI scent'.

This article is from WeChat public account 'Xinzhiyuan', author: ASI Revelation

Trending Cryptos

Related Questions

QAccording to the article, what percentage of recent arXiv papers in computer science were flagged as having 'AI flavor' by the detection tool?

A65% of recent arXiv papers in computer science were flagged.

QWhy might the detection tool show a very low flag rate (0.7%) for mathematics papers compared to other fields?

AThe tool analyzes text structure and language, but mathematics papers are dominated by symbols, formulas, and theorems, with very little prose. The limited text often uses formal, standardized phrasing that the tool may not recognize as typical scientific English, potentially making it 'blind' to AI use in this context.

QWhat does the Stanford study mentioned in the article suggest about the correlation between AI use and research field competitiveness?

AThe Stanford study suggests that the most competitive research fields, where researchers publish preprints most frequently, show the highest levels of LLM-assisted writing. It has become an 'arms race' where using AI provides a speed advantage.

QWhat is a key limitation of the AI-text detection tool discussed in the article?

AA key limitation is that the tool detects a general 'AI flavor' or style in the text but cannot distinguish between partial AI use (like grammar polishing) and fully AI-generated content. It also has high false positive rates, as it sometimes flags human-written academic prose.

QWhat broader societal concern regarding writing does the article raise in its conclusion?

AThe article raises the concern that 'AI flavor' is becoming a new kind of 'original sin' in writing. Well-written, clear, and structured text is now often met with suspicion of being AI-generated, which undermines the trust and value traditionally placed on carefully crafted human writing.

Related Reads

Why P2P and Exchangers Are Becoming Obsolete, and What Will Replace Them

Titled "Why P2P and Exchanges Are Becoming Obsolete, and What Will Replace Them," this article discusses the evolution of stablecoins, particularly USDT, from a trading tool to a global payment method. However, converting crypto back to fiat for daily expenses remains a challenge. The process of using P2P platforms or crypto exchanges has become increasingly risky and complex due to stricter banking anti-fraud measures, new legislation in Russia, and rampant fraud schemes like "triangles," where sellers can inadvertently receive stolen funds. The article highlights that while P2P was convenient, banks now scrutinize frequent peer-to-peer transfers, often freezing accounts. Exchanges also pose risks like unfavorable rates and unclear compliance. The market is therefore shifting towards integrated services that eliminate the manual conversion step. It cites OneSix as an example—a crypto wallet accessible via Telegram and web that embeds currency conversion directly into payment and withdrawal scenarios. Users can pay bills or send money as if using a bank app, with the crypto-to-fiat exchange happening seamlessly in the background. It also offers invoicing for freelancers and conducts AML checks, returning suspicious transactions instantly instead of freezing them. In conclusion, as crypto integrates into everyday finance, the demand for manual, risky exchange methods is declining. The future lies in unified platforms that combine storage, compliance, conversion, and payments into a single, secure user experience.

cryptonews.ru50m ago

Why P2P and Exchangers Are Becoming Obsolete, and What Will Replace Them

cryptonews.ru50m ago

The CEO of MARA Holdings Compares AI Operations to Bitcoin Mining! Which is More Profitable?

Fred Thiel, CEO of MARA Holdings (a major Bitcoin mining company), states that the rapid growth of the artificial intelligence (AI) sector is transforming mining companies' business models. He claims that powering data centers for AI is significantly more profitable than Bitcoin mining. In a recent interview, Thiel explained why many mining firms are diversifying into AI infrastructure. According to Thiel, the energy demands of data centers are surging, especially due to the spread of generative AI applications. This creates new revenue opportunities for Bitcoin mining companies with robust power infrastructure. MARA Holdings is among those monitoring this shift and aims to develop its energy and infrastructure services for AI data centers. Thiel emphasized that this move toward AI does not mean the end of Bitcoin mining. He stated that Bitcoin mining remains a sustainable business model, especially for miners in regions with low electricity costs, and it is still a crucial field for utilizing excess or idle power capacity. In recent years, many Bitcoin mining companies have begun using their energy-intensive infrastructure not just for block production but also for high-performance computing (HPC) and AI applications. This strategy aims to diversify revenue sources and increase resilience against cryptocurrency market volatility. Analysts note that the growing energy demand in the AI sector presents significant transformation opportunities for mining companies. The need for high-power, uninterrupted electricity supply, particularly for large data centers, gives Bitcoin miners with existing energy infrastructure expertise a considerable advantage.

cryptonews.ru1h ago

The CEO of MARA Holdings Compares AI Operations to Bitcoin Mining! Which is More Profitable?

cryptonews.ru1h ago

Trading

Spot

Hot Articles

Discussions

Welcome to the HTX Community. Here, you can stay informed about the latest platform developments and gain access to professional market insights. Users' opinions on the price of AI (AI) are presented below.

活动图片