"AI Burning Books" Is Actually a Misunderstanding

marsbit发布于2026-08-19更新于2026-08-19

文章摘要

"AI Book-Burning" Is Actually a Misunderstanding Recent reports about AI companies purchasing used books, scanning them, and then destroying the physical copies have sparked widespread outrage. Terms like "AI is devouring human knowledge" have become common, fueled by dramatic visuals of books being cut and shredded. However, the actual facts reveal a more nuanced story. While companies like Anthropic have indeed spent millions to buy and "destructively scan" several million books for AI training, this volume is a small fraction of the global second-hand book market. The core act—digitizing content and then discarding the physical object—is the opposite of historical book-burning, which aimed to erase knowledge. A key point of contention is the purchase of rare or out-of-print books. Yet, if these books were legally for sale on the open market, the buyer (whether an AI firm or an individual) has the right to do with them as they wish. The real question is whether society has adequate systems to protect books of genuine cultural heritage *before* they are sold. Expecting profit-driven companies to self-regulate on this is unreliable; the solution lies in establishing public rules, such as protected lists for rare editions or granting libraries priority purchase rights. Much of the intense public reaction stems not from the scale of actual harm, but from the powerful symbolism. The image of books being fed into machines taps into deeper anxieties about AI: fears of job disp...

By Zimubiao

In Silicon Valley, the "AI burning books" saga just keeps getting wilder.

In a recent report, a media outlet collaborated with second-hand book dealers, placing an AirTag in a batch of books purchased by a mysterious buyer. They tracked it all the way to an Amazon factory in Las Vegas.

Factory employees indicated this was part of an Amazon project called VGT3, specifically dedicated to cutting and scanning books.

A second-hand book ended up following the plotline of a crime drama.

At this point, the situation has become somewhat surreal.

Over the past month, reports about "AI burning books" have emerged one after another, yet they haven't revealed much new information—AI companies are indeed buying large quantities of old paper books, cutting off the spines, high-speed scanning them to digitize the content, and then disposing of the original books. When Anthropic's "Project Panama" was exposed earlier this year, the most sensational elements were already laid out on the table: millions of books, tens of millions of dollars in procurement, and destructive scanning.

Later, when second-hand booksellers worldwide noticed unusual orders, ISBNdb was exposed for sourcing paper books specifically for AI companies, and now AirTags have been tracked into an Amazon warehouse—these developments have merely been confirming the same fact repeatedly: yes, this is really happening, and more than one company is doing it.

Clearly, the facts remain the same, but the public sentiment around "AI burning books" is intensifying.

"AI companies are devouring books," "Human civilization is being fed to machines," "They won't even spare out-of-print books"—countless human fingers fly across social media leaving horrified comments, with videos and photos of books being cut being particularly provocative.

But to be fair, at this stage, there's a significant disconnect between the known facts and the meaning ascribed to them.

AI companies are indeed destroying books; let's not sugarcoat that.

But exactly how many are they destroying? Does destroying a physical book equate to destroying knowledge? When an out-of-print book is bought and then cut up, is the responsibility solely on the AI companies?

It's time to debunk some myths.

01

Let's start with the conclusion here. Much of the current anger towards "AI destroying books" is directed at the symbolic meaning of the act, rather than the actual harm that has occurred so far.

Of course, the two are easily conflated. After all, headlines like "AI companies are buying and destroying antique books" and "AI companies are shredding millions of books after scanning them" are genuinely frightening, and the visual impact of books being cut is powerful. It's hard to find a more fitting image to represent "AI devouring human civilization."

But if we zoom out a little, the situation isn't quite as apocalyptic as it appears in those images.

The largest known scale of this activity comes from Anthropic. After launching "Project Panama" in 2024, aiming to scan all the world's books, it spent tens of millions of dollars over roughly a year to purchase and destructively scan millions of physical books, sometimes buying tens of thousands at a time.

However, "millions of books" placed within the entire second-hand book market is not as staggering as one might imagine.

Take the large second-hand bookseller World of Books alone: in 2024, it sold 31 million books globally, maintained a long-term inventory of about 7 million books, and its website typically lists around 2 million different titles.

In other words, Anthropic's processing volume, compared to the normal circulation scale of a major second-hand bookseller, is just a few months' worth of their business.

So, "AI companies have scanned and destroyed millions of books" is true, but the leap from there to "AI is exhausting the supply of second-hand books" is a long one.

Furthermore, there's another issue easily loaded with emotional significance that needs clarification: the disappearance of a physical book does not equal the disappearance of the knowledge within it.

Anthropic bought these books precisely to preserve their content. In the copyright lawsuit against Anthropic by writer Andrea Bartz and others, court documents revealed that after each physical book is disassembled and scanned, it becomes a PDF containing full-page images and machine-readable text, entering Anthropic's own digital library.

If we only discuss the process of "buying a book, cutting it open, scanning it, and destroying the original," it is almost the opposite action of "burning books" in a historical sense.

Burning books aims to make content disappear. AI companies, by cutting books open, specifically aim to preserve the content. Moreover, once used for model training, it will become a living resource in new forms. Judges have expressed similar views: the content from books used to train models is transformed into new capabilities during the training process.

Those cut-up physical books are indeed painful to look at, but human knowledge hasn't been sold for scrap along with them.

02

There is a point worth noting, and it's one of the strongest arguments in the current wave of criticism. AI companies' procurement doesn't consist solely of common old books available everywhere; it does include some rare second-hand books, even out-of-print ones.

This issue cannot simply be glossed over with "the physical book is gone, but the knowledge remains." Because the value of some books isn't only in the text. Special editions, unique bindings, author inscriptions and annotations, even the provenance of a book—whose collection it passed through, what marks it left—can make the physical object itself irreplaceable.

But first, we need to distinguish one thing: scarcity does not equal preciousness, and being out-of-print does not equal being cultural heritage.

Your university professor self-published 300 copies of "Handbook of County Sewage Pipeline Construction" in 1994. Today, perhaps only a dozen copies remain worldwide—this is certainly scarce, even long out of print. But whether it is a treasure of human civilization that must be protected at all costs is probably self-evident.

The second-hand book market is full of such items. The reason they are hard to find might simply be that they were printed in small numbers or never reprinted. It doesn't mean that the loss of each copy represents a loss to public culture.

Secondly, even if we narrow our focus to those books whose physical forms are genuinely valuable, the problem is still not that simple.

As long as a book is still freely circulating as a commodity in the second-hand market, it can be legally bought by a "crazy person."

This buyer could buy it to enshrine in a climate-controlled glass cabinet, or use it as a table leg coaster. Even crazier, they could tear out two pages to soak in milk for breakfast. As long as the book isn't under special legal protection, others can only call them wasteful; they can't stop them.

So why does the situation suddenly change when the buyer is an AI company? It's as if after Anthropic or Amazon pays, even though the book nominally belongs to them, they are actually obligated to preserve it for all of humanity.

This kind of debate has happened before among humans themselves.

In 2010, Christie's auctioned a Book of Hours made in France in the 1460s. It had gold leaf decoration, exquisite calligraphy, and 17 full-page illustrations, and was largely intact at auction, finally selling for £25,000.

A few years later, Elaine Treharne, a Stanford University scholar of medieval literature, acquired the remnants of this book. Originally 254 pages, only 7 remained.

She later discovered that after the Christie's auction, the buyer—a German antiquarian bookseller—deliberately disassembled the book. The pages were sold off one by one, with ordinary pages fetching a few hundred dollars and beautiful illustrations selling for more.

But the antiquarian's subsequent defense was quite interesting:

When Christie's held the public auction, museums could have bought it, large collecting institutions could have bought it, wealthy collectors could have bought it. Everyone saw it. No one was willing to pay a higher price in the end; I paid for it and bought it. Why is it only after the transaction concludes that you suddenly tell me this thing actually belongs to all humanity and I must preserve it for you?

Well, the words might be crude, but the reasoning is sound—and the more you think about it, the more intriguing it becomes.

The same applies to AI companies today.

If a particular book is truly so precious that cutting a single page constitutes an irreparable cultural loss, then where were the museums? The libraries? The public collecting institutions? Those humans who supposedly care more about humanity's precious assets?

Why is it only after the AI company actually pays and buys it that everyone gathers around to say: You are not allowed to destroy this; it belongs to all of humanity.

You didn't buy it, yet you want to control what I do with it after I buy it. Isn't that a bit excessive?

Of course, this doesn't mean that AI companies destroying valuable old books is praiseworthy. "Having the legal right to do so" and our belief that "doing so is awful" can absolutely coexist.

But if a type of book is truly important enough that its owner shouldn't be allowed to dispose of it freely, then what really needs to change probably isn't the moral standards of AI companies.

03

To clarify the logic earlier, the language was a bit extreme.

From the perspective of humanity's common interest, AI companies' large-scale procurement of second-hand books, especially their focus on obscure, out-of-print, and small-language books, is indeed a cause for concern.

As mentioned, Anthropic buying a few million books a year isn't much in the context of the entire second-hand book market. However, a procurement volume that's a drop in the bucket for the overall market can still cause large-scale depletion if heavily concentrated on scarce titles within that niche.

So, "AI is buying up all the second-hand books" is an exaggeration, but "AI companies' procurement may be concentrated on consuming some already fragile book stocks" is not groundless.

The problem is, don't expect AI companies to be the saints in this scenario.

They certainly can be.

But "it would be nice if they did" and "society can only rely on them to do so" are two different things.

Companies are, after all, companies. Today, a certain company might think protecting rare books is important and be willing to spend more on non-destructive scanning. Tomorrow, with a different manager who finds this process 30% more expensive, they could very well cut it.

Relying on commercial companies to long-term safeguard cultural heritage for society based on conscience is a bit like expecting Sun Wukong to put the tightening band on his own head and recite the incantation himself—of course, it's great if he occasionally has an epiphany, but system design better not count on that kind of miracle.

If AI companies' procurement truly begins to threaten books with public value, then this is no longer a question of "whether a company has good ethics," but rather a question of public interest.

Since it involves public interest, the solution should also come from public rules.

For example, we could establish a more comprehensive registration system for rare book titles and editions. If a book is confirmed to have extremely few surviving copies and its physical form possesses clear historical, edition-specific, or artifact value, it enters a protected list. AI companies could buy it but not perform destructive scanning. Or, after purchase, they must prioritize non-destructive digitization.

We could also give libraries, museums, and public collecting institutions some form of right of first refusal. When commercial procurement systems detect suspected unique copies or editions with extremely few surviving copies, they first submit the information to public institutions, giving them a certain period to decide whether to collect them. If the institutions don't want it, then you can buy and process it.

Second-hand book platforms aren't entirely without responsibility either. If algorithms today can accurately detect "only three copies of this 1987 book remain online," platforms can certainly trigger alerts when the same book is suddenly subject to concentrated bulk purchasing by multiple large buyers.

We could even require large-scale procurers to bear some informational obligations. If you're destructively scanning millions of books a year, you should at least record what editions you processed, which books were exceptionally scarce, and contribute this data to public bibliographic systems.

Which of these solutions is best is debatable; perhaps none will be used in the end. But at the very least, there needs to be a rule.

04

Having said all this, there might be one final question left.

When people furiously accuse AI companies of "destroying humanity's books" and "swallowing human knowledge," what are they really accusing? When seeing books being cut open and fed into scanners, where does that intense discomfort and fear truly come from?

The increasingly ambivalent feelings humans now harbor towards AI itself cannot be ignored. People are becoming more reliant on AI, while simultaneously becoming more afraid of it, even growing to hate it.

A Pew survey this year shows the proportion of U.S. adults using ChatGPT has risen from 18% in 2023 to 44% in 2026, with 38% of the employed population already using AI for work tasks. Yet simultaneously, 63% of Americans believe AI is developing too quickly, and 40% expect it to have negative impacts on society in the future.

This contradiction is particularly evident among young people.

A Gallup survey this year of Americans aged 14 to 29 found that 51% use generative AI at least once a week, with 22% using it daily. But compared to a year ago, their excitement about AI dropped 14 percentage points, hope dropped 9 points, while anger rose 9 points.

Now, 31% explicitly state they feel angry about AI, and 42% feel anxious.

Even more interestingly, even among those who use AI daily, excitement and hope have significantly declined compared to last year.

In other words, using it more hasn't automatically led to liking it more.

What happened at American university graduation ceremonies this year is almost a live demonstration of this sentiment.

In May, when former Google CEO Eric Schmidt mentioned at Arizona State University's commencement that AI would enter all industries, the audience booed. Similar scenes subsequently occurred at multiple universities. One survey claimed about 70% of college students worry AI threatens their employment prospects.

This is likely the real situation many face with AI today. Not using it risks falling behind, but the more they use it, the clearer they become about the potential consequences.

The negative associations triggered by "book cutting" are almost inevitable—machines are growing by consuming everything humanity has created in the past.

And then they will come back to take your job.

Thus, copyright anxiety, employment anxiety, distrust of tech companies, and even that fear and dread of "is humanity in the process of creating its own replacement"—all are triggered by the visual symbol of "book cutting."

This also explains why the public sentiment surrounding this issue has far outpaced the actual losses.

That's not just a book being cut open; that is me being sliced by the knife, and then turning into useless garbage.

热门币种推荐

相关问答

QWhat is the main argument of the article regarding the 'AI book-burning' controversy?

AThe article argues that the public outrage over 'AI book-burning' is largely based on its symbolic meaning rather than the actual scale of harm. It points out that the number of books destroyed by AI companies is small relative to the global second-hand book market, and that the knowledge within the books is preserved digitally, not erased. The core issue is misdirected anger stemming from broader anxieties about AI's societal impact.

QAccording to the article, how does the scale of Anthropic's book scanning compare to the broader second-hand book market?

AThe article states that Anthropic scanned millions of books over about a year. However, it compares this to a single large second-hand book dealer, World of Books, which sold 31 million books in 2024 alone and maintains millions in inventory. Therefore, Anthropic's activity represents just a few months' business for one major dealer, indicating it is not depleting the global market.

QWhat point does the article make about the destruction of rare or out-of-print books by AI companies?

AThe article argues that while the destruction of rare books is concerning, responsibility should not fall solely on AI companies. It uses the analogy of a rare manuscript being legally purchased and dismembered by a human collector. If a book is truly a cultural treasure, public institutions like museums or libraries had the chance to acquire it first. The problem highlights a lack of public rules, not just corporate ethics.

QWhat potential solutions does the article suggest for protecting culturally valuable books from destructive scanning?

AThe article suggests establishing public rules rather than relying on corporate goodwill. Solutions could include: creating a registry of rare books that require non-destructive scanning; granting public institutions a right of first refusal for rare books; requiring platforms to flag bulk purchases of scarce titles; and mandating that large-scale scanners contribute data to public bibliographic systems.

QHow does the article explain the intense public emotional reaction to images of books being cut and scanned?

AThe article explains that the visceral reaction stems from deeper, widespread anxieties about AI, not just the physical destruction. It cites surveys showing rising use of AI alongside growing fear, anger, and anxiety about its negative impacts on jobs and society. The image of books being 'consumed' by machines becomes a powerful symbol for fears of human creativity being digested to create a potential replacement for humans themselves.

你可能也喜欢

美国商品期货交易委员会主席:若CLARITY法案未通过,该机构将继续推进加密货币监管工作

美国商品期货交易委员会(CFTC)主席迈克尔·塞利格表示,即使《数字资产监管清晰法》(CLARITY)未获国会通过,该机构也将继续推进加密货币监管工作。塞利格称,此举旨在帮助特朗普总统兑现其承诺。他已指示工作人员制定新规,允许注册及未注册机构提供杠杆或保证金加密货币交易,并研究对开发者的保护措施。 目前,CLARITY法案的审议因美国参议院休会而暂停,预计9月恢复后将进行程序投票。该法案在参议院需获得60票才能通过,之后将送回众议院,并可能提交总统签署。 塞利格的言论是在他与特朗普及加密行业领袖于白宫会晤后一天发表的。特朗普曾呼吁国会通过一项“公平版本”的CLARITY法案,以使美国在相关领域领先于中国。然而,国会许多民主党人要求加强法案中的道德条款,特别是涉及特朗普家族加密投资的问题。法案能否获得足够支持尚不确定。 与此同时,美国证券交易委员会(SEC)也于本周发布了数字资产监管的拟议规则,旨在为加密公司提供安全港待遇。 塞利格目前是CFTC唯一获参议院批准的领导层委员,自去年12月以来一直主导该机构议程。此外,CFTC咨询委员会在会议上还讨论了人工智能和预测市场相关议题,塞利格声称CFTC对预测市场拥有“专属管辖权”。

cryptonews.ru1小时前

美国商品期货交易委员会主席:若CLARITY法案未通过,该机构将继续推进加密货币监管工作

cryptonews.ru1小时前

研究人员将Rust语言arrayref库遭黑客入侵事件与朝鲜黑客联系起来

网络安全公司Wiz的研究人员将Rust库arrayref遭供应链攻击事件归因于朝鲜黑客。该恶意更新在构建脚本中隐藏了窃取凭证的后门,任何在周四编译了相关项目的用户都可能中招,导致计算机和敏感信息泄露。 研究人员指出,攻击载荷中的控制频道与朝鲜黑客组织Sapphire Sleet(又名UNC1069)关联的Mastra行动相同,且使用了相同的IP地址和托管服务商Hostwinds。攻击手法隐蔽,仅在arrayref、internment和append-only-vec三个流行Rust包的依赖列表中,添加了一个拼写错误的恶意依赖项“proc-macro1”(仿冒合法库proc-macro2)。该恶意库内含真实代码能通过编译,但构建脚本会在编译时触发攻击,窃取Chrome、Brave、Edge等浏览器的保存密码,并在Windows、Mac和Linux系统上建立持久化。 据安全公司Aikido评估,这是按下载量计最大的Rust库泄露事件。arrayref被广泛用于Solana和以太坊相关工具,累计下载约2.44亿次。漏洞存在约86分钟后被撤销。Rust官方团队认为原开发者账户可能被盗用,而非其本人恶意行为。 报告同时提及,亚马逊和TRM Labs均指出近期多起npm库攻击与朝鲜有关联,且朝鲜黑客在2026年4月以来盗取了约5.77亿美元加密货币,占同期加密货币黑客攻击总额的76%。其常用手法包括假借工作机会诱骗开发者安装恶意依赖。

cryptonews.ru1小时前

研究人员将Rust语言arrayref库遭黑客入侵事件与朝鲜黑客联系起来

cryptonews.ru1小时前

交易

现货

热门文章

美股TradFi:传统金融在AI IPO浪潮下的稳健锚点

2026年,美股IPO市场重回高热度。本文梳理即将上线或受关注的热门赛道龙头,分析具备投资潜力的交易标的及其逻辑,并探讨宏观趋势与相关风险。

2.8k人学过发布于 2026.07.08更新于 2026.07.08

美股TradFi:传统金融在AI IPO浪潮下的稳健锚点

相关讨论

欢迎来到HTX社区。在这里,您可以了解最新的平台发展动态并获得专业的市场意见。以下是用户对AI(AI)币价的意见。

活动图片