Looking Back at Information Hiding in the GenAI Era from Claude's Invisible Watermark: How Can Embedded Watermarks Ensure Losslessness?

marsbit发布于2026-08-17更新于2026-08-17

文章摘要

On August 11, 2026, Anthropic announced that its new Claude models will embed invisible, machine-readable watermarks into generated text, allowing AI-generated content to be traced even after copying and redistribution, without affecting semantic quality or readability. This move follows global initiatives to watermark AI outputs, highlighted by the World Economic Forum and futurist Kevin Kelly. A core challenge is ensuring watermarking does not degrade model performance, leading to the key research goal of "provably lossless" generation. The theoretical foundation lies in provably secure steganography, which requires that watermarked data be indistinguishable from normal data in distribution. Early theoretical work was limited by the inability to precisely sample natural data distributions. The rise of generative AI models around 2018 provided a breakthrough, as these models learn and sample from explicit distributions. Pioneering work by a University of Science and Technology of China (USTC) team proposed frameworks for provably secure steganography using generative models, mapping encrypted messages to drive content generation. Subsequent research expanded the practicality of these techniques. Advancements include public-key steganography for asymmetric extraction, "box-less" extraction where the receiver doesn't need the original model, and "grey-box" scenarios for resource-constrained devices. A key innovation, SyncPool, resolved tokenization ambiguities in LLMs to ens...

On August 11, 2026, Anthropic announced that starting from August 2nd, the newly released version of the Claude model will directly embed invisible, machine-readable watermarks into generated text. This will enable AI-generated content to be identified and traced by software even after being copied and disseminated.

The official documentation states: "When a supported Claude model generates text, it embeds a subtle watermark directly into the text. You will not see it, and it does not change the semantics, quality, or readability of Claude's responses."

This mechanism covers Claude's web interface, API, Claude Code, and all product lines. It takes effect for global users and cannot be disabled. Generated images and other files will also include digital signature and traceability metadata compliant with the C2PA standard.

This is a landmark move by Anthropic following its signing of the EU's Artificial Intelligence Act, Article 50, Transparency Code of Conduct.

"Watermarking" AI-generated content is evolving from corporate initiatives to a global consensus. In June 2025, the World Economic Forum released the "Top 10 Emerging Technologies of 2025" at the Summer Davos in Tianjin, listing "Generative Watermarking" in second place. Futurist Kevin Kelly predicted in "2049: The Possible Next 10,000 Days" that the AI era necessitates "redefining reality" and "adding watermarks or similar markers to distinguish authenticity on AI-generated images and videos."

Figure 1: Summer Davos Forum and Kevin Kelly's Prediction

However, getting watermarks "right" is not easy.

On the day of Anthropic's announcement, users' most concentrated concern was precisely: will watermarking harm the model? Will it reduce generation quality?

"Provable performance losslessness" thus became the core proposition of information hiding research in the GenAI era. To understand this proposition, we must return to its theoretical origin—provably secure steganography.

Provably Secure Steganography: Twenty Years Waiting for Generative Models

Steganography and watermarking both belong to information hiding. Traditional steganography has long remained at the level of "empirical security": using steganalysis algorithms to test security strength, but unable to provide mathematical guarantees. Provably secure steganography requires strict proof: the distribution of data carrying secrets is indistinguishable from that of normal data.

This theoretical line has a long history.

In 1949, Shannon pointed out the difficulty of establishing a theory for covert communication; in 1998, Cachin introduced relative entropy, providing an information-theoretic security definition; in 2002, Turing Award winner Manuel Blum guided Hopper et al. in establishing the computational security steganography framework, further developing public-key steganography and covert key agreement in 2004.

However, all constructions of provably secure steganography rely on a stringent premise—the carrier distribution must be precisely sampleable. Natural data distributions are uncontrollable, causing the theory to lie dormant for years.

Figure 2: A Brief History of Provably Secure Steganography Development

In 2018, AI generative models brought a turning point: models first learn a distribution, then sample according to it, naturally providing an "explicit distribution or perfect sampler."

A team from the University of Science and Technology of China (USTC) pioneered the idea of generative provably secure steganography internationally. They proposed two frameworks: "black-box sampling" and "compression-reversible sampling" (IWDW 2018, arXiv 2018), reducing steganographic security to cryptographic algorithm security—encrypting a message into pseudo-random ciphertext, using the ciphertext to drive sampling for content generation, and allowing the receiver to recover the ciphertext via inverse sampling.

At that time, generative AI had not yet exploded, and this idea seemed "ahead of its time," initially finding it difficult to gain peer recognition. As the quality of generative models like text and audio improved rapidly, related achievements gradually gained acknowledgment. The International Workshop on Digital Forensics and Watermarking (IWDW) invited the USTC team to deliver a keynote report titled "When Provably Secure Steganography Meets Generative Models."

Figure 3: The Earliest Papers on Generative Provably Secure Steganography

Around 2021, AI-generated data gradually became a major form of online content. Provably secure steganography gained widespread attention globally: Tsinghua University proposed provably secure linguistic steganography based on sample grouping (Findings of ACL 2021); Boston University and Johns Hopkins University proposed the cryptographically secure steganography scheme Meteor for real distributions (ACM CCS 2021); Oxford University proposed perfectly secure steganography based on minimal entropy coupling (ICLR 2023).

The USTC team proposed the DisCop construction based on "distribution replicas" (IEEE S&P 2023), significantly improving embedding payload rate. In the same year, Quanta Magazine listed "perfectly hiding secret information in generated data" as one of the seven international breakthroughs in computer science for the year, heralding the entry of information hiding into the "provable" era.

From Symmetric to Asymmetric, from "Boxed" to "Box-less"

Early generative provably secure steganography schemes were all "symmetric key" systems, requiring the sender and receiver to pre-share a key and relying on white-box extraction, limiting their application scenarios.

Additionally, subword ambiguity could lead to extraction failure. The USTC team pushed the technical route towards public-key steganography, box-less/grey-box scenarios, and eliminated subword ambiguity, making generative hiding practical.

Public-key steganography (IEEE TIFS 2024): Proposed a provably secure public-key steganography scheme combining Elliptic Curve Cryptography (ECC) with generative models and introduced a steganographic key exchange protocol. Solved the problems of steganographic key agreement and "asymmetric" hidden information extraction.

Box-less steganography (IEEE TMM 2026): Traditional methods rely on "white-box" extraction, requiring the receiver to possess the exact same language model as the sender. The team proposed Disreo, achieving "box-less extraction" through token position randomization and output probability reorganization, enabling the receiver to recover the message without accessing the underlying model, providing greater convenience for practical deployment.

Grey-box steganography (ACM CCS 2026): Between white-box and box-less lies the "grey-box" scenario—unequal resources between sender and receiver; the receiver may only have the capability to run a small model (e.g., on mobile devices). SpecStega is based on speculative sampling: using a small model shared by both parties to embed and extract messages, then refining the output with the target large model. This ensures the final stego-text aligns with the large model's output distribution, maintaining high quality and security, while allowing efficient decoding at the receiver end using only the small model. The payload rate is improved over 20 times compared to existing black-box schemes—offering a third path between "boxed" and "box-less" for provably secure steganography.

Resolving subword ambiguity (IEEE TDSC): Steganography based on large language models commonly faces token decoding ambiguity—the same text segment can be tokenized into different subword sequences, causing extraction failure. The team proposed SyncPool, which groups tokens with prefix relationships before embedding, eliminating ambiguity in principle and achieving reliable extraction for provably secure steganography.

From Steganography to Watermarking: Bringing "Provable Losslessness" to Industry

First, AIGC large models learn the distribution of natural data, then sample according to that distribution to generate text, images, audio, video, etc. If one can: embed a watermark in the generated content without affecting the sampling distribution, it means the watermark embedding does not impact generation quality.

Provably secure steganography theoretically guarantees that the distribution of data carrying secrets is indistinguishable from that of normally generated data—meaning embedding information does not alter the model's sampling distribution. This is precisely the definition of "provably generation-quality-lossless watermarking." Watermarking does not require the large capacity of steganography, thus capacity can be traded for robustness.

In the text domain, the provably generation-quality-lossless watermarking system developed by the USTC team is plug-and-play, requiring no modification of model parameters. It supports single-bit robust discrimination and multi-bit model attribution, and has been applied to platforms like the Spark large model and Secure GPT, serving 13,000 developers.

Meanwhile, international research on provably lossless generative text watermarks is also advancing rapidly:

In 2024, teams from the University of Maryland and others' "Unbiased Watermark for Large Language Models" (ICLR 2024 Spotlight) defined unbiased watermarks and provided a general construction.

The same year, Stanford University proposed robust unbiased watermarks resistant to distortion (TMLR 2024).

In 2025, Multi-Channel Unbiased Watermark MCmark (ACL 2025) improved watermark extraction robustness while maintaining strict unbiasedness.

In the image domain, the USTC team proposed Gaussian Shading (CVPR 2024), mapping an encrypted, randomized watermark to Gaussian latent variables indistinguishable from normal generation, then acting on the entire latent space via the diffusion process—training-free, plug-and-play, and provably performance-lossless.

Gaussian Shading++ and T2SMark (NeurIPS 2025) further addressed robustness, generation parameter variation, and generation diversity issues in real-world deployment.

TAG-WM (ICCV 2025) introduced a dual mechanism of "template watermark + information watermark," achieving both tamper localization and ownership verification under the lossless premise.

SemBind (ICML 2026) binds the watermark to image semantics through a semantic masker, resisting black-box forgery attacks.

Figure 4: Provably Lossless Generative Image Watermark Gaussian Shading (CVPR 2024)

From content watermarking to model watermarking. Models themselves are important digital assets, also requiring reliable provenance. Model watermarking always faces one question: will it affect the model's normal use?

Drawing inspiration from the "provably undetectable backdoor" construction proposed by Turing Award winner Shafi Goldwasser et al. at FOCS 2022, the USTC team proposed a provably performance-lossless black-box model watermarking protocol (IEEE TDSC 2026): using unforgeable message authentication codes to construct branch indicators, making the probability of normal users triggering the watermark branch computationally negligible, thereby reducing performance losslessness to cryptographic security.

Thus, "provable losslessness" extends from generated content to the models themselves.

From "Empirical Security" to "Provable Security"

Anthropic's invisible watermark sparked controversy on its launch day: writers worried about attribution, developers worried "will the watermark reduce generation quality?" Industry pioneers chose "engineering fast runs," while the goals of "accurate marking, no harm, withstand rewriting and erasure, and provide definitive evidence" call for solid theoretical support.

From Shannon's conceptual challenge to Blum et al.'s theoretical frameworks, from the first generative provably secure steganography algorithm in 2018 to the Spark model watermark and Claude's watermark, information hiding has taken seventy-seven years to finally upgrade from "empirical security" to "provable security," turning "worrying that embedding watermarks harms the model" into "mathematically proven losslessness"—whether for generated content or the model itself—providing theoretically guaranteed technical support for AI content governance.

Watermarking: Attack and Defense

Designing watermarks is hard; erasing them is easy. Or perhaps not entirely. Designing watermarks requires pursuing both quality losslessness and robustness simultaneously, while attackers also need to erase watermarks under quality constraints.

Without constraints, watermarks are easily erased. Rewriting sentence by sentence to erase watermarks—does it still count as content generated by the original model? Perhaps it would be easier to just use another open-source model for generation.

Like other security fields, attackers and defenders engage in a game under their respective constraints. The existence of both attack and defense keeps a technology direction vibrant.

And "losslessness" is the constraint and pursuit for both sides.

Defenders do not want to sacrifice generation quality by embedding watermarks; attackers also do not want to degrade quality by erasing watermarks.

Absolute security never exists. The significance of defense is to create the greatest possible cost for attack. The essence of attack and defense is a game of cost, and the same holds for watermarking.

This article is from the WeChat public account "New Zhiyuan," author: New Zhiyuan.

热门币种推荐

相关问答

QWhat is the core concept behind 'provably secure steganography' as discussed in the article?

AThe core concept of 'provably secure steganography' is to mathematically prove that the distribution of stego-data (data containing hidden information) is indistinguishable from the distribution of normal, cover data. This ensures that embedding information does not alter the statistical properties of the generated content, providing a theoretical guarantee of security and 'losslessness' in terms of generation quality.

QWhat key advantage do generative AI models provide for provably secure steganography, according to the article?

AGenerative AI models provide a key prerequisite that was historically difficult to achieve: an explicit or perfect sampler of a data distribution. Since these models learn a distribution and then sample from it to generate content, they naturally offer the 'perfect sampler' required for the theoretical constructions of provably secure steganography, enabling the shift from 'empirical security' to 'provable security'.

QName three practical challenges in steganography/watermarking that the research from USTC addressed, as mentioned in the article.

A1. Public Key Steganography: They proposed schemes combining Elliptic Curve Cryptography with generative models to solve asymmetric key management and extraction. 2. Boxless Extraction (Disreo): Enabling message recovery without the receiver needing access to the original generative model. 3. Subword Token Ambiguity (SyncPool): Solving extraction failures caused by the same text being tokenized into different subword sequences by grouping related tokens.

QHow does the concept of 'provable performance losslessness' apply to both generated content and AI models themselves?

AFor generated content (text, images), 'provable performance losslessness' means embedding a watermark does not change the model's sampling distribution, guaranteeing no degradation in generation quality. For AI models, it refers to model watermarking protocols where the probability of a normal user triggering the 'watermark branch' of the model is computationally negligible. This is achieved by using cryptographic constructs like unforgeable message authentication codes, thereby extending the guarantee of 'no performance impact' from the content to the model asset itself.

QAccording to the article, what is the fundamental nature of the battle between watermarking defense and attack?

AThe fundamental nature is a cost博弈 (game of cost/博弈). There is no absolute security. The goal of defense (watermarking) is to impose the highest possible cost on the attacker (e.g., significant quality degradation or computational effort required for removal). Conversely, attackers operate under the constraint of maintaining content quality while removing the watermark. 'Losslessness' is a shared constraint and pursuit for both sides in this ongoing博弈.

你可能也喜欢

债市惊魂,一场回购如何引爆黄金和比特币?

8月18日,美国30年期国债收益率飙升至5.337%,创2007年以来新高。次日,美国财政部紧急宣布,将长端国债流动性支持回购操作的规模至少翻倍,以应对市场流动性紧张。此举旨在买回交易稀薄的老旧债券,释放现金以改善长端市场的交易环境,而非直接印钞。 消息公布后,市场反应剧烈:30年期国债收益率应声急跌,黄金单日暴涨超125美元,比特币亦大幅拉升近8.7%。其核心传导逻辑在于:财政部出手干预,向市场暗示其对长端收益率存在“痛阈”和隐性天花板,导致市场对利率上行的预期减弱。实际利率预期下降降低了持有黄金的机会成本,从而推高金价。同时,市场也强化了“财政主导”的叙事,即高赤字与高债务压力下,政府可能被迫进行更多市场干预,这对黄金等资产构成长期支撑。 比特币的暴涨则叠加了加密市场特有的轧空机制。在宏观利好刺激下,价格快速突破关键点位,触发了大量空头仓位的强制平仓,形成了上涨的自我强化。 此次事件暴露了美国债市在巨额赤字、高额到期债务与“买家罢工”背景下的结构性压力。财政部虽未明确设定收益率上限,但其在关键点位的干预行为,已向市场揭示了政策底线。未来,市场或将反复试探这一底线,而财政部能否有效稳住长端利率,将成为影响黄金、比特币等资产走势的关键。

marsbit32分钟前

债市惊魂,一场回购如何引爆黄金和比特币?

marsbit32分钟前

交易

现货

热门文章

如何购买S

欢迎来到HTX.com!我们已经让购买Sonic(S)变得简单而便捷。跟随我们的逐步指南,放心开始您的加密货币之旅。第一步:创建您的HTX账户使用您的电子邮件、手机号码注册一个免费账户在HTX上。体验无忧的注册过程并解锁所有平台功能。立即注册第二步:前往买币页面,选择您的支付方式信用卡/借记卡购买:使用您的Visa或Mastercard即时购买Sonic(S)。余额购买:使用您HTX账户余额中的资金进行无缝交易。第三方购买:探索诸如Google Pay或Apple Pay等流行支付方法以增加便利性。C2C购买:在HTX平台上直接与其他用户交易。HTX场外交易台(OTC)购买:为大量交易者提供个性化服务和竞争性汇率。第三步:存储您的Sonic(S)购买完您的Sonic(S)后,将其存储在您的HTX账户钱包中。您也可以通过区块链转账将其发送到其他地方或者用于交易其他加密货币。第四步:交易Sonic(S)在HTX的现货市场轻松交易Sonic(S)。访问您的账户,选择您的交易对,执行您的交易,并实时监控。HTX为初学者和经验丰富的交易者提供了友好的用户体验。

3.6k人学过发布于 2025.01.15更新于 2026.08.06

如何购买S

相关讨论

欢迎来到HTX社区。在这里,您可以了解最新的平台发展动态并获得专业的市场意见。以下是用户对S(S)币价的意见。

活动图片