Anthropic Details How Claude's New Watermarking System Will Work

Anthropic published a blog post Friday to address common questions about its plans to watermark text generated by its AI chatbot Claude. The company explains how the watermarking will function, whether it can be circumvented through editing, and its impact on code generation. This move follows Anthropic's announcement earlier this week that it will implement watermarking to comply with the EU AI Act's Transparency Code, which mandates that AI companies adopt systems capable of identifying AI-generated content. The decision has sparked debate among Claude users, with reactions ranging from criticism to support. For instance, a Reddit user characterized the watermark as a "conspiracy against innocent Claude users," while another argued, "The only reason you wouldn't want this is to lie to people." Additionally, Business Insider reported that "dozens" of users on X have claimed to cancel their Claude subscriptions in response. In its blog post, Anthropic begins with a general overview of watermarking, explaining that when Claude makes "low-stakes choices"—such as selecting between words like "overcast" or "grey" to describe weather—it can embed a pattern in its responses that is "undetectable to the reader, but is detectable to anyone who has a key that encodes it." The company assures that "watermarking does not impact the quality of Claude's output," stating, "To a reader, a watermarked response is indistinguishable from an unwatermarked one." More specifically, Anthropic plans to use the SynthID-Text approach, originally outlined by Google DeepMind in 2024, and will release a watermark detection API. The company also clarifies that watermarking differs from AI detection methods offered by firms like Pangram, which analyze writing for stylistic "tells" (e.g., phrases like "This isn't [X], it's [Y]") to flag AI usage. Anthropic notes, "Picking up on these patterns is fundamentally different from checking for a watermark." Regarding potential circumvention, Anthropic acknowledges that a complete rewrite—replacing every word—could remove the watermark, but "light editing probably won't remove the watermark completely." The company adds, "In the latter case, of course, it's arguable whether the text can any longer be described as AI-generated." As for text that Claude only proofreads or lightly edits, watermark detectability will depend on "the length of the text and how heavily Claude has edited it." If lightly edited, "nearly all the words" would be authored by the human, leaving "very little (if anything) for the watermark to attach to." Code, however, will carry a lighter watermark than standard text. Since Claude must generate functional code, it lacks the freedom to choose among equally valid alternatives, reducing the opportunity to embed watermarks without compromising functionality.

via TechCrunch AI

Related