Anthropic has released more details on how it will watermark text generated by its Claude chatbot. The company explained that the watermarking process will create patterns in responses that are undetectable to readers but can be identified by those with the correct key. According to the company, the watermarking does not impact the quality of Claude’s output, and a watermarked response is indistinguishable from an unwatermarked one to the reader.

The company plans to use the SynthID-Text approach, which was outlined by Google DeepMind in 2024, and will also release a watermark detection API. Anthropic emphasized that its watermarking method is different from AI detection tools like Pangram, which look for writing patterns that may indicate AI use. The company clarified that while light editing may not fully remove the watermark, a complete rewrite where every word is replaced would likely eliminate it. However, it noted that such a rewrite could make the text less describable as AI-generated.

Anthropic said the effectiveness of the watermark in edited text depends on the length of the text and how heavily Claude has edited it. If the text is only lightly edited, there will be very little for the watermark to attach to. Code, on the other hand, is expected to have a weaker watermark because the model must create working code and has fewer choices to make. The company also mentioned that other major model developers have signed the same Code of Practice and will implement their own watermarks.

Source: techcrunch