Connect with us

Digital

How Claude will leave a hidden trail in AI-generated text, according to Anthropic

Invisible system will identify likely Claude involvement without changing how text reads

Published

on

MUMBAI: Anthropic is pulling back the curtain on Claude’s text watermark. The AI company has explained how its new system will identify whether Claude was likely involved in generating a piece of text, while keeping the watermark invisible to readers and avoiding any change to the quality or content of its responses.

The watermark will be introduced in future Claude models as Anthropic moves to comply with requirements under the EU AI Act. The company said it will initially apply the system globally because it does not yet have a reliable way to restrict the technology by region.

Rather than adding a visible mark, hidden characters or extra text, the system works through the way Claude makes choices while generating words. Anthropic said it will not require additional tokens and will have a negligible impact on model speed or cost.

Claude, like other large language models, generates text one word or token at a time. When several words could work equally well in a sentence, the model has some flexibility over which one to choose.

The watermark uses these low-stakes decisions to create a statistical pattern across a passage.

Instead of using an arbitrary source of randomness to make those choices, Claude uses a secret key together with words that appeared earlier in the text. The choices remain effectively random to the reader, but the resulting sequence can later be checked against the key.

If the pattern matches what would be expected from Claude’s watermarking system, a detector can assign a probability that Claude was involved in generating the text.

Anthropic said the method is based on the SynthID-Text approach developed by Google DeepMind and published in Nature in 2024.

Anthropic said watermarking is designed to have no practical impact on the quality, creativity or readability of Claude’s responses.

The system does not force Claude to select unusual words or steer it towards particular phrases. It only operates where there are multiple reasonable choices.

This means a watermarked answer should look and read the same as an unwatermarked one.

The company also cited testing by Google DeepMind, which found no statistically significant difference in user ratings between watermarked and unwatermarked Gemini responses. Human raters in a separate controlled study also found no noticeable difference in quality.

The watermark is less active when a passage contains information for which there is only one correct answer.

For example, if Claude is completing a factual statement and only one term is accurate, there is no meaningful choice for the watermarking system to influence. The same principle applies to technical material where changing a particular word could make the information incorrect.

Code is similarly less heavily watermarked because programming often requires exact terms and structures. An alternative token could alter how a program works or cause it to fail.

There can still be room for watermarking in areas such as code comments, where several words or phrases may communicate the same thing, but Anthropic said the effect on actual code should be negligible.

Anthropic said the watermark is attached only to words Claude chooses.

That makes light editing particularly difficult to detect. If a user gives Claude human-written text and asks it to correct grammar and punctuation, most of the original words remain untouched.

There may therefore be too few new decisions for the watermark to produce a strong enough signal.

The more extensively Claude rewrites a piece, the more opportunities there are for the watermark to attach and the more confident detection can become.

Short passages also present a challenge because there are fewer word choices from which a watermarking pattern can be established.

The watermark is designed to identify likely Claude involvement, not the person behind it.

Anthropic said the watermark contains no identifying information and cannot be used to trace content back to a particular user, organisation or conversation.

It also does not prove that Claude wrote an entire piece. A positive result can only indicate that Claude was likely involved at some point, including through substantial editing.

The system is also specific to Claude. It cannot determine whether text was generated by another AI model, which could use a different watermarking method or key.

The move is primarily about transparency around AI-generated content.

Anthropic said it and several other major AI providers signed the EU Code of Practice on Transparency of AI-Generated Content in July 2026. The framework requires AI providers serving the European market to mark AI-generated content.

The company is applying the watermark globally at launch because it does not yet have a durable way to limit it to particular regions.

Anthropic also plans to offer a watermark detection API, although details of the service are still being worked out.

For supported files such as PNG, JPG and SVG, Claude will use content credentials rather than the text watermark.

The credentials will follow the C2PA open industry standard and will be stored as cryptographically signed information in the file’s metadata. They will indicate that Claude was involved in creating or processing the file without revealing information about the user.

Anthropic acknowledged that watermarking is not impossible to defeat. Light editing may not remove it completely, while rewriting every word can eliminate the signal. But in that situation, the resulting text has arguably been substantially transformed from the original AI-generated material.

For Claude users, the practical change is therefore largely invisible. The text will look the same, cost the same and take roughly the same time to generate. The difference is that, after the fact, there will be a new way to assess whether Claude was likely part of the writing process.

Click to comment

Leave a Reply

Your email address will not be published. Required fields are marked *

Advertisement News18
Advertisement
Advertisement Whtasapp
Advertisement Year Enders

Indian Television Dot Com Pvt Ltd

Signup for news and special offers!

Copyright © 2026 Indian Television Dot Com PVT LTD