Claude's New Text Watermark Changes What 'AI-Written' Means
Anthropic now embeds an invisible claude text watermark in new model output under the EU AI Act. Here is what it covers, what it proves, and what it doesn't.
Anthropic has started embedding a claude text watermark in output from its newest models, a change tied directly to the EU AI Act rather than a voluntary feature rollout. The company published the details in a support article explaining how it is meeting its obligations under Article 50(2) of the Act's Code of Practice on Transparency of AI-Generated Content. If you generate text, code, or images with Claude and pass that output along without editing it heavily, it now carries a mark you cannot see and, in most cases, cannot remove.
This is worth understanding in some detail, because the mechanism is easy to describe wrong. It is not a disclaimer stapled to the end of a document, and it does not mean every piece of AI-assisted work is now automatically flagged. It is closer to a forensic signal: present under some conditions, absent under others, and not conclusive proof of anything on its own.
What the watermark actually is
Anthropic uses two separate techniques, not one. For plain text, Claude weaves what the company describes as an imperceptible watermark directly into the output itself. It survives copy and paste, which is the main design goal, since most AI-generated text gets pasted into an email, a doc, or a CMS before anyone reads it. It can degrade or disappear under heavy editing, so a first draft carries the mark more reliably than a heavily rewritten final version.
For images in supported formats (SVG, PNG, and JPG), Claude attaches signed provenance metadata instead, following the C2PA standard, the same Coalition for Content Provenance and Authenticity framework that camera makers and news organizations have been adopting for photo authenticity. That metadata is separate from the text watermark mechanism and travels with the file rather than the pixels.
Which products and models are covered
The cutoff is specific: Claude models launched in the EU on or after August 2, 2026 support machine-readable marking starting at launch. Anthropic says it is working to backfill marking support into models released before that date, but as of now, older model output is not reliably marked.
Coverage spans the Claude Platform API, the Claude consumer app, Claude Code, Claude Cowork, and Claude Tag, plus the same models running through third-party infrastructure, AWS, Google Cloud, and Microsoft Foundry, wherever those partners support it. In practice, that means the marking behavior is not something you opt into or configure. If you are running a new-enough Claude model through any of those surfaces, marking is already active in the background.
What a detected mark does and doesn't prove
Anthropic is explicit about the limits here, and the limits matter more than the headline feature. A detected mark means content "may have been processed by Claude." It does not confirm that Claude was the sole or original author, and it does not establish full provenance for content that passed through multiple tools or multiple rounds of editing.
The reverse case matters just as much. An undetected mark does not mean content wasn't AI-generated. Anthropic lists several reasons a mark can be missing even when Claude produced the text: the model predates the August 2026 cutoff, the content was heavily edited after generation, the passage is too short to carry a reliable signal, metadata was stripped during file conversion, or the file type isn't one of the three currently supported for provenance metadata. The company's own framing is that marks provide "important signals about content" while remaining "not fully conclusive" about where that content came from.
That two-sided caveat is the actual news here. It means the mark shifts the burden of proof without fully resolving the underlying question of AI content provenance. Detection tools built on top of this signal will produce true positives, false negatives, and edge cases that require judgment, not a clean yes-or-no answer.
What this means if you build with Claude
For anyone running Claude in production, whether for content, code, or client deliverables, three things follow directly from the mechanism as described.
First, a habit of manually disclosing AI use in a caption or footer no longer controls the full picture. The content can now carry its own signal independent of what you say about it, in either direction: your careful disclosure doesn't create the mark, and your silence doesn't prevent one from already being there in a first draft you later heavily rewrote.
Second, editing depth is now functionally relevant to provenance, not just quality. Output you paste through with light touch-ups is more likely to still carry a detectable claude text watermark than output you've substantially rewritten. If your workflow already involves a real editing pass before anything ships, that pass is doing double duty now.
Third, this is a live instance of a broader pattern worth watching: platforms are moving from trusting the person's disclosure to making the content argue its own case, the same shift that pushed currency printers toward embedded security features once photocopiers got good enough to make disclosure-by-honesty unreliable. Organizations that haven't yet written a policy on how they handle AI-generated content, in code or in text, are going to find that question harder to defer. Oracle's own team recently landed on two contradictory answers to a similar question inside one company, which is a useful reference point for anyone drafting a policy of their own.
The EU AI Act context
Article 50(2) of the EU AI Act's Code of Practice on Transparency requires providers of AI systems that generate synthetic content to mark that output in a machine-readable format, detectable as artificially generated. Anthropic's watermarking rollout is a direct compliance response to that requirement, not a feature Anthropic chose to ship independently of regulation. The August 2, 2026 date lines up with when Anthropic's new-model marking support went live, and the company's public framing treats this as the first phase of a rollout that will eventually extend backward to older models.
The support article's own caveats are the most useful part of the announcement, more useful than the mechanism itself. A watermark that admits its own limits, in both directions, is a more honest signal than one presented as definitive, and it is the version operators actually need to plan around.
Source: Anthropic, AI-generated content marking under the EU AI Act, accessed August 2026.
Join the newsletter
AI workflows and systems, straight to your inbox.
No spam. Unsubscribe anytime.