FISTA Solutions does not load Google Analytics until you accept. Rejecting keeps optional analytics off. Read the Cookie Policy.

All field notes

Glossary · 5 minute read

What Is AI Watermarking? Marking Generated Content Explained

AI watermarking embeds a statistically detectable signal into generated content so that its origin can later be established. Image watermarking is reasonably robust; text watermarking degrades quickly under paraphrasing and editing. Detection is probabilistic, which limits what it can support in consequential decisions.

By FISTA Solutions· AI-Native Engineering Team·
What Is AI Watermarking? Marking Generated Content Explained article cover

Watermarking is frequently proposed as the answer to identifying AI-generated content, and it is a genuinely useful technique with limits that matter for anyone considering it as policy. The gap between what it can do and what people assume it can do is wide enough to cause real harm, particularly in academic and employment settings. This explainer covers both sides. It complements what is content provenance and ai governance framework, and reflects FISTA Solutions' approach in AI enablement delivery. This article is general guidance, not legal advice.

How does text watermarking work?

By influencing token selection during generation. At each step the vocabulary is pseudorandomly partitioned using a secret key, and generation is biased toward one partition. The resulting text reads naturally, and across enough tokens the bias is statistically detectable by someone holding the key.

The detection is a statistical test, not a flag lookup, which is the source of most of its limitations.

PropertyText watermarkingImage watermarkingProvenance metadata
Survives editingPoorlyModeratelyBreaks, detectably
Survives paraphraseNoN/AN/A
Works on short contentNoYesYes
Proves presenceStatisticallyYesCryptographically
Proves absence means humanNoNoNo
Requires provider adoptionYesYesYes

Why are text watermarks fragile?

Because the signal lives in the specific tokens chosen, and any process that replaces tokens removes it. Paraphrasing removes most of it. Translation removes nearly all. Passing the text through a second model to rewrite it removes it almost completely, and that takes one prompt.

Length compounds the problem: the statistical test needs enough tokens, so short passages — a paragraph, a social post, an answer — cannot be detected reliably even unmodified.

How does image watermarking differ?

More favourably. Signals embedded across an image's frequency domain survive resizing, compression, and cropping better than token-level marks survive editing, because the transformations people apply to images are less destructive to the signal than rewriting is to text.

It is not unbreakable — targeted removal works — but it is meaningfully more robust for ordinary handling.

What can detection actually prove?

That a watermark is present, probabilistically, with a confidence that depends on length and modification. That is a real signal and a limited one.

What it cannot prove is absence of AI generation. Most models do not watermark, open-weight models can be run without any watermarking, and watermarks can be stripped. A negative result is evidence of nothing at all, and treating it as evidence of human authorship has already produced serious unfairness in academic settings.

Where does this go wrong in practice?

When detection output drives a consequential decision about a person. Academic misconduct proceedings and employment decisions based on detector output have repeatedly been shown to be unreliable, with false positives concentrated among non-native writers and formulaic styles.

Any policy built on detection needs to state explicitly that a detector result is not evidence on its own, and to require corroboration. Most policies do not.

What is the alternative?

Provenance metadata: cryptographically signed records of how content was created and modified, attached to the file and verifiable by anyone. Industry standards exist and are being adopted by camera manufacturers, editing software, and some model providers.

The logic is inverted and stronger: rather than inferring generation from absence of a human signal, provenance asserts and verifies what actually happened. Stripping the metadata is possible and is itself detectable as a gap. See what is content provenance.

What should organisations do?

Mark their own generated content where disclosure is required or useful, adopt provenance standards for content where origin matters, and avoid building any policy that treats a detector's output as proof. For internal governance, recording what was generated at the point of generation is far more reliable than trying to detect it afterwards.

What should you do first?

Check whether any of your policies rely on AI detection to make a decision about a person. If so, that policy needs revising, because the technique does not support the weight being placed on it, and the resulting decisions are difficult to defend.

What about regulatory requirements?

Several jurisdictions have introduced or proposed obligations to disclose that content is machine-generated, particularly for synthetic media depicting real people. Those obligations generally require marking at generation rather than detection afterwards, which aligns with where the technique actually works.

Organisations generating public-facing content should therefore plan for marking and provenance as a production requirement rather than treating it as an optional signal, and should track how the requirements differ by market since they are not converging quickly.

How does this interact with internal AI use?

Internally the problem is usually simpler, because the organisation controls generation. Recording at the point of creation which content was model-generated, by what system, with what prompt and inputs, gives a complete and reliable record that no detection method could match.

That record also serves review, audit, and incident investigation, which makes it worth building for reasons beyond disclosure. Teams that defer it end up with a body of content whose origin nobody can establish.

How FISTA Solutions helps

FISTA Solutions records generation provenance at the point of creation rather than relying on later detection, adopts provenance metadata standards where content origin matters, and advises against policies that treat detector output as evidence about individuals, through AI enablement, AI agents, and forward deployed engineers. The record behind the approach is 150+ projects for 50+ companies with 99.9% uptime.

To handle AI content provenance defensibly, message FISTA on WhatsApp, or read ai governance framework.

Share-ready article cover

Download the generated social format.

Download cover

Clear answers

Questions raised by this field note.

Straightforward guidance for evaluating scope, fit, and the next step.

01How does text watermarking work?

By biasing token selection during generation toward a pseudorandom subset, creating a statistical pattern detectable by someone holding the key. The text reads normally, and the pattern is measurable across a sufficient span of tokens.

02Why are text watermarks fragile?

Because the signal lives in token choices, and paraphrasing, translating, or editing replaces those choices. Passing watermarked text through another model removes the mark almost entirely, and short passages carry too few tokens to detect reliably.

03How reliable is detection?

Probabilistic, and dependent on length and modification. A long unedited passage detects reliably; a short or edited one does not. Every detector has false positives, which matters enormously when the output is used to accuse someone.

04Does absence of a watermark mean human-written?

No. Most models do not watermark, open-weight models can be run without it, and any watermark can be stripped. Absence is evidence of nothing, and treating it as evidence of human authorship is a common and damaging error.

05What works better than detection?

Provenance metadata: cryptographically signed records of how content was created and edited, attached to the file and verifiable. It proves what is present rather than inferring what is absent, which is a fundamentally stronger position.

Start with the hard problem

Need the outcome owned, not merely analyzed?

Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.

Start a project