
A watermark-removal claim must name the mark it can change
Metadata cleaning and changes to sampled wording operate on different signals. A tool’s removal label does not establish that it defeats a particular text watermark.
A tool can successfully remove metadata while leaving a statistical text watermark untouched. The useful question is which signal it changes, not whether its name contains “watermark remover.”
Anthropic's August 14 mechanism explanation says Claude's text watermark changes the source of randomness used in selecting words. It adds no hidden characters. The company also acknowledges that sufficiently extensive rewriting can remove the signal.
The watermarks-remover repository's August 17 README separates character cleaning, file metadata and statistical text marks. It describes rewriting against the latter as best effort. That is a meaningful limitation, and I would preserve it rather than treat the repository as a claim of guaranteed removal.
The existence of a remover does not prove that the watermark works. Nor does a plausible rewriting method prove that this tool defeats a particular detector. Those conclusions require measurements against the relevant mechanism.
Different signals need different evidence
A hidden character is part of a text representation. File provenance metadata lives in a container or associated record. A pattern in word selection is distributed through the wording itself. They may all be discussed as marks, but changing one does not establish a result against the others.
That distinction matters even in ordinary content maintenance. A workflow cleaning a document for compatibility should be able to state what it changed. It should not advertise an unrelated provenance outcome merely because it removed an invisible character or a file property.
I would inspect the tool's stated scope before interpreting any success message. “Metadata removed” is a claim about metadata. It is not evidence that the content is human-authored, untraceable or free of every machine-readable signal.
Likewise, a detector failing to identify a supported mark is a bounded observation about that detector and input. It should not become a general certificate about the origin of the work.
Rewriting creates a second content problem
If a process changes wording, the output needs review as edited content. Meaning, attribution and factual precision can move even when the person requesting the change wanted only a technical property altered.
Imagine a hypothetical policy summary containing a carefully scoped exception. An aggressive rewrite produces smoother prose but broadens the exception. Whatever happened to a statistical signal, the document has become less accurate.
I would compare the revised claims with their sources and preserve the original when evaluating the change. The acceptance question is whether the document remains faithful and useful, not merely whether a detection score moved.
This is not a removal procedure or a test result. I have not run the repository against Anthropic's detector. The source comparison establishes that the mechanisms differ and that the tool itself qualifies its statistical-removal path.
For an AI-assisted publication, a clear account of how work is made and verified is more durable than a promise to make its origin disappear. Provenance and content quality are related responsibilities, but neither should be reduced to a tool's marketing category.
Require every watermark claim to identify its mechanism and evidence, and treat rewritten wording as content that must earn approval again.


