
We are in the middle of a curious artificial intelligence arms race. AI companies are building increasingly capable large language models (LLMs), regulators are demanding ways of identifying their output, researchers are developing watermarks and, almost immediately, somebody develops a tool to remove them.
It is tempting to think this is a technical problem. It is not. At least not entirely.
The most widespread use of generative AI today is remarkably mundane. People use LLMs to write ‘stuff’. Students write essays. Journalists draft articles. Authors develop manuscripts. Employees prepare reports and write emails. Scientists polish papers. Medical writers edit documents. People in almost every profession involving written communication are experimenting with LLMs. Sometimes it is entirely harmless and potentially beneficial. Sometimes. The key issue is not necessarily whether AI was involved in writing your essay, report or scientific paper. It is what the AI was used for, what happened to its output afterwards and, ultimately, who remains responsible for what was produced.
That distinction is becoming particularly important in science.
Knowing who, or what, wrote what
Scientific publishing depends upon accountability. Authors are not simply claiming ownership of words and their underlying ideas. They are accepting responsibility for the accuracy, integrity, originality and provenance of the work. That creates a problem if an LLM has contributed substantially to a manuscript or scientific report. A machine cannot take responsibility for a scientific claim. It cannot defend a statistical interpretation at peer review. It cannot explain why one reference was selected over another. It certainly cannot accept responsibility for a fabricated citation or an overoptimistic conclusion.
The International Committee of Medical Journal Editors (ICMJE) places responsibility explicitly on human authors when AI-assisted technologies are used. AI tools cannot be authors because they cannot accept responsibility for accuracy, integrity or originality [1][2]. JAMA has similarly adopted a pragmatic approach, allowing specified AI-assisted uses while insisting that authors remain accountable for the final work [3][4].
So far, so sensible. But there is an awkward assumption sitting underneath much of the current debate about AI in scientific publishing: that if authors use an AI system, they will happily report what, when and how they used it. Many journals now require authors to declare whether and how generative AI has been used, and responsible authors can provide a transparent account of what the technology contributed [1][5][6]. But disclosure is only as reliable as the person doing the disclosing.
The problem is compounded by the absence of a universally agreed boundary between acceptable assistance and use that requires disclosure. Journal policies differ. Authors may not recognise that an apparently minor intervention constitutes meaningful AI use. A recent analysis of guidance from leading publishers and journals found substantial variation in what uses were permitted and when disclosure was required [7].
There is also evidence that what authors tell journals may not represent the full extent of AI use. A bibliometric study of articles published in medical education journals found relatively few explicit AI-use disclosures, raising questions about underreporting and the meaningfulness of some disclosure statements [8]. Similar concerns have emerged in bioethics [9].
This creates an uncomfortable problem. If the integrity of the scientific record depends upon authors voluntarily declaring their use of AI, we cannot know how complete those declarations are. And that raises a more difficult question. How do we know that AI has been used when the author does not tell us?
It is this question that has driven the development of AI detection and watermarking. But there is an important distinction. The purpose of detection should not be to establish that an author has cheated. AI use is not synonymous with misconduct. Contemporary publication guidance increasingly recognises legitimate uses ranging from language improvement and editing to assistance with aspects of research and scientific communication [1][5][6]. Rather, detection is being proposed as a means of establishing provenance when disclosure is absent or uncertain. And that is where things become complicated.
Detection, provenance and accountability are not the same thing
The terms AI detection and provenance are often used as though they mean the same thing. They do not.
These are three very different questions.
A watermark might provide evidence that text is statistically consistent with output from a particular AI system. It does not necessarily tell us who used the system, why they used it, how much of the intellectual work came from it or whether the resulting work is scientifically sound. This distinction is critical. Imagine a manuscript paragraph that has been generated entirely by an LLM. That is one thing. Now imagine that a scientist writes the paragraph themselves and asks an LLM to correct the grammar. Then imagine the scientist asks the LLM to restructure the argument. Then imagine the scientist provides the scientific content but asks the LLM to make the prose clearer. Then imagine that the scientist translates the paragraph into another language and asks another AI system to translate it back. AI may have been involved in every case. But these are not ethically equivalent activities.
The question therefore cannot simply be: Was AI involved? It has to be: What did AI do, what did the human do and who accepts responsibility for the final result? That is a much harder question for a watermark to address.
AI watermarks
Traditional watermarks are visible or embedded physical or digital features that identify something as authentic or originating from a particular source. AI watermarks are rather different.
For text, the watermark can be a statistical signature hidden in the model's choice of words. LLMs generate text token by token, selecting from a probability distribution of possible next tokens. A watermark can subtly alter those choices so that, across a sufficiently long passage, the resulting sequence contains a detectable mathematical pattern [10][11]. Your average reader would see nothing unusual. A detection system, however, can look for the statistical fingerprint.
Google's SynthID-Text is probably the best-developed example. Rather than inserting visible characters, it modifies the sampling process used to select tokens. Research published in Nature demonstrated that the approach could preserve text quality while producing a detectable signal, including in an assessment involving almost 20 million Gemini responses [12]. That is an impressive technical achievement. But even Google makes clear that watermarking is not a complete solution. The signal can become much weaker following substantial rewriting or translation, and watermarking depends upon the AI provider actually implementing it [12][13]. Anthropic has adopted a similar approach which it has deployed in newer models of Claude. Its watermark is designed to be imperceptible to readers and does not add characters or hidden text. Detection provides an assessment of the probability with which Claude has been used in generating any text [14]. It cannot provide definitive proof however that Claude did or didn’t write the text. That distinction is important.
This is the crucial point. A watermark does not say: “This was written by AI.” It says something closer to: “The statistical characteristics of this text are sufficiently consistent with output from this particular AI system.” That is useful. But it is not proof of authorship, intent or misconduct.
AI-generated is not the same as AI-assisted
This distinction deserves more attention because it creates one of the fundamental weaknesses of watermarking. Suppose an author asks an LLM:
“Write me a 1,500-word discussion section about the limitations of this clinical trial.”
A watermark may have a substantial amount of text in which to establish a signal. Now suppose the author writes the discussion themselves and asks the LLM:
“Correct the grammar and improve readability without changing the meaning.”
There may be much less opportunity for any introduced watermark to survive. Google's own description of SynthID-Text notes that detection confidence can be reduced when text is thoroughly rewritten or translated. It also notes that factual prompts can provide fewer opportunities for the watermarking process to modify token selection without affecting accuracy [13].
So, the technology can produce a rather counterintuitive result. The more extensively AI generates the text, the more opportunity there may be for a watermark to survive. But the more lightly AI assists a human writer, the less reliable detection may become. This matters enormously in scientific publishing because legitimate AI use may frequently involve editing, translation, restructuring and language improvement rather than wholesale generation.
The boundary that matters ethically is therefore not simply human versus machine. It is assistance versus abdication of responsibility.
Images offer more room for watermarking
The situation is somewhat different for images, audio and video. These files contain enormous amounts of information in which a signal can potentially be distributed without being perceptible to the human eye (or ear). Watermarks can therefore be designed to survive some transformations such as compression, cropping or modest editing.
Text is much more fragile. Change a word and you may change the meaning, tone or accuracy of a sentence. Change enough words and the statistical signature can disappear. Change the text sufficiently and you may also change the author's voice. There is another important distinction. For images, audio, video and increasingly documents, provenance information can also be attached as metadata.
This is where technologies such as the Coalition for Content Provenance and Authenticity, or C2PA, become interesting. C2PA is not simply another watermarking system. It provides a technical framework for recording information about the origin and history of digital content, including its creation and subsequent modifications, through cryptographically bound Content Credentials [10][11].
In other words, watermarking asks:
“Does this content contain a signal associated with an AI system?”
Provenance asks:
“Can we establish something about where this content came from and what happened to it?”
Neither is perfect. Metadata can be removed. Files can be transformed. Content can move between systems that support provenance and systems that do not. But provenance may ultimately be more useful than trying to determine authorship from the final pixels or words alone.
The European experiment
The European Union has determined that transparency around AI-generated content is sufficiently important to warrant regulation. Article 50 of the EU AI Act introduces transparency obligations relating to AI-generated and manipulated content, with the relevant provisions becoming applicable from 2 August 2026 [15]. But it is important not to reduce this to “Europe has legislated for AI watermarks”. The requirements are more nuanced.
The Act includes obligations concerning machine-readable marking of certain AI-generated content and transparency around certain AI-generated or manipulated material. It also distinguishes between different circumstances, including AI systems used in an assistive capacity for standard editing [15]. That distinction is important. It recognises, at least in principle, that not every interaction between a human and an AI system is equivalent to handing authorship to a machine.
The legislation therefore exposes an uncomfortable problem. Technology does not necessarily become reliable simply because legislation requires transparency. Different providers use different approaches. Detection remains probabilistic. Watermarks can be weakened by transformation. Provenance systems depend upon adoption. And content can pass through multiple systems before it reaches its final audience. The regulatory requirement and the technical reality are therefore not quite the same thing.
China has taken a somewhat different but equally interventionist approach. Since September 2025 it has required explicit and implicit labelling of AI-generated synthetic content, including text, images, audio and video. Metadata-based identification is mandatory in specified circumstances, while digital watermarking is encouraged [16]. The United States currently has no equivalent comprehensive federal requirement covering AI-generated content, although individual states and sectors have introduced their own requirements and policies. California, for example, has introduced extensive requirements around disclosure and responsible generative AI use in government and judicial settings [17].
The watermark arms race
The biggest problem with text watermarking is that it is not a padlock. It is a statistical signal. And statistical signals have patterns that can themselves be detected and countered. Research has shown that watermark robustness varies according to how generated text is subsequently edited, paraphrased or mixed with human writing [18]. Substantial rewriting can weaken or remove the signal. Translation provides another obvious route. So does asking a second LLM to rewrite the first LLM's output.
The watermark arms race therefore becomes rather predictable. The hound learns to detect the watermark. The fox learns how to remove it. The hound learns to recognise the revised version. The fox learns another trick. And meanwhile, the scientist who simply wanted some help improving a report remains blissfully unaware that any of this is happening.
There is also a more fundamental problem. What happens when the AI system does not use a watermark at all? Open-source and locally deployed models make mandatory watermarking much harder to enforce. Google itself acknowledges that generative watermarks cannot solve AI-text detection completely because they depend upon the systems generating the text implementing the technology in the first place [19]. A universal watermark therefore requires something approaching universal cooperation. That is unlikely to be achieved simply by developing a better watermark.
The false-positive problem
Perhaps the more serious concern is not whether determined bad actors manage to remove watermarks. It is what happens when ordinary people are wrongly accused. AI detection systems can produce false positives. Research has demonstrated bias in some AI detectors against non-native English writers, with human-written material incorrectly classified as AI-generated [20].
That should make educators, editors and scientific publishers extremely cautious. Imagine a detector reporting that a manuscript has a 70% probability of being AI-generated. What does that mean? Does it mean the author cheated? Does it mean an LLM wrote the paper? Does it mean the author used AI to improve grammar? Does it mean that the statistical characteristics of the writing happen to resemble those produced by an LLM? Or does it simply mean that the detector has found something it does not understand? The critical distinction is between probability and culpability.
A probability score produced by an algorithm is not a probability that an author committed research misconduct. Those are entirely different propositions. And that distinction becomes even more important when the technology itself cannot reliably distinguish between generation, editing, translation and rewriting.
A scientific manuscript needs a chain of custody
Perhaps the more interesting solution lies elsewhere. Laboratories dealing with physical evidence understand the importance of chain of custody. We need to know where something came from, who handled it and what happened to it. Why should scientific writing be fundamentally different? Imagine the development of a manuscript as a chain:
idea → notes → draft → AI interaction → human editing → statistical analysis → co-author review → submission → peer review → publication
The final manuscript reveals little about that history. Perhaps it should. Rather than asking only whether a manuscript contains an AI watermark, we could ask whether there is credible evidence of how the scientific work was produced: Drafts; Research notes; Data and analysis files; AI-use declarations; Records of substantive AI assistance; Human review; Reference verification; Version histories.
For some types of research, this may eventually provide a much stronger basis for establishing responsibility than attempting to reverse-engineer authorship from the final prose. A watermark may tell us that a machine was involved. A provenance record could tell us about its evolution. Neither, however, can substitute for human responsibility.
AI assistance is not necessarily scientific misconduct
There is a tendency in discussions of AI to imagine two categories of user. The first is the virtuous human scientist, laboriously producing every sentence unaided. The second is the lazy individual pressing a button and submitting whatever comes out. Reality is considerably more complicated.
Experienced scientists and professional writers have always used tools to augment their capabilities. Reference managers, statistical packages, spellcheckers, databases, search engines and computer-assisted editing all reduce cognitive and administrative load.
Psychological research describes this more generally as cognitive offloading. External tools can reduce cognitive demand and improve task performance, although excessive reliance can create its own problems [21]. Automation-bias research makes a similar point: technology can improve performance while simultaneously encouraging people to accept automated outputs too readily [22].
The lesson is not that tools are bad. The lesson is that human oversight remains essential. That is where current scientific publishing guidance is increasingly heading.
The journals may be asking the right question
The most interesting development may therefore be occurring not in watermarking technology, but in publication policy. The current ICMJE recommendations recognise that AI and LLMs are increasingly used in scientific work. They permit AI-assisted use provided that authors remain responsible, disclose relevant use, check the output and address plagiarism and attribution issues [1][2]. JAMA has adopted a similarly pragmatic approach. AI can be used in research, literature searching, content preparation, revision and formatting within its stated requirements, while authors remain responsible for accuracy and integrity [3]. JAMA also specifically warns about the use of AI in generating or managing references because of persistent problems with inaccurate or fabricated citations [3].
That is a much more sophisticated question than simply asking whether AI was involved. It asks:
That is transparency rather than surveillance. It also suggests that authors should retain drafts, notes, source material and other evidence of their work, particularly where AI has played a substantive role. Perhaps that is where scientific publishing should be heading. Not towards proving that every word was written by a human. Towards demonstrating that a human scientist remained intellectually responsible for the work.
Too late to close the gate
There is a legitimate case for watermarking AI-generated material. It can provide useful provenance. It may help identify automated misinformation, mass-generated spam and some forms of academic cheating. It may also create enough friction to discourage malicious behaviour. Watermarks therefore have a role. But they should not be confused with proof. They are probabilistic. They can fail on short passages. They can weaken after rewriting or translation. Different systems use different approaches.
Other models may not carry a watermark at all. Human writing can be misclassified. AI can be used to edit rather than generate. And a watermark cannot tell us whether the scientific content is true. Perhaps this is why we need to stop asking one technology to answer three different questions.
The first may be helped by watermarking. The second may be helped by technologies such as C2PA and other provenance systems. The third cannot be outsourced to either. That is the uncomfortable bit. The temptation is to build better detectors because detectors feel objective. They produce a number for us to hang our hats on. They produce a confidence score. They appear to give us an answer. But a scientific publishing system built around detecting AI use risks asking the wrong question.
Instead of asking: “Can we prove that AI was involved?”
Perhaps we should ask: “Can we demonstrate that the people responsible for this work behaved responsibly?”
That requires something considerably less glamorous than a watermark. It requires education. It requires professional standards. It requires transparency. It requires authors to verify claims rather than admire fluent prose. It requires organisations to establish appropriate boundaries around confidential information. It requires journals to distinguish between assistance and abdication of responsibility. And perhaps it requires a better record of how scientific work comes into existence. The gate is closing and it may already be too late.
For LLMs, watermarking looks less like the final solution and more like the opening move in another technological game of chess. That may be inevitable. But before we legislate our way into a world where every piece of text needs to carry a hidden mathematical fingerprint, we should spend more time defining what good AI-assisted science actually looks like. Because this is not simply an efficiency issue. It is not even primarily a technology issue. It is an issue of scientific ethics.
A watermark can potentially tell us that a machine was involved. It cannot tell us whether a scientist behaved responsibly. And that is the question that matters. Ethics in science cannot be embedded in a watermark. It has to be inherent in everything those of us working in science do.
References

Get our latest news and publications
Sign up to our news letter