Home Top Stories Watermarking Of AI Text Might Increase The Chances Of Getting AI Hallucinations
Top Stories

Watermarking Of AI Text Might Increase The Chances Of Getting AI Hallucinations

Share
Watermarking Of AI Text Might Increase The Chances Of Getting AI Hallucinations
Share

In today’s column, I examine an interesting twist associated with the latest trend of AI makers opting to include text-based watermarking in their latest versions of generative AI and large language models (LLMs). The twist is that new research suggests that there is a chance of increasing the odds of getting AI hallucinations due to the efforts of the watermarking process. Yes, believe it or not, the internal efforts of devising hidden text-oriented watermarks seem to also come with the potential for generating AI hallucinations.

You might already know that AI hallucinations are the bane of AI makers and all users of generative AI. An AI hallucination is when the AI generates a made-up, ungrounded response to your prompt. For example, you might ask when Halloween takes place each year, and the AI responds that it occurs annually on October 13 (the correct answer is October 31). The AI has produced a confabulation, something that is not based on actual facts. All users of AI need to be alert to the possibility of AI hallucinations. Now, the advent of text-based watermarking could worsen the AI hallucinations predicament by providing a new avenue for producing unwanted and potentially unsafe AI hallucinations.

Let’s talk about it. This analysis of AI breakthroughs is part of my ongoing Forbes column coverage of the latest in AI, including identifying and explaining key AI complexities (see the link here).

Watermarking Is Challenging

First, some foundational aspects of the topic of watermarks (for my in-depth analysis of AI watermarking, see the link here). We are all aware of watermarking when it comes to paper-based materials and likewise for any tangible artifact that exists in a definitive physical form. A dollar bill can contain a watermark, allowing the naked eye to tell whether it is real or counterfeit. Watermarks can also be hidden from visual inspection, requiring some other means to detect the watermark.

Watermarking for digital photographs and graphical images is more readily accomplished than with text since you can embed all sorts of digital ones and zeros that won’t impact the picture, but that can be detected by inspecting the binary representation. It is possible to use sophisticated mathematical algorithms to populate the bits in a manner that almost no one other than someone armed with the algorithm can later detect as being part of a special pattern.

Trying to watermark digital text is a beast of a different kind. Anything that is done to the text will potentially alter the words we see and impact the meaning of the text. If you had a watermarking algorithm that simply said to replace the word “of” with the word “horse”, the resulting text, which is now presumably discernible as AI-written due to the excessive use of the word “horse”, is going to be nonsensical for human use. Likewise, if the watermark consisted of embedding special characters or the use of emojis, you could quickly find those and remove them easily.

Example Of How It Works

An ingenious way to infuse watermarking is to do so by selecting suitable words that can be viably chosen during the AI writing process. Here’s how that works. Envision that AI is generating a response to a prompt, doing so one word at a time. Each word is carefully chosen by the AI. The choice of which word to use is made from several possible words at each step.

Suppose the prompt was asking the AI how to make a ham sandwich. The AI might start assembling the response word-by-word and could have arrived at these choices: “Place a slice of ham onto a bagel and add mustard.” Each word was selected on a one-at-a-time basis, going from the start of the sentence to the end of the sentence.

When the AI got to the word about the bread, in this instance the word selected was “bagel,” but there were several other options available, such as saying “flatbread” (statistical second choice), “wheat bread” (statistical third choice), “white bread” (statistical fourth choice), and other possibilities. Assume that the word “bagel” was the statistically top-ranked choice overall and therefore chosen accordingly.

Aha, in the realm of watermarking, the AI might opt to intentionally choose the second choice rather than the top-ranked choice; thus, the sentence comes out as “Place a slice of ham onto a flatbread and add mustard.” If the AI consistently keeps picking the second choice for many of the words that are being chosen, this becomes a handy pattern for the AI. A human looking at the sentence doesn’t realize that the second choice is being chosen. They see a sentence that looks completely normal.

Detecting The Watermark

I think you can see that this statistical uplift is going to be quite hard to detect by conventional means. Humans are unlikely to discern the watermark by looking for any patterns in the wording. All the sentences are still going to make sense and abide by whatever the topic at hand is. The subtlety of picking the second statistically viable word on numerous occasions is a hidden way of producing the watermark.

How does an authorized detection tool figure out if the watermark is present?

That is the added trickery. The chances of any usual automated detection method ferreting out the watermark are low. It won’t realize what the watermark method is or how to ferret it out. Meanwhile, a detection tool that is built knowing the method can examine the sentences and compare the word choices to the pattern of word choices that the AI would normally make. If the second word choice is consistently being encountered in the examined text, this is a strong indicator that the AI indeed generated that content.

We can make this method much stronger. Instead of always choosing the second choice, the watermark process does something else. Suppose that 50% of the time the second choice is made, 30% of the time the third choice is made, and 20% of the time the fourth choice is made. This makes things even harder for anyone else to crack and find the watermark. An even better method includes having a secret cryptographic key that guides the watermarking process toward the preferred token patterns.

The Specter Of AI Hallucinations

You might naturally assume that when AI is watermarking text, this is accomplished on a “no harm” basis. In other words, though the sentences will differ due to the word choices being made, by and large the resulting sentences will be sensible and accurate. One way in which the generation of text can be inaccurate is when AI hallucinations arise. An AI hallucination is when the LLM produces sentences that contain made-up falsehoods. Rather than referring to these as AI hallucinations, which seem somewhat anthropomorphizing of AI, some prefer to say that these are AI confabulations.

Most users are not on the lookout for AI hallucinations. They get their output from AI and tend to assume that everything is perfectly fine. The especially tricky aspect of AI hallucinations is that we don’t know when they will arise; they can arise in the subtlest of ways, people aren’t expecting them, and the confabulations can slip through a casual inspection by a user who isn’t on their toes. Ongoing research to try and prevent AI hallucinations, or at least automatically detect them when they happen, is avidly underway; see my discussion at the link here.

Putting Two Things Together

Consider then these two situations:

  • (1) AI hallucinations that ordinarily arise. We already know that AI hallucinations happen from time to time, and that they aren’t particularly predictable as to when or if they will happen.
  • (2) AI hallucinations that arise specifically during watermarking. New research suggests that AI hallucinations can be spurred by the AI watermarking process, such that an AI hallucination can come along while AI watermarking occurs.

I would bet that many of you might not have thought about this somewhat unique second case. We already have AI hallucinations that occur during the normal activities of an LLM. Now, we have the special case of AI hallucinations that happen when AI watermarking is taking place. It’s a double whammy and worsens an already undesirable phenomenon.

Research Revelation

In a recently posted research study entitled “Invisible Ink, Visible Lies: How Production Watermarking Causes LLMs to Hallucinate” by Haocheng Ye, Aoting Hu, Xinwei Zhang, Xunzhu Tang, Shuchao Pang, Jason Xue, arXiv, October 4, 2026, these salient points were made:

  • “This paper studies factual degradation caused by watermarking itself. We identify watermarking hallucination as a failure mode distinct from ordinary hallucination and surface-quality loss: watermarked text can remain fluent and detectable while changing a context-supported fact.”
  • “Using a controlled retrieval-augmented generation setting, we compare unwatermarked and watermarked generations under the same context, query, and decoding configuration, and quantify their factual accuracy decrease.”
  • “Across six representative watermarking methods, including KGW, SWEET, DiPmark, GumbelSoft, Gumbel-Max, and SynthID watermarking, we consistently observe watermark-induced hallucination.”
  • “Overall, watermarking should not be assumed harmless in fact-critical applications: factual reliability must be evaluated alongside detectability, robustness, and output quality.”

You can see that the researchers have coined this new finding as watermarking hallucinations. This provides a distinction: there are conventional AI hallucinations, and now we are also faced with AI watermarking hallucinations. In that sense, AI watermarking appears not to follow the Hippocratic oath, namely, the presumption is to first do no harm. It will be important to see if other research further replicates the watermarking hallucinations machinations.

Example Of AI Watermarking Hallucination

I’ll give you a quick example of an AI watermarking hallucination. This will offer insights into what this added scourge of AI hallucinations is about.

Suppose that the AI is responding to a prompt about when Halloween annually occurs. First, consider when the AI isn’t using watermarking. The response generated might be this:

  • AI-generated response (no watermark): “Halloween takes place on October 31 of each year.”

Let’s now turn on watermarking. The AI is going to select words in a manner that will probabilistically allow us to later examine the sentence and try to determine whether the watermarking algorithm was making the wording choices. Assume that the watermarked version opted to reword “takes place” with the single word “occurs”. The watermarked response then is this:

  • AI-generated response (watermarked): “Halloween occurs on October 31 of each year.”

So far, so good.

AI Hallucination During Watermarking

What can happen is that while the watermarking is taking place, an AI hallucination ends up arising in the flow of the words being generated. It could be that an AI hallucination happens in this case regarding the date of Halloween. For example, consider this response:

  • AI-generated response (watermarked, AI hallucinated): “Halloween occurs on October 13 of each year.”

The date says October 13 but should say October 31. What happened? A local substitution in the generated sentence can change aspects such as numbers, dates, names, and other seemingly clear-cut factual identifiers. Observe that the fluency of the sentence is still maintained. The sentence reads like a normal sentence, and there isn’t anything that immediately catches our eye.

We would almost be luckier if fluency also were undermined by an AI hallucination in this context. Why so? It might be a better telltale sign that an AI hallucination has happened. For example, we might want this instead:

  • AI-generated response (watermarked, AI hallucinated): “Halloween putrefies on October 13 of each year.”

I would assume that even a casual user would notice that the sentence says, “Halloween putrefies,” and wonder what in the heck is going on. This would get their Spidey-sense tingling and possibly then catch that the October 13 date is also incorrect.

AI Watermarking Is A Trade-Off

Let’s ponder the ramifications of this added means of generating AI hallucinations.

One interesting question is whether the use of AI watermarking is worth the chances of inducing AI hallucinations. It now presents a trade-off in considerations. You might be tempted to say that AI watermarking is such a vital contribution to society that any rise in AI hallucinations is worth it. The problem there is that suppose the rise in AI hallucinations was humongous. Would you still support AI watermarking? I believe we would be accepting as long as the chances of AI hallucinations arising during AI watermarking is sufficiently low.

Another fascinating and crucial question is how much AI watermarking leads to AI hallucinations versus how much conventional AI generation leads to AI hallucinations. If the rate of AI hallucination is the same, well, we probably would write this off as the same old, same old. In that way of thinking, the ball game is still the same.

An outlier thought about this is whether conventional AI hallucinations can arise amidst the AI watermarking hallucinations. In essence, suppose that without watermarking there is a chance X that AI hallucinations might arise, and that there is a chance Z during AI watermarking that an AI hallucination might arise. Are the probabilities of X and Z independent of each other? Can they both occur in the same sentence? Does this raise the overall chances of AI hallucinations, therefore regrettably boosting AI hallucinations exceedingly?

The World We Are In

The researchers explored various means to try to eliminate or mitigate AI watermarking hallucinations. You can anticipate that AI makers are going to work hard to try and devise AI watermarking that has zero chances of generating AI hallucinations. They certainly don’t want people to get upset about AI watermarking that seemingly undermines their generated results. If there are AI hallucinations induced by or corresponding with AI watermarking, it means that the “cure” of watermarking comes with a corresponding malady.

Of course, since we are already familiar with the bane of AI hallucinations, maybe some will just shrug off that AI hallucinations can occur during AI watermarking. From their perspective, AI hallucinates, and it doesn’t really matter when or why. Their view is that we need to stamp out AI hallucinations, and until then, they are a sordid fact of life.

A final thought for now. The famous painter Pablo Picasso made this pointed remark: “Life is full of surprises, some good, some not so good.” It is perhaps a bit of a surprise that there appears to be a link between AI watermarking and the inducing of AI hallucinations. Hopefully, the next surprise will be that we can completely alter AI watermarking so that it never contributes to AI hallucinations. That would be a very pleasant surprise.

Source link

Share

Leave a comment

Leave a Reply

Your email address will not be published. Required fields are marked *