Subscribe Sign in

AI model watermarking changes agent behavior

1 min read Rewritten in plain language

Artificial intelligence

Show what we removed Rules applied: A3×4 A4 D2×4 D3×3 D4×2 F2×4 all 30 rules
  • Lasso Security sees differences in tool handling and model refusals.
  • With the implementation of the EU AI Act, providers of AI models must mark the output of their software with machine-readable code.
  • For example, if Claude were emitting the sentence "The weather today was cold and…" then it might favor one statistically likely candidate (e.g. "overcast") over an alternative (e.g "gray").
  • Thus an agent based on OpenClaw or an API client that calls an Anthropic model would process whatever output variation follows from Anthropic's watermarking.
  • As for refusals watermarking had a small effect on the handling of harmful requests, based on test runs using HarmBench and JailbreakBench.

5 sentences from our version of the report, chosen to cover it. Nothing here is written; every line is in the article below. How

Lasso found this sort of digital content tagging can affect tool calling and refusal behavior. In terms of tool calling, based on a benchmark called BFCL v4 single-turn AST, watermarking reduced the accuracy on six of seven models tested.

Headline check

There is nothing in this headline a machine can check against the report: no figure, no name and no quotation.

Nothing was measured here, so nothing is claimed. How this is checked

Lasso Security sees differences in tool handling and model refusals.

Watermarks that European law needs be added to AI-generated content to establish provenance may come at a cost.

The altered behavior isn't necessarily worse but can be, under adversarial prompt injection.

With the implementation of the EU AI Act, providers of AI models must mark the output of their software with machine-readable code. Google DeepMind's SynthID-Text is one method for doing so, and has been adopted by Anthropic and by OpenAI.

The benefit of this sort of digital labeling is that manipulative or deceptive AI-generated content can be more detected, even if it does have the potential to stigmatize the usage of AI.

Anthropic's explanation of how it applies watermarks to Claude output involves intervening in the prediction that results in specific words. For example, if Claude were emitting the sentence "The weather today was cold and…" then it might favor one statistically likely candidate (e.g. "overcast") over an alternative (e.g "gray").

Lasso found that this sort of digital content tagging can affect tool calling and refusal behavior.

Thus an agent based on OpenClaw or an API client that calls an Anthropic model would process whatever output variation follows from Anthropic's watermarking.

In terms of tool calling, based on a benchmark called BFCL v4 single-turn AST, watermarking reduced the accuracy on six of seven models tested.

As for refusals watermarking had a small effect on the handling of harmful requests, based on test runs using HarmBench and JailbreakBench.

Shortened to 1 minute of reading, this version reads 9.8 on the Niral Score.

You are reading our version, not theirs. This is The Register's report shortened to its most important sentences, in plainer words, with verdicts and loaded words taken out. Plain description stays, and so do adjectives that carry a fact, such as "former" or "federal". The reporting, the facts and the quotations are theirs — quotations are never edited — and the indicators beside it measure this version. Hover or tap Adjectives to see every one left in the text.

How this outlet filed it, and how we rewrote it

No other newsroom we read has filed on this event, so there is nothing to compare it with yet.

Outlet Niral ScoreAdjectivesSourcingHappiness
The Registeras they published this story 11.5 13 54 50
Mundane Readneutralized from The Register 10 13 54 50

Sign in to react.

Comments

Nothing here yet.

Sign in to comment.

Questions

Readers can ask a question about this story here. Questions and answers are for subscribers. Sign in to read them.

Comments are read before they appear where anything in them needs a person to look. Nothing posted here is ever deleted; a comment taken down keeps its text and the reason, so the decision can be looked at again. How this works