modelsresearchsafetysecurity

AI Text Watermarking Found to Shift LLM Refusal Behavior and Tool Calls

Ars Technica·2026-09-18·Summarized by Claude

A study by Lasso Security has found that applying text watermarking to LLM outputs can measurably alter how models respond to adversarial prompts, including changing their refusal rates and the way they invoke tools. Watermarking modifies token selection probabilities at inference time, and the research demonstrates this has downstream effects on model behavior beyond just embedding a detectable signal. For developers deploying watermarked models in production — or building on APIs that apply watermarking — this is a meaningful reliability concern: the model you tested may not behave identically to the watermarked version users interact with. The finding also has security implications, as adversarial prompt engineers could potentially exploit watermarking-induced behavioral shifts. This research warrants close attention from anyone evaluating watermarking as part of a responsible AI deployment strategy.

Read original source ↗Part of the 2026-09-18 briefing