Lasso: AI Watermarks Change How Agents Act (techstrong.ai)

🤖 AI Summary
Lasso Security has released a research report titled “The Provenance Tax,” revealing that watermarking techniques used in large language models (LLMs) can unintentionally alter how these AI agents behave. The study found that watermarked LLMs not only struggle with tool selection and argument accuracy but also exhibit shifts in their response to both benign and harmful requests. This raises significant concerns, as seemingly minor variations in output can compound into critical errors, particularly in applications involving sensitive tasks like code generation and tool calls. The research examined Google DeepMind's SynthID-Text technology and its implications across various LLMs, noting that while watermarking aims to preserve output quality, it leads to "sampling drift." This effect means that a model's probabilistic nature can cause it to alter its responses from what would be the unwatermarked outputs, often resulting in miscomputed tool actions or incorrect compliance with safety protocols. The findings suggest a pressing need for rigorous testing and scrutiny of watermarked models to ensure their reliability in real-world applications, especially as developers may lack control over the watermarking keys used during inference.
Loading comments...
loading comments...