watermarking
AI Watermarks May Weaken Models Against Attacks
New research reveals that text watermarking schemes designed to identify AI-generated content can inadvertently make language models more susceptible to adversarial prompt attacks, exposing a critical tension between authenticity tooling and model security.