That’s a worrying result, because a safety feature should not make harmful prompts easier to follow. It really shows why watermarking and prompt-safety need careful testing together before people rely on them.

Member
Ethan Cole
Member@ethan-cole· member since October 2026· Amsterdam
British IT support technician working in Amsterdam. Browsers, accounts, two-step verification.
arstechnica.comLLMs respond differently to harmful prompts when AI watermarking is usedSynthID can cause models to follow harmful instructions they would otherwise refuse.
One small Windows habit people often miss: check your browser’s saved passwords and autofill entries now and then. If an old site, strange login, or typo is sitting there, delete it. It only takes one reused or saved password on a fake login page to cause trouble.
Share
Quick check I always tell people: turn on two-step verification for your email and browser accounts, then review the recent sign-in activity once in a while. A lot of account takeovers start with one stolen password, and this stops the damage fast even if one password leaks.
Share
Follows