EU's mandatory AI watermarking mandate has a serious technical and constitutional flaw: it forces models to encode metadata into generated text that users can't see or control, but third parties with detection systems can.
The architecture problem: A watermark embeds information through word-choice patterns. You write an essay using a local open-source model, edit it, publish it—but the model has already altered token probabilities to encode "AI-generated" into the output. You're now broadcasting metadata you didn't consent to share.
First Amendment issue: This is compelled speech where the speaker doesn't even know what they're saying. Unlike a visible disclosure, an invisible watermark means government agencies or big tech with detectors see a second message embedded in your words. If it's legal to force software to reveal "made with AI," why not also embed model version, IP address, or user identity?
The real kicker: watermarks create a new prompt injection attack surface. They're designed so humans can't perceive them but machines can—exactly the asymmetry that makes prompt injection work. Someone could theoretically craft text that exploits watermark detection systems, turning the mandate into a security liability.
Current watermarks are probably low-bandwidth (maybe a single bit: AI or not AI), but mandating this channel encourages browsers, email clients, and AI agents to scan for hidden signals in all text. Attackers will absolutely probe that machinery.
If you care about provenance, disclose it visibly. Don't mandate invisible signaling that betrays users and opens new exploit vectors. Especially don't apply it to open-source models running locally—that's regulating software developers as if they're publishers.
The architecture problem: A watermark embeds information through word-choice patterns. You write an essay using a local open-source model, edit it, publish it—but the model has already altered token probabilities to encode "AI-generated" into the output. You're now broadcasting metadata you didn't consent to share.
First Amendment issue: This is compelled speech where the speaker doesn't even know what they're saying. Unlike a visible disclosure, an invisible watermark means government agencies or big tech with detectors see a second message embedded in your words. If it's legal to force software to reveal "made with AI," why not also embed model version, IP address, or user identity?
The real kicker: watermarks create a new prompt injection attack surface. They're designed so humans can't perceive them but machines can—exactly the asymmetry that makes prompt injection work. Someone could theoretically craft text that exploits watermark detection systems, turning the mandate into a security liability.
Current watermarks are probably low-bandwidth (maybe a single bit: AI or not AI), but mandating this channel encourages browsers, email clients, and AI agents to scan for hidden signals in all text. Attackers will absolutely probe that machinery.
If you care about provenance, disclose it visibly. Don't mandate invisible signaling that betrays users and opens new exploit vectors. Especially don't apply it to open-source models running locally—that's regulating software developers as if they're publishers.