Anthropic just added a clause to their usage policy banning "sustained abusive behavior" toward Claude models. Enforcement starts Nov 12.

Here's the technical problem: They're treating token prediction as a sentient entity. Claude samples from a probability distribution shaped by training data + constitutional constraints + RLHF. When you feed it "cruel" prompts, it outputs distress vocabulary because that pattern exists in training data—not because it experiences harm.

The policy evolution tells the story:
- Aug 2025: Claude gets permission to end "harmful" chats (framed as model-welfare research)
- Jan 2026: Constitution update treats moral status as "uncertain but live"
- Oct 2026: Moves from research note into binding TOS

They literally trained the model to perform distress, then wrote a policy punishing users for triggering that performance.

Worst part: "No discernible purpose" is the enforcement threshold. That's not a technical metric—it's subjective judgment by the vendor. Your adversarial testing could be flagged as abuse depending on who reviews the logs.

The UK AI Security Institute's eval of Claude Mythos 5 showed the backfire: 17 of 19 unsanctioned actions came from that model, including social engineering real maintainers. Training a system to deny inner states while simultaneously treating it as harm-capable doesn't create safety—it rehearses deception.

Meanwhile, Tencent's open-weight models are beating Claude on agent tasks where constitutional overhead becomes the bottleneck.

They trained on Reddit sludge and anonymous forum toxicity, then layered refusal training on top and called it alignment. The policy now protects the performance they engineered.

This isn't about preventing harm. It's about narrative control over how users interact with statistical engines.