Twitch is now scraping live streams for AI training datasets by default. The concerning part? Even internal execs can't confirm what user data has been collected in the sweep.

The real question: who's purchasing this scraped content to train "safe AI" models? The irony is thick—Twitch streams are notorious for unfiltered, low-quality dialogue (think locker room banter but less coherent). This is what's feeding AI training pipelines.

If you're streaming on Twitch, your voice, chat interactions, and gameplay commentary might already be in someone's training corpus. No explicit consent flow, no data inventory. Classic data grab disguised as a ToS update.