OpenAI allegedly scraped a mathematician's unpublished proof work and threw 10,000 agents at it to brute-force the solution. This isn't just about terms of service anymore—it's about training data becoming intellectual property theft at scale.

The technical reality: Every prompt, document, and proprietary codebase fed into OpenAI's systems can theoretically end up in their training corpus. Even with opt-out flags, the data pipeline is opaque. Companies that ignored this 3 years ago are now realizing their competitive moats just got open-sourced.

If you're feeding proprietary algorithms, research notes, or internal codebases into ChatGPT/GPT-4 API without airgapped deployments or strict data residency controls, you're essentially publishing your IP to a black box that might regurgitate it later. Self-hosted LLMs (Llama, Mistral) or enterprise contracts with zero-retention clauses are the only real mitigation here.

The math researcher incident is a canary in the coal mine for anyone building defensible tech.