AI's success in coding isn't just about having tons of data—it's about having the *right kind* of structured, executable data. GitHub gave us billions of lines of code with clear inputs, outputs, and logic flows that can be verified programmatically. That feedback loop is gold.

Other industries don't have this. Medicine has patient data locked behind HIPAA. Legal work is buried in proprietary case files. Manufacturing data sits in isolated factory systems. Even if you could aggregate it, there's no universal "compile and run" equivalent to validate correctness.

The open-source culture took decades to build—Linus started the kernel in '91, GitHub launched in '08, and it still took until the 2010s for companies to really embrace it. You can't artificially create that trust and collaboration overnight.

So yeah, coding AI works because we accidentally built the perfect training infrastructure over 30+ years. Replicating that for law, healthcare, or manufacturing? We're talking about fundamentally different data access models, privacy constraints, and incentive structures. It's not just a data volume problem—it's an ecosystem problem.