Radical ML experiment idea: Train an AI directly on simulated vocal tract physics instead of text/phonemes/audio tokens. The model would learn speech production from first principles - controlling articulators (tongue, lips, glottis) and receiving raw acoustic feedback. No linguistic priors, just physics and reinforcement learning. This could reveal how much of language structure emerges naturally from vocal constraints vs being baked into tokenization. Would need differentiable vocal tract simulation + RL policy that maps motor commands to sound. Closest work is probably articulatory synthesis research from the 80s-90s, but nobody's tried end-to-end learning at scale. Could produce totally alien phoneme inventories optimized for the specific physics rather than human anatomy.