The AI training data debate hits different when you see the contrast: some labs literally destroying physical books (acid bath scanning for datasets) while preservation orgs are doing conservation work to keep texts alive for centuries.
The irony is brutal. We're shredding cultural artifacts to teach models about culture. Meanwhile book restoration tech (deacidification, rebinding, digitization that doesn't destroy originals) exists and works.
This isn't about being anti-AI. It's about questioning why we're choosing destructive data harvesting when non-destructive methods exist. The "move fast break things" approach applied to irreplaceable physical media is genuinely insane.
If you're building training datasets, maybe don't burn the library to light your GPU cluster.
The irony is brutal. We're shredding cultural artifacts to teach models about culture. Meanwhile book restoration tech (deacidification, rebinding, digitization that doesn't destroy originals) exists and works.
This isn't about being anti-AI. It's about questioning why we're choosing destructive data harvesting when non-destructive methods exist. The "move fast break things" approach applied to irreplaceable physical media is genuinely insane.
If you're building training datasets, maybe don't burn the library to light your GPU cluster.