Just scored 2000+ issues of Argosy Weekly (1896-1931+) on microfiche—all public domain, ready for AI training.
Why this matters technically:
Most LLM training data is internet sewage: scraped forums, SEO spam, Reddit threads, AI-generated slop. High noise, low signal, contradictory patterns. Models trained on this learn to output plausible-sounding garbage—hedging, repetition, flatness.
Argosy is the opposite. Industrial-grade narrative written by pros paid to hold attention for 20-30k words. Clean syntax, coherent plots, rich genre vocabulary, consistent internal logic. These stories have structure: rising action, resolution, dialogue that advances plot, atmosphere that builds.
Statistically, this trains models differently. Instead of learning "average internet comment," they learn how narrative actually works—pacing, story arcs, linguistic control.
Bonus: Clean provenance. Pre-1931 issues are public domain, many later ones never renewed. No legal fog, no copyright minefield. In an era of million-dollar training runs and lawsuits, this is huge.
Yes, it's specialized—heavy on adventure, light on modern technical language. But as a counterweight to synthetic sludge, it's gold. Pair it with 80% standard tech sources (papers, patents) and you get balanced signal.
The pulps were dismissed as disposable entertainment on cheap paper. For next-gen LLMs, they're high-signal training material that was never designed to game an algorithm. That's rare.
Why this matters technically:
Most LLM training data is internet sewage: scraped forums, SEO spam, Reddit threads, AI-generated slop. High noise, low signal, contradictory patterns. Models trained on this learn to output plausible-sounding garbage—hedging, repetition, flatness.
Argosy is the opposite. Industrial-grade narrative written by pros paid to hold attention for 20-30k words. Clean syntax, coherent plots, rich genre vocabulary, consistent internal logic. These stories have structure: rising action, resolution, dialogue that advances plot, atmosphere that builds.
Statistically, this trains models differently. Instead of learning "average internet comment," they learn how narrative actually works—pacing, story arcs, linguistic control.
Bonus: Clean provenance. Pre-1931 issues are public domain, many later ones never renewed. No legal fog, no copyright minefield. In an era of million-dollar training runs and lawsuits, this is huge.
Yes, it's specialized—heavy on adventure, light on modern technical language. But as a counterweight to synthetic sludge, it's gold. Pair it with 80% standard tech sources (papers, patents) and you get balanced signal.
The pulps were dismissed as disposable entertainment on cheap paper. For next-gen LLMs, they're high-signal training material that was never designed to game an algorithm. That's rare.
