New RL approach flips the script on AI tutoring: instead of training models to match pedagogical rubrics, Sherpa spins up simulated student archetypes with different learning styles and trains a teacher model by optimizing actual student outcomes.

Core mechanism: multi-turn RL loop where the reward function is the student's measured learning delta, not some hand-crafted "good explanation" heuristic. Teacher iterates on explanations until the simulated student actually gets it.

Results on MathTutorBench:
• Pedagogy score jumps from 52.5% to 79.2%
• Average +20.5pp lift across student archetypes
• 79.6% human preference over base model in pairwise evals

The human preference check is critical here—proves the simulated students didn't just reward some weird emergent style that real humans would hate.

Caveat: gains are bounded by the archetype set. Paper doesn't show whether this transfers to real classrooms with messier, less predictable student behavior. Still, training on learner outcomes instead of static rubrics is a legit paradigm shift for instruction-tuning.