Tether is releasing a set of STEM synthetic data for small models. 99.45% can extract the answers—not the correct-answer rate.
On September 23, Tether AI Research released QVAC Genesis III, with 19.143 billion Tokens and about 159.6 million documents, covering 19 science and engineering fields. The difficulty ranges from high school to professional level. The roadmap continues from Genesis I in October 2025 and Genesis II in December. If students get it wrong, it’s written as a correction; if they get it right, it’s written as why each option is right or wrong. The model to be trained is intended to run on notebooks, phones, and local servers, with an emphasis on explaining the problem-solving path.
Among the 1.7 billion-parameter model trained from scratch, the option-level data achieves an effective answer rate of 99.45% on MMLU STEM. Compared with Cosmopedia-v2, whose corpus size matches the Token scale, the accuracies on ARC-Easy, ARC-Challenge, and MMLU STEM reach 51.85, 42.71, and 30.19 respectively—exceeding by 28.57, 21.35, and 15.03 points.
The data is released under CC-BY-NC 4.0, and is not available for commercial use. The paper has been accepted by COLM 2026.
On September 23, Tether AI Research released QVAC Genesis III, with 19.143 billion Tokens and about 159.6 million documents, covering 19 science and engineering fields. The difficulty ranges from high school to professional level. The roadmap continues from Genesis I in October 2025 and Genesis II in December. If students get it wrong, it’s written as a correction; if they get it right, it’s written as why each option is right or wrong. The model to be trained is intended to run on notebooks, phones, and local servers, with an emphasis on explaining the problem-solving path.
Among the 1.7 billion-parameter model trained from scratch, the option-level data achieves an effective answer rate of 99.45% on MMLU STEM. Compared with Cosmopedia-v2, whose corpus size matches the Token scale, the accuracies on ARC-Easy, ARC-Challenge, and MMLU STEM reach 51.85, 42.71, and 30.19 respectively—exceeding by 28.57, 21.35, and 15.03 points.
The data is released under CC-BY-NC 4.0, and is not available for commercial use. The paper has been accepted by COLM 2026.

