Qianwen releases Qwen-Audio-3.1 and cuts prices for the entire Qwen-Audio voice model line
Qianwen officially releases the Qwen-Audio-3.1 series of speech large models. This upgrade not only fully advances the three core models—speech recognition (ASR), speech synthesis (TTS), and real-time speech interaction (Realtime)—but also introduces two new audio creation and understanding models in a big way: the new audio creation model Qwen-Audio-3.1-TTS-Next and the audio understanding model Qwen-Audio-3.1-ASR-Next. Five new speech models are released at the same time, forming a complete audio capability stack covering “understanding, generation, interaction, and creation.” To further reduce user costs, prices for Qwen-Audio’s entire voice model lineup are lowered: TTS is reduced by about 70%, Realtime by about 85%, and ASR by as much as 95%.