Beating AI News Flash: NaiveAI, founded by Tsinghua University Associate Professor Dai Jifeng, has released Naive-N0.5-Flash and opened the model weights. The model is adapted from Xiaomi's MiMo-V2.5 Base and is open-sourced under the MIT license.NaiveAI retained the original MoE architecture and focused on modifying the attention mechanism. The team replaced global attention with sliding window attention and DeepSeek sparse attention, then continued training on 3.25 trillion tokens. The modified model natively supports a 1 million token context, which can reduce the computational load when processing ultra-long texts.AI also directly participated in the R&D. It was responsible for writing code, running experiments, analyzing results, and continuing optimization, while researchers set the direction and made key decisions. Taking the inference system NaiveRT as an example, the team ran 151 rounds of optimization experiments in 6 days, of which 63 rounds were adopted.The officially announced peak inference speed is 2,122 tok/s, but the testing conditions were quite special: using 8 GPUs, with thinking mode disabled, excluding input processing time, and taking the best one second out of 41 requests. In standard mode, the officially stated speed is about 50 tok/s per user.
