GPU seller Tinygrad posted that, according to reliable reports, the GLM 5.2 model can achieve an inference speed of 120 tokens per second on two networked Blackwell-architecture tinybox systems. This configuration is priced at $150,000 and can be either a dual standard tinybox setup or a single tinybox Pro, both delivering the performance mentioned above. Tinygrad is promoting this as a selling point, focusing on a private deployment route with the tagline “one-time purchase, never paying cloud fees,” directly targeting pay-as-you-go cloud inference services. As of now, the message has not been confirmed by GLM, and Tinygrad has not disclosed additional technical details.---------------------------------Click the link to the original article below to join the Beating · Feishu AI News channel and get 24/7 continuous monitoring of global AI hotspots and news.