According to Sina Finance, Nvidia said Groq 3 LPX cabinets have entered mass production, marking the commercialization of technology obtained through the company’s largest acquisition. Nvidia senior director Dion Harris said the Groq cabinets will be deployed alongside Vera CPUs and Rubin GPUs in Nebius data centers and will go live later this year.

Nvidia said it is accelerating production of the chips and delivering them to customers, underscoring the importance of low-latency inference. Harris said the product allows token-output providers to offer premium services to latency-sensitive customers. Nvidia also said cloud providers can charge more for this type of high-speed token output service.

The company said it acquired Groq-related assets for $20 billion in December last year, its largest acquisition to date. Nvidia said each LPX cabinet integrates 256 Groq 3 chips, and cited Artificial Analysis benchmark data showing output of 3,400 tokens per second.

Nvidia said low-latency chips will not replace GPUs. Harris said they are intended for the decode phase of model serving, while GPUs remain the main chips for both training and inference. Nvidia is also expanding shipments of its Vera Rubin system, which began production this year. Nvidia will report earnings on Wednesday.