**OpenAI** stated that its custom chip, the ‘Jalapeño’, outperformed Nvidia’s GB300 in two 700-watt environmental inference benchmarks. The company plans to roll out this chip in stages after 2026 to lower costs and improve response times.

Key takeaways

  • OpenAI said that, based on its test results, Jalapeño showed an advantage over Nvidia GB300 in terms of throughput per unit power and response speed.

  • Designed as a 700-watt-class inference chip to deliver AI model services faster while reducing data center operating costs.

  • Hailapeno is not intended for training (training), and it was not directly compared to the latest Nvidia “Vera Rubin” generation, whose shipments have already begun.

OpenAI benchmark results

OpenAI said its custom processor, Hailapeno, recorded higher performance than Nvidia’s GB300 in tests in terms of the amount of AI work that can be processed per unit of power and the response return speed metrics.

Richard Ho, OpenAI’s head of semiconductor operations, said that Hailapeno “can deliver high throughput and low latency at the same time, giving customers a wider range of choices between cost-optimized models and ultra-fast response models.”

According to Ho, Hailapeno, co-developed with **Broadcom**, achieved these benchmark results under a 700-watt power consumption condition. Power costs are one of the largest items in the total cost of running a data center.

Also worth seeing: RWA market surges 48.7%, but it’s just one token that explains nearly everything

Comparison with Nvidia, and the limitations are also clear

This comparison has clear limitations.

First, Hailapeno has not directly gone head-to-head with Nvidia’s latest Vera Rubin-generation chip, whose shipments have already begun. Also, the design objectives are different. Hailapeno is optimized for inference rather than training AI models, so it is not at a level that immediately threatens Nvidia’s position in the large-scale training market. OpenAI is still keeping Nvidia as a key supplier, too.

OpenAI is also using **Cerebras Systems** technology for some inference workloads, especially for small models. Ho said, “Hailapeno can handle bigger models,” adding that because computing demand is so enormous, the situation forces them to use multiple hardware platforms in parallel. “We need compute so much,” he said.

This public testing included one of OpenAI’s small open-source models, as well as DeepSeek and the model from **Moonshot AI**. In particular, the explanation was that Hailapeno’s relative strengths were more evident in Moonshot AI’s large “Kimi” model.

Work on developing the second-generation chips is already well underway, and a tape-out—expected within a few months, referring to the stage after design is completed when it is handed over to production lines. At the same time, OpenAI has already begun work on its third-generation concept.

This chip is an extension of the roadmap that OpenAI and Broadcom announced on June 24. At the time, the two companies said that thanks to design automation using OpenAI’s own models, the first-generation Hailapeno reached the manufacturing tape-out stage—nine months after design kickoff. The chip released this time is that first-generation product.

Next article: Nvidia to supply GPUs for SpaceX’s 2027 orbital AI plans…shares fall in tandem