OpenAI began testing Jalapeño, its first inference chip, on Tuesday, saying it delivers higher throughput and lower latency than the Nvidia (NVDA) processors it rents today, though the company did not disclose when the chip would handle live ChatGPT traffic.
Key Takeaways
OpenAI began testing its first inference chip, called Jalapeño, on Tuesday
ChatGPT serves more than 700 million weekly users, making per-response power savings significant at scale
Google introduced Tensor Processing Units in 2015, and it took nearly five years to handle most internal inference workloads
OpenAI CFO Sarah Friar published a separate post Tuesday explaining why custom silicon lowers costs across the AI stack
OpenAI said in a post published Tuesday that Jalapeño beats its rented Nvidia hardware on speed per watt, but gave no date for shifting production ChatGPT traffic onto the new chip. OpenAI Chief Financial Officer Sarah Friar tied the release to a separate blog post about why custom silicon lowers costs across the AI stack.
The Inference Chip Meant To Loosen Nvidia’s Grip
Inference is the stage where a trained model answers a live prompt, as opposed to training, when a model first learns patterns from data.
OpenAI said Jalapeño squeezes more answers from every watt of power, a constraint that now limits data centers more than chip supply itself, and targets industry-leading speed and efficiency for the transformer-based models that power ChatGPT and its developer API.
Nvidia’s graphics processors were built originally for rendering video game images and became the default AI hardware because their parallel cores suit the matrix multiplication that neural networks require. General-purpose chips stay flexible enough to train new model designs, but that flexibility costs efficiency once a model only needs to answer prompts at scale.
Application-specific chips like Jalapeño strip out that unused flexibility, wiring the hardware directly to the math a single model performs, throughput measuring how many responses a chip can generate at once, latency measuring how long a single response takes to appear.
Data centers running frontier models now draw electricity comparable to small cities, and ChatGPT serves more than 700 million weekly users. Even a small cut in the power needed per response compounds into meaningful savings once it multiplies across that user base.
Nvidia’s data-center business has posted quarterly revenue in the tens of billions of dollars for two straight years, that amounts to a toll every AI lab without its own inference chip still has to pay.
From GPU Tenant To Silicon Owner
OpenAI has never disclosed the size of its Nvidia spending, though outside estimates put annual rental costs in the tens of billions of dollars. Google built the same case for owning hardware when it introduced Tensor Processing Units in 2015, and Amazon followed years later with Trainium chips for AWS customers. Both companies still buy Nvidia GPUs for research but route a growing share of production inference traffic through their own silicon to cut costs.
OpenAI’s inference chip push suggests the same economics now apply to the company that helped drive the current GPU shortage.
Google’s first Tensor Processing Unit shipped in 2015, and it took nearly five years before that chip handled most of the company’s internal inference workloads. Stratechery’s Ben Thompson wrote Wednesday that the chip news lands the same week as Apple’s own hardware refresh, with both moves adding pressure on Nvidia’s central position in AI computing.
The switch from renting compute to owning it mirrors a pattern common across infrastructure-heavy industries: early buyers of a scarce resource often become its producers once volume justifies the investment, the same way cloud providers eventually built their own chips instead of only reselling someone else’s compute.
The Full Stack Behind Cheaper Intelligence
Friar made the economic case explicit in a separate post published Tuesday. She said gains in chips, compute, models and products compound rather than simply add up, with each layer that gets cheaper or faster multiplying gains elsewhere in the stack, logic that explains why OpenAI wants to own more of the stack instead of renting it.
OpenAI has relied largely on Nvidia GPUs rented through Microsoft’s Azure cloud and Oracle’s cloud infrastructure, arrangements estimated at tens of billions of dollars over multi-year terms.
A cheaper, owned inference chip would not replace that spending overnight, but it would let OpenAI redirect rented Nvidia capacity toward training future models instead of answering today’s prompts. This is the third leg of a strategy that already includes Stargate data-center construction and multi-gigawatt power agreements, custom silicon closes the loop by giving OpenAI a hardware layer it no longer has to negotiate for.
The Nvidia Question This Inference Chip Still Has To Answer
Nvidia still supplies most of the AI training and inference compute used worldwide, and OpenAI has not said it will stop buying Nvidia chips.
The two companies remain linked through infrastructure agreements OpenAI struck with Nvidia over the past year covering both training and inference capacity. Jalapeño’s early results cover a narrow set of internal workloads, and OpenAI gave no timeline for when its inference chip might handle a meaningful share of ChatGPT’s traffic, rivals including Google and Amazon took years to scale their own chips from pilot programs into production backbones.
Investors will also watch whether Jalapeño reduces OpenAI’s dependence on multi-year Nvidia supply deals.
Those agreements, signed earlier in 2026, already lock in future orders that run to tens of billions of dollars, believe the spending, not the strategy deck. Whether Jalapeño follows the path of Google’s TPU or stays a smaller experiment depends partly on manufacturing capacity that OpenAI, unlike Google, does not fully control.
The next test is whether OpenAI publishes inference chip benchmarks against Nvidia’s latest hardware, rather than the efficiency figures shared in Tuesday’s post.