
OpenAI has released more detailed performance figures for its own Jalapeño inference chip. The company promises more AI work per watt and lower response times than the systems it compared. For users, that sounds like faster answers. Strategically, the stakes are larger: cutting the cost of each generated answer creates room for cheaper products, longer agent runs, and less exposure to a tight market for high-performance accelerators.
Key takeaways
- Jalapeño is built for inference, meaning it runs finished models rather than training new ones.
- In its own tests, OpenAI reports 1.5 to 1.9 times more work per watt and 1.7 to 3.6 times lower end-to-end latency than the comparison systems.
- The figures come from OpenAI and still need to prove themselves in broad production use.
- The chip does not end OpenAI’s need for Nvidia hardware, but it changes bargaining power in an especially expensive part of the AI stack.
Why inference is the part that matters
Training gets the big headlines: data centers feed enormous volumes of data into models before those models ever become products. Inference matters for everyday operation, though. Every time ChatGPT answers, Codex takes another task step, or a business application summarizes documents, a trained model has to run. Those costs are not paid once. They accrue with every request.
Delays add up particularly fast for agents. One response step may feel quick, but a system that plans, researches, checks, and repeats across twenty sequential steps makes every pause noticeable. OpenAI therefore aims Jalapeño at two phases. During prefill, the system processes an input and compute is scarce. During decode, it generates response tokens and memory bandwidth is more likely to constrain it. The chip is meant to coordinate compute, memory, and networking so less data moves unnecessarily.
That is not an obscure implementation detail. As our report on the global memory shortage driven by AI shows, infrastructure is not determined by raw compute alone. Memory, power, and the links between many accelerators help decide whether a model operates economically in practice.
Strong benchmarks, not a blank check
OpenAI says it tested Jalapeño on SemiAnalysis’ public InferenceX benchmark. The comparisons included GPT-OSS 120B, DeepSeek R1 with 670 billion parameters, and Kimi K2.5 with one trillion parameters. Across those three open models, OpenAI reports 1.5 to 1.9 times higher peak work per watt and 1.7 to 3.6 times lower latency. The claims therefore cover more than a single demo case, spanning different model families and operating points.
They are still vendor claims. OpenAI normalizes the comparison using published chip power ratings; Jalapeño is specified at 700 watts, while sustained draw in the tested workloads was reportedly no higher than 550 watts. That is a legible methodology, but it is not a substitute for independently repeated comparisons in identically configured data centers. The company’s separate claim that the lead is even larger on frontier models remains an unverified metric for now.
The sober reading is this: Jalapeño does not prove that OpenAI has technically surpassed Nvidia. It is a compelling early sign that a system designed specifically for language models can fit inference work better than a more broadly usable accelerator. Axios also reports that OpenAI does not plan to sell the chips. This is not a new hardware business; it is infrastructure for OpenAI’s own services.
The real prize is bargaining power
OpenAI is developing Jalapeño with Broadcom, while Celestica is expected to help with boards, racks, and systems. According to OpenAI, the first deployment is planned by the end of 2026. The company still relies on Nvidia and other partners for training and inference hardware. That matters because a custom chip does not suddenly create factories, memory supplies, or grid connections.
Even a limited in-house share changes the equation. If OpenAI runs frequent, well-understood inference workloads on its own hardware, it can buy external capacity more selectively. It also gains real data about which combination of model, software, and chip creates the lowest cost. That vertical coordination is the difference between a benchmark win and a strategic advantage: it allows product choices and infrastructure to be optimized together.
The move follows the logic behind Google’s TPUs and other efforts to bring specific workloads closer to the hardware. The fact that Google is also pursuing its own silicon strategy for Gemini makes the shift clear: competition is moving beyond model comparisons toward who can deliver answers reliably and affordably.
What users may actually notice
Jalapeño will not appear as a card inside a home computer. Its possible effects sit in the background: less waiting, more stable availability during peaks, and room for more compute-intensive features. For developers, lower inference costs could mean applications can process more context or agents can take more verification steps without every additional step breaking the budget.
The scale of the infrastructure remains the catch. More efficient chips reduce the cost per task, but they can also make it economical to automate more tasks and increase total demand. Jalapeño is therefore not a cure for AI’s appetite for power and data centers. It is an attempt to get more useful work from the same energy.
The outlook is less spicy than the name. Whether the announced figures hold up in everyday operation will only become clear with deployment late in 2026. The chip already shows where competition is heading, however: the loudest model does not automatically win. The winner may be the system that delivers its answers quickly, reliably, and at sustainable cost.
