OpenAI’s Jalapeño Claims Up to 1.9x Efficiency Gain, but Leaves Nvidia Training Intact
The Broadcom co-developed inference processor is headed for deployment by year-end. Its published gains are meaningful, but they exclude model training and Nvidia’s Vera Rubin platform.
Listen to this story
The audio brief
Story brief
3 key pointsOpenAI is positioning Jalapeño as a specialized inference accelerator rather than a replacement for Nvidia. In SemiAnalysis’s InferenceX tests, the Broadcom co-developed ASIC used 216 GiB of HBM4 and delivered its strongest results against GB200 and GB300 under specified single-token workloads. OpenAI still plans to use Nvidia and other suppliers for training and inference, while Jalapeño is not training-capable and...
- 01
Jalapeño combines six HBM4 stacks for 216 GiB of memory and 15.4 TB/s bandwidth.
- 02
Package power was 700W, versus 1,200W for GB200 and 1,400W for GB300; all-in utility comparisons were narrower.
- 03
Against a GB300 using multi-token prediction, Jalapeño’s peak efficiency advantage fell to about 1.5x.
OpenAI says Jalapeño, its custom chip co-developed with Broadcom, delivered up to 1.9 times higher throughput per kilowatt and 1.7 to 3.6 times lower end-to-end latency than Nvidia GB200 and GB300 systems in published tests. The result gives OpenAI a new inference option, not a verdict on Nvidia’s broader position: Jalapeño does not train models, was not tested against Nvidia Vera Rubin, and OpenAI expects to keep deploying Nvidia and other partners’ accelerators.
Jalapeño is an in-house inference ASIC, or application-specific integrated circuit, that OpenAI announced with Broadcom in June. OpenAI plans to begin deploying it in its compute infrastructure by the end of 2026.
The benchmark measures a defined slice of inference
The tests used SemiAnalysis’s public InferenceX suite and covered GPT-OSS 120B, DeepSeek R1 670B and Moonshot AI’s Kimi K2.5. They compared Jalapeño with Nvidia’s GB200 and GB300 rack systems, rather than across every current Nvidia platform or workload.
Memory bandwidth is the stated design target
Each Jalapeño package combines a compute die with six HBM4 memory stacks, totaling 216 GiB and 15.4 TB/s of bandwidth. OpenAI’s Hot Chips presentation said the architecture is designed to expose aggregate HBM bandwidth rather than simply add more of it.
Two ways to read the power comparison
- Jalapeño is rated at 700W, versus 1,200W for GB200 and 1,400W for GB300 accelerators in the cited comparison.
- OpenAI said Jalapeño’s sustained power stayed at or below 550W during testing.
- Its all-in utility-power comparison used 1.18kW for Jalapeño and 2.55kW for GB300, producing narrower gaps than the package-power comparison.
A custom option alongside Nvidia systems
The competitive limits are explicit. Jalapeño was not tested against Vera Rubin and does not support model training. The major tests also used single-token prediction on both Jalapeño and GB300; when compared with a GB300 using multi-token prediction, Jalapeño’s peak efficiency lead fell to roughly 1.5 times.
OpenAI says it expects to widely deploy Nvidia and other partners’ accelerators for both training and inference. Nvidia has also said it will provide up to $105 billion in credit support for an OpenAI data-center project in Ohio that will exclusively use Nvidia compute. Jalapeño’s near-term test is whether its published inference advantage persists as OpenAI moves it into its own infrastructure.
Editorial analysis
Our Read
Our read: Jalapeño’s strategic value is not that it settles the GPU contest. It gives OpenAI a processor tailored to an inference design target: making aggregate HBM memory bandwidth available to workloads. The production test now matters more than the headline benchmark. OpenAI plans deployment by year-end, and its figures change when power is counted at the utility level or when GB300 uses multi-token prediction. Watch whether OpenAI discloses comparable production results across those configurations while it continues to deploy Nvidia systems, including the Nvidia-only Ohio project.
Sources
- tomshardware.comOpenAI’s 700W Jalapeño ASIC outpaces 1,400W Nvidia flagship GPU — claims up to 1.9x throughput per kilowatt and 3.6x lower latency, co-developed with Broadcom
- cnbc.comOpenAI says its Broadcom custom chip is a winner. What does that mean for Nvidia?