Productspublished

OpenAI’s Jalapeño Claims Up to 1.9x Efficiency Gain, but Leaves Nvidia Training Intact

The Broadcom co-developed inference processor is headed for deployment by year-end. Its published gains are meaningful, but they exclude model training and Nvidia’s Vera Rubin platform.

By 3 min read
OpenAI’s Jalapeño Claims Up to 1.9x Efficiency Gain, but Leaves Nvidia Training Intact

Listen to this story

The audio brief

About 1:45
0:001:45
Read transcript
OpenAI says its new Jalapeño chip delivered up to 1.9 times more throughput per kilowatt, with latency as much as 3.6 times lower than Nvidia’s GB200 and GB300 systems. But this is an inference result, not a broad Nvidia replacement. Jalapeño is an in-house ASIC, or application-specific integrated circuit, co-developed with Broadcom, and OpenAI plans to begin deploying it in its infrastructure by the end of 2026. The tests used SemiAnalysis’s public InferenceX suite, covering GPT-OSS 120B, DeepSeek R1 670B, and Moonshot AI’s Kimi K2.5. They focused on defined workloads, mainly single-token prediction, and compared Jalapeño with only the GB200 and GB300 rack systems. The chip combines six HBM4 memory stacks, totaling 216 GiB, with 15.4 terabytes per second of bandwidth. Its package power was 700 watts, versus 1,200 for GB200 and 1,400 for GB300. But the broader utility comparison was narrower: 1.18 kilowatts for Jalapeño versus 2.55 for GB300. And when GB300 used multi-token prediction, Jalapeño’s peak efficiency lead fell to roughly 1.5 times. The larger constraint is that Jalapeño cannot train models, and OpenAI has not tested it against Nvidia’s Vera Rubin platform. OpenAI still expects to deploy Nvidia and other partners’ accelerators widely. The key question is whether Jalapeño’s inference advantage survives real infrastructure deployment, while Nvidia remains central to training and to OpenAI’s Ohio project.

Story brief

3 key points

OpenAI is positioning Jalapeño as a specialized inference accelerator rather than a replacement for Nvidia. In SemiAnalysis’s InferenceX tests, the Broadcom co-developed ASIC used 216 GiB of HBM4 and delivered its strongest results against GB200 and GB300 under specified single-token workloads. OpenAI still plans to use Nvidia and other suppliers for training and inference, while Jalapeño is not training-capable and...

  1. 01

    Jalapeño combines six HBM4 stacks for 216 GiB of memory and 15.4 TB/s bandwidth.

  2. 02

    Package power was 700W, versus 1,200W for GB200 and 1,400W for GB300; all-in utility comparisons were narrower.

  3. 03

    Against a GB300 using multi-token prediction, Jalapeño’s peak efficiency advantage fell to about 1.5x.

OpenAI says Jalapeño, its custom chip co-developed with Broadcom, delivered up to 1.9 times higher throughput per kilowatt and 1.7 to 3.6 times lower end-to-end latency than Nvidia GB200 and GB300 systems in published tests. The result gives OpenAI a new inference option, not a verdict on Nvidia’s broader position: Jalapeño does not train models, was not tested against Nvidia Vera Rubin, and OpenAI expects to keep deploying Nvidia and other partners’ accelerators.

Jalapeño is an in-house inference ASIC, or application-specific integrated circuit, that OpenAI announced with Broadcom in June. OpenAI plans to begin deploying it in its compute infrastructure by the end of 2026.

The benchmark measures a defined slice of inference

The tests used SemiAnalysis’s public InferenceX suite and covered GPT-OSS 120B, DeepSeek R1 670B and Moonshot AI’s Kimi K2.5. They compared Jalapeño with Nvidia’s GB200 and GB300 rack systems, rather than across every current Nvidia platform or workload.

Memory bandwidth is the stated design target

Each Jalapeño package combines a compute die with six HBM4 memory stacks, totaling 216 GiB and 15.4 TB/s of bandwidth. OpenAI’s Hot Chips presentation said the architecture is designed to expose aggregate HBM bandwidth rather than simply add more of it.

Two ways to read the power comparison

  • Jalapeño is rated at 700W, versus 1,200W for GB200 and 1,400W for GB300 accelerators in the cited comparison.
  • OpenAI said Jalapeño’s sustained power stayed at or below 550W during testing.
  • Its all-in utility-power comparison used 1.18kW for Jalapeño and 2.55kW for GB300, producing narrower gaps than the package-power comparison.

A custom option alongside Nvidia systems

The competitive limits are explicit. Jalapeño was not tested against Vera Rubin and does not support model training. The major tests also used single-token prediction on both Jalapeño and GB300; when compared with a GB300 using multi-token prediction, Jalapeño’s peak efficiency lead fell to roughly 1.5 times.

OpenAI says it expects to widely deploy Nvidia and other partners’ accelerators for both training and inference. Nvidia has also said it will provide up to $105 billion in credit support for an OpenAI data-center project in Ohio that will exclusively use Nvidia compute. Jalapeño’s near-term test is whether its published inference advantage persists as OpenAI moves it into its own infrastructure.

Editorial analysis

Our Read

Our read: Jalapeño’s strategic value is not that it settles the GPU contest. It gives OpenAI a processor tailored to an inference design target: making aggregate HBM memory bandwidth available to workloads. The production test now matters more than the headline benchmark. OpenAI plans deployment by year-end, and its figures change when power is counted at the utility level or when GB300 uses multi-token prediction. Watch whether OpenAI discloses comparable production results across those configurations while it continues to deploy Nvidia systems, including the Nvidia-only Ohio project.

Sources

  1. tomshardware.comOpenAI’s 700W Jalapeño ASIC outpaces 1,400W Nvidia flagship GPU — claims up to 1.9x throughput per kilowatt and 3.6x lower latency, co-developed with Broadcom
  2. cnbc.comOpenAI says its Broadcom custom chip is a winner. What does that mean for Nvidia?