Back to Blogs

What Is OpenAI's Jalapeño Chip? Benchmarks, Broadcom Deal & Inference Cost Impact

  • AI agent

What Is OpenAI's Jalapeño Chip? Benchmarks, Broadcom Deal & Inference Cost Impact

At the Hot Chips conference in Silicon Valley on Aug. 25, 2026, OpenAI VP of Hardware Richard Ho presented concrete performance figures for Jalapeño, turning what had previously been little more than a press release claim into a measurable piece of custom silicon. Jalapeño is OpenAI's first custom chip developed with Broadcom, designed specifically to run trained models faster and more efficiently than the GPUs OpenAI currently rents.

The Jalapeño Test results are worth examining beyond the headline numbers. Whether you're making infrastructure decisions, evaluating a vendor's roadmap, or trying to distinguish a genuine technology trend from a temporary press cycle, understanding what was actually measured and what the numbers really mean is essential, especially if you're weighing these choices as part of a broader AI strategy consulting effort.

More importantly, these results offer useful context for teams building production AI agents on top of frontier models.


The Short Version

Jalapeño is an application specific inference chip (ASIC), rather than a GPU or a training accelerator. OpenAI and Broadcom co-developed it, TSMC fabricated it, and Celestica integrated it into racks and systems.

The chip was designed to reach tape-out in nine months, an unusually short timeline for custom silicon. At Hot Chips 2026, OpenAI released its first benchmarking data comparing Jalapeño with Nvidia's Blackwell generation GB200 and GB300 systems.

The results showed strong performance in throughput per watt and latency, but they also revealed several important caveats that should be considered before using the headline figures in a board presentation.


Why OpenAI Is Building Its Own Chip

The demand-side rationale was outlined by Greg Brockman on CNBC on the same day Jalapeño was introduced. OpenAI has struggled to secure enough compute through traditional channels, while its six major compute customers expect demand to remain high through at least 2026. That persistent demand is a key part of the supply-side pressure driving OpenAI toward custom silicon.

The cost side of the equation, which has a lot more relevance for those who have to actually build production LLM workloads, is that Broadcom's CEO told Bloomberg the accelerator is delivering roughly 50% of the cost when compared to traditional AI GPUs in initial tests. The difference compounds in inference, since the recurring cost that all AI products pay on each user query goes on infinitely into the future, while training is a one-time capital cost.

The subtext is not hard to see: OpenAI apparently burned through well over $13 billion providing its service to users in 2025, and a custom inference chip is the way to transform that spend into sustainable margins in advance of an expected public offering.

There's also a strategic dimension worth stating: decreasing dependency on Nvidia. OpenAI is far from leaving Nvidia behind, since it still plans to commit to a massive amount of its own GPU compute resources, including an entire gigawatt of Nvidia's Vera Rubin platform for 2026, but it now owns a pricing lever and protection against Nvidia's possible supply problems and margin stacking on every single token processed. This kind of hardware and infrastructure planning often overlaps with broader digital transformation initiatives inside large organizations.


What Jalapeño Actually Is

A few technical facts, drawn from OpenAI's own announcement and the Hot Chips presentation, that separate this from vaporware:

It's an inference-only ASIC. OpenAI has said outright that Jalapeño isn't built to replace Nvidia hardware for training, the far more compute-intensive process of building models. Jalapeño runs already-trained models; it doesn't teach new ones.

700W TDP, roughly 550W sustained. OpenAI disclosed at Hot Chips that the chip is rated at 700 watts, with measured sustained power staying at or below 550W across tested workloads. That's well below Nvidia's flagship parts, which run in the 1,200 to 1,400W range, a gap that matters enormously for data center power and cooling budgets, since 700W parts can typically stay air-cooled where 1,200W-plus parts usually require liquid cooling.

Rack- and pod-scale deployment. OpenAI's planned deployment unit is a 128-chip rack delivering roughly 1.7 exaflops of 4-bit compute and 27.5TB of HBM4, with a full pod scaling to 2,048 ASICs and 15.4TB/s of memory bandwidth per package.

Memory density. Each package pairs its compute die with six HBM4 stacks, totaling around 216 GiB, roughly 50% more memory per watt of rated power than Nvidia's GB300, which carries 288GB of HBM3E at a 1,400W rating.

Raw compute vs. Nvidia's newest die. Independent analysis from SemiAnalysis put Jalapeño's single compute die at about 13.4 PFLOPs of MXFP4 on TSMC's N3P node, versus roughly 17.5 PFLOPs of dense NVFP4 on a similarly sized Rubin die on the same node, so on raw per-die throughput, Nvidia's newest silicon still leads. Jalapeño's edge shows up in the system-level, power-normalized numbers instead.

A multi-generation platform, not a one-off. OpenAI has framed Jalapeño as the first entry in a platform where AI products, models, chips, and memory are developed together going forward, the same full-stack co-design logic behind Apple's M-series chips and Google's TPUs. It's a similar principle to how strong generative AI development services aim to align infrastructure and model choices from the start rather than bolting them together later.


The Benchmarks: What Was Actually Measured

This is the section that requires careful reading, as the two phrases are distinct: "50% cheaper" and "beats Blackwell." OpenAI expressed a more concrete version of the second claim at Hot Chips.

Test results were collected using InferenceX, the public power-normalized benchmark suite created by SemiAnalysis for power comparisons. Normalized to package TDP, Jalapeño's efficiency at 700 watts was compared to Nvidia's 1.2 kilowatts with the GB200 and 1.4 kilowatts with the GB300.

The headline results, corroborated across multiple outlets covering the presentation:

  • 1.5x to 1.9x more throughput per kilowatt than Nvidia's GB200 and GB300 rack systems, and 1.7x to 3.6x lower end-to-end latency, on the InferenceX benchmark suite.
  • On highly interactive, low-latency traffic, the kind ChatGPT actually generates in normal use, Jalapeño ran roughly 2 to 4 times faster than the systems it was tested against.
  • SemiAnalysis, which ran the benchmark alongside OpenAI engineers in OpenAI's own lab, said the part outperformed every competing accelerator it has tested from Nvidia, AMD, and Google.

Three open-weight models were used to stress the chip. Instead of relying on one easy internal option, the team ran OpenAI's GPT-OSS-120B, DeepSeek's R1, and Moonshot AI's trillion-parameter Kimi K2.5.

OpenAI said none of the three models were built for Jalapeño. Engineers also loaded and ran all three on the chip during a brief period, the window between when the engineering silicon reached the lab and when the Hot Chips talk started. OpenAI used that timing as a point against picking a model that had been tuned for this exact hardware. Even so, OpenAI still decided which models to show.


Read the Caveats Before You Quote the Multiplier

A few things temper the headline numbers, and any technical buyer should hold onto them:

It wasn't tested against Nvidia's newest platform. Jalapeño was tested against GB200 and GB300. It was not compared to Vera Rubin, the Nvidia platform meant to run the first gigawatt of Nvidia systems that OpenAI plans to roll out in the second half of 2026. The key point is that this is a look at the prior Nvidia setup, not a comparison to the next generation.

Decoding method mismatch. Jalapeño's benchmark runs relied on single-token prediction. The Nvidia baselines they were compared with often used multi-token prediction in real production systems, a setup that can raise usable throughput by a few times on its own. So the comparison is not made up, but it is still not a perfect match. It probably makes the gap look larger on the slide than what you would see with Nvidia hardware in everyday use.

It's still lab silicon. OpenAI has been clear that final performance testing is ongoing and that a fuller technical report will follow in the coming months. These are engineering sample numbers, not shipping hardware numbers at scale.

Volume is small and slow. OpenAI's own hardware lead estimated very small deployment volumes by the end of 2026, with meaningful volume arriving in 2027. This isn't something enterprises will be renting compute on next quarter.

Training is untouched. Nvidia's GPU position in model training is unaffected by any of this. Jalapeño doesn't compete there, and OpenAI isn't claiming otherwise.

None of that erases the result. A 700W part beating 1,200 to 1,400W parts on throughput-per-watt, even with a decoding-method asterisk, is a strong outcome for a first-generation chip built in nine months. It just means "OpenAI's chip beats Nvidia" is a headline, not the full technical picture.


Jalapeño vs. Nvidia GB200/GB300: Quick Comparison

Jalapeño (OpenAI/Broadcom)Nvidia GB300Nvidia GB200
Package TDP700W rated, ~550W sustained1,400W1,200W
Memory~216 GiB HBM4, 15.4 TB/s per package288GB HBM3EHBM3E
Workload targetInference onlyTraining + inferenceTraining + inference
CoolingAir-cooled feasibleTypically liquid-cooledTypically liquid-cooled
Throughput/kW (InferenceX)Efficiency frontier~0.5 to 0.65x of Jalapeño~0.5 to 0.65x of Jalapeño
Deployment stageEngineering samples, limited 2026 volumeShipping at scaleShipping at scale
Decoding method testedSingle-token predictionMulti-token prediction (typical in production)Multi-token prediction (typical in production)

What This Means for Inference Costs and For You

Here's what matters for teams building AI applications: inference cost is becoming a major part of AI spending. As providers improve their chips and infrastructure, the cost of running models may decrease over time.

But CTOs should not simply wait for cheaper AI. The bigger opportunity is to build efficient AI systems today. Model routing, caching, batching, and smart token usage, often implemented through solid API development practices, can reduce costs significantly.

AI agents can waste money through unnecessary context, repeated tool calls, and retry loops. Better architecture, the kind you'd get from dedicated AI agents services, can often save more than simply choosing a cheaper model. Teams looking to tighten these workflows may also benefit from automation services that cut down on redundant processing before it ever reaches the model.

At RejoiceHub, we help teams decide between frontier APIs, self-hosted models, or a hybrid approach. If you're comparing these options, our machine learning development services team can walk you through self-hosted AI vs. API models.

Ready to Grow?

Accelerate Your Workflows with Custom AI

Book a free consultation session with RejoiceHub. We'll map out a tailored automation roadmap for your company.

Jalapeño in Context: OpenAI's Full-Stack Bet

Jalapeño is not just a side project. OpenAI sees it as part of its larger AI platform, covering products, models, and custom chips. The goal is to support AI workloads at very large data center scales.

This follows a strategy similar to Apple's M-series chips and Google's TPUs. By controlling the chip, model, and product, companies can optimize the entire AI system instead of relying fully on third-party hardware, an approach that mirrors how thoughtful generative AI solutions are built when performance and cost both matter.

But the bigger point is that custom AI chips also depend on key hardware partners. Companies such as Broadcom help design and build these chips, while TSMC provides advanced manufacturing and packaging. These partnerships are becoming critical as AI infrastructure grows, and they underline why solid DevOps consulting services matter just as much as the model itself when scaling AI workloads reliably.


Conclusion

Jalapeño is a real, tested AI inference chip that has shown strong performance against Nvidia's previous-generation flagship chips. However, the results depend on factors such as the benchmark used, decoding methods, and deployment scale.

It is not currently available for enterprises to use directly, so it will not reduce your AI costs immediately. But it shows an important industry trend: AI companies are focusing more on inference speed, power efficiency, and cost.

Over the next 12 to 24 months, this could affect API pricing, model selection, and AI infrastructure decisions.

If you're planning to scale your AI agents, RejoiceHub can help you review model choices, API vs. self-hosted options, and ways to reduce unnecessary token usage through our AI integration and AgentKit builder services. Talk to RejoiceHub's AI strategy team.

Frequently Asked Questions

1. What is OpenAI's Jalapeño chip?

Jalapeño is OpenAI's first custom chip built with Broadcom. It is an inference only ASIC, meaning it runs already trained AI models instead of training new ones, and it was fabricated by TSMC.

2. Is Jalapeño better than Nvidia's GPUs?

Jalapeño beats Nvidia's GB200 and GB300 on throughput per watt and latency in early tests. It was not tested against Nvidia's newer Vera Rubin platform, so this is not a full win over Nvidia yet.

3. How much cheaper is Jalapeño than Nvidia GPUs?

Broadcom's CEO said the chip delivers around 50% lower cost compared to traditional AI GPUs in initial tests. This mainly applies to inference costs, not the one time cost of training models.

4. Does Jalapeño replace Nvidia GPUs for AI training?

No, Jalapeño does not replace Nvidia for training. It is built only for inference, which means running models that are already trained. Nvidia GPUs still handle all of OpenAI's model training work.

5. How much power does the Jalapeño chip use?

Jalapeño is rated at 700 watts, with sustained power staying around 550 watts. This is much lower than Nvidia's GB200 and GB300, which use between 1,200 and 1,400 watts.

6. Can Jalapeño chips be air cooled?

Yes, because Jalapeño runs at 700 watts or less, it can typically use air cooling. Nvidia's higher power GPUs usually need liquid cooling, which adds extra cost and complexity to data centers.

7. When will Jalapeño chips be available at scale?

OpenAI expects only small deployment volumes by the end of 2026. Meaningful, larger scale volume is expected to arrive sometime in 2027, so enterprises cannot rent this hardware right now.

8. Who makes the Jalapeño chip for OpenAI?

Broadcom co-developed Jalapeño with OpenAI, TSMC fabricated the chip, and Celestica handled integrating it into racks and full data center systems for testing and deployment.

9. Why is OpenAI building its own AI chip?

OpenAI wants to lower inference costs, reduce reliance on Nvidia, and secure more compute for growing demand. A custom chip also gives OpenAI more control over pricing and supply.

10. Will Jalapeño lower AI API prices for businesses?

It could, but not immediately. Since Jalapeño is still in limited testing with small deployment volumes, any effect on API pricing will likely take shape over the next 12 to 24 months.

Amrendra kumar profile

Amrendra kumar (Technical Content Writer | AI, Coding & Automation)

Technical Content Writer at RejoiceHub, creating AI, automation, AI agents, coding, and SEO-focused content that makes complex topics clear, useful, and search-friendly.

Published August 26, 202695 views