Short answer: OpenAI’s new Jalapeño chip beat Nvidia’s Blackwell-generation systems on performance-per-watt and latency in OpenAI’s own inference tests, shown at Hot Chips on August 25, 2026. But Jalapeño won’t ship in volume until 2027, and the more honest comparison is against Nvidia’s newer Rubin platform, not Blackwell. So the win is real, but it’s not the knockout the headlines make it sound like.
I’ve spent the past few days reading through OpenAI’s own posts, the SemiAnalysis teardown, and what Nvidia-side analysts are saying. Here’s what actually matters, without the hype.
What Is the OpenAI Jalapeño Chip?
Jalapeño is OpenAI’s first custom AI chip, built specifically for inference — the part of AI where a model reads your prompt and writes a response, rather than the training phase where it learns. OpenAI built it with Broadcom and Celestica, and it’s the first product from a deal the two companies signed in October 2025 to co-develop 10 gigawatts of custom accelerators.
This isn’t a side project. It’s the first piece of a multi-generation compute platform where OpenAI’s models, chips, and memory get designed together instead of bolted on afterward. Richard Ho, OpenAI’s VP of hardware, said the design went from initial RTL to tapeout in roughly nine months, with OpenAI’s own models helping write and debug the chip design. That’s fast for custom silicon. A typical chip project takes two to three years.
Jalapeño is an ASIC — application-specific silicon — but it’s not locked to OpenAI’s own models. It’s a general-purpose inference chip that ran GPT-OSS-120B, DeepSeek R1, and Kimi K2.5 in early tests, plus a version of Doom ported over using Codex prompts. That last part is a party trick, but it does show the chip isn’t a one-model gimmick.
What Is Nvidia Blackwell, and Why Compare Against It?
Blackwell is Nvidia’s current-generation AI GPU architecture, sold in rack-scale systems like the GB200 NVL72 and GB300 NVL72. These racks are what most large AI labs, including OpenAI itself, have been running on for the last two years. Blackwell uses HBM3e memory and is the dominant chip behind most commercial LLM inference today.
Nvidia is already moving past it. Its next platform, Rubin, uses HBM4 memory and started shipping to customers around the same time Jalapeño’s benchmarks came out. That timing matters a lot for how you should read this comparison — more on that below.
The Benchmark Numbers, Straight
OpenAI shared its first real performance data at Hot Chips 2026, tested against SemiAnalysis’s InferenceX benchmark suite. Here’s what came back:
| Metric | Jalapeño vs Nvidia Blackwell |
|---|---|
| Peak throughput (“AI work” per watt) | 1.5x to 1.9x more |
| End-to-end latency | 1.7x to 3.6x lower |
| Ultra-low-latency interactive inference | 2.1x to 4.1x faster |
| Tokens/sec/user (DeepSeek R1, low concurrency) | 700+, using single-token prediction only |
| Per-rack compute | 1.7 exaFLOPS (4-bit), 128 accelerators |
| Per-rack memory | 27.5 TB HBM4, ~2 PB/s bandwidth |
That single-token-prediction number is the detail I keep coming back to. Jalapeño hit those speeds without speculative decoding or prefill-decode disaggregation — tricks other chipmakers lean on to boost their numbers. SemiAnalysis called this the more apples-to-apples test, and on it, Jalapeño beat every competitor it was tested against.
Nvidia’s current racks aren’t standing still, though. AMD and Nvidia’s latest systems still deliver 1.46x to 2x more raw compute and up to 12% more total memory than a Jalapeño rack. Jalapeño’s edge is efficiency and latency, not brute-force power. That’s an important distinction if your workload is training-heavy rather than inference-heavy.
The Catch Nobody’s Headline Mentions
Here’s where I have to be honest about what this comparison actually is. SemiAnalysis itself called the Jalapeño-vs-Blackwell matchup “somewhat incomplete and unfair,” for a simple reason: Blackwell uses older HBM3e memory, while Jalapeño uses HBM4. Nvidia’s Rubin platform also uses HBM4 and is already shipping to real customers. Jalapeño, by contrast, is still at the engineering-sample stage.
So the fairer fight is Jalapeño vs Rubin, not Jalapeño vs Blackwell. When SemiAnalysis pushed on that comparison, Jalapeño still held up well against Rubin on efficiency, but the gap shrinks and the story gets more nuanced. Semiconductor analyst Dylan Patel put it bluntly: it’s not just Blackwell that’s been challenged, Rubin has too — but that claim deserves the same scrutiny as OpenAI’s own numbers, since these are still self-reported results from a chip that hasn’t shipped.
A few other caveats worth knowing before you repeat this stat at a dinner party:
- Timing is not neutral. OpenAI announced these results one day before Nvidia’s Q2 FY2026 earnings call. That doesn’t make the numbers false, but it does explain why the story spread so fast.
- The models tested aren’t OpenAI’s newest frontier models. InferenceX runs GPT-OSS-120B, DeepSeek R1, and Kimi K2.5 — useful benchmarks, but not the largest, newest models on the market.
- Volume production is 2027. Small-scale deployment inside OpenAI’s own infrastructure starts by the end of 2026. If you’re a business deciding what to buy this year, Jalapeño isn’t a purchasable option yet — it’s an internal OpenAI tool.
- Nvidia isn’t going anywhere. Even the analysts most bullish on Jalapeño agree Nvidia still owns the “vast majority” of AI compute worldwide, plus the CUDA software lock-in that keeps developers on Nvidia hardware.
Why This Still Matters, Even Before It Ships
I’d argue the benchmark numbers matter less than what Jalapeño proves is possible. For years, the assumption was that only Nvidia’s silicon and CUDA ecosystem could realistically serve frontier-scale inference. Jalapeño is evidence that a hyperscaler can design a competitive inference chip from scratch in under a year, with meaningful help from AI tools during the design process itself.
That’s the part that should worry Nvidia’s margins, not the specific 1.9x number. Inference is the fastest-growing segment of AI spending right now, and it’s also the segment where a company like OpenAI has the most incentive to cut its own costs. Every dollar OpenAI doesn’t spend on Nvidia GPUs for inference is a dollar it keeps. Yole Group’s Adrien Sanchez told CNBC this is a genuine threat to Nvidia’s inference margins specifically, even if training workloads stay firmly on Nvidia hardware for now.
Nvidia’s own hardware chief has publicly maintained the company remains an important partner to OpenAI, and OpenAI’s hardware lead Richard Ho said the same — this isn’t a breakup, it’s a hedge. OpenAI still needs enormous amounts of Nvidia compute for training and won’t walk away from that relationship anytime soon.
Jalapeño vs Blackwell: Quick Comparison Table
| OpenAI Jalapeño | Nvidia Blackwell (GB200/GB300) | |
|---|---|---|
| Built for | Inference only | Training + inference |
| Status (Aug 2026) | Engineering samples | Shipping, in production |
| Memory | HBM4 | HBM3e |
| Partner ecosystem | Broadcom, Celestica | CUDA, broad third-party support |
| Best case shown | 1.5x–1.9x perf/watt over Blackwell | Higher raw compute per rack |
| Availability to outside buyers | Not sold externally — internal OpenAI use | Sold to any customer |
| Fair rival | Nvidia Rubin | — |
Is Jalapeño Actually Better Than Blackwell?
For inference workloads specifically, on the metrics OpenAI chose to test, yes — Jalapeño came out ahead. But “better” depends entirely on what you need. If you need raw training throughput, or you need a chip you can actually buy today, Blackwell (or Rubin) is still the real-world answer. If you’re OpenAI, running your own inference at massive scale, cutting cost-per-token by even a small percentage saves an enormous amount of money over a year — and that’s the problem Jalapeño was built to solve.
Our Take
I don’t think this is a “Nvidia is finished” moment, and I’m skeptical of anyone framing it that way. It’s a “the moat is shrinking” moment. Custom silicon from Google, Amazon, and now OpenAI keeps chipping away at the idea that Nvidia is the only realistic option for serious AI infrastructure. That’s good news if you’re paying inference bills and bad news if you’re holding Nvidia stock hoping margins never compress.
What I’d actually watch next: OpenAI’s second and third chip generations, which Ho confirmed are already in development, and how Nvidia’s Rubin platform performs once it’s fully in production rather than just “starting to ship.” That comparison, not this one, is the one that will actually settle the argument.
If you’re tracking AI infrastructure costs for your own product or business, don’t make purchasing decisions off this benchmark alone — none of this hardware is available to buy yet outside OpenAI’s own data centers. Watch for OpenAI’s promised technical report in the coming months, and Nvidia’s Rubin production numbers, before drawing conclusions.
FAQ SECTION:
Q1: What is the OpenAI Jalapeño chip used for? A1: Jalapeño is built for AI inference — running trained models to generate responses — not for training new models. OpenAI designed it to cut latency and power use when serving ChatGPT and API requests at scale, working alongside its existing Nvidia hardware rather than replacing it entirely.
Q2: Is Jalapeño faster than Nvidia Blackwell? A2: In OpenAI’s benchmarks on the InferenceX test suite, Jalapeño delivered 1.5x to 1.9x more throughput per watt and up to 3.6x lower latency than Nvidia Blackwell systems. Analysts note the comparison favors Jalapeño somewhat, since Blackwell uses older HBM3e memory than Jalapeño’s HBM4.
Q3: When will Jalapeño be available? A3: OpenAI plans small-scale deployment inside its own infrastructure by the end of 2026, with broader rollout in 2027. It is not being sold to outside companies — it’s built for OpenAI’s internal compute needs only.
Q4: Does Jalapeño replace Nvidia GPUs at OpenAI? A4: No. OpenAI’s hardware lead has said Nvidia remains an important partner, and OpenAI still relies heavily on Nvidia GPUs for training. Jalapeño targets inference workloads specifically, reducing but not eliminating OpenAI’s dependence on Nvidia.
Q5: Why did OpenAI compare Jalapeño to Blackwell instead of Rubin? A5: Blackwell is Nvidia’s current production hardware running most real-world workloads today, so it’s a practical comparison. But SemiAnalysis pointed out that Rubin, Nvidia’s newer HBM4-based platform, is the more fair technical match since it uses the same memory generation as Jalapeño.
Q6: Who makes the Jalapeño chip with OpenAI? A6: Jalapeño was co-developed with Broadcom, which handled the silicon and networking design, and Celestica, which handles systems integration. The partnership stems from a 10-gigawatt custom accelerator deal OpenAI and Broadcom signed in October 2025.
Q7: Will Jalapeño hurt Nvidia’s stock or business? A7: Analysts see it as a threat to Nvidia’s inference-market margins specifically, since that segment is growing fastest. Nvidia still controls the vast majority of AI compute and benefits from CUDA’s ecosystem lock-in, so the near-term business impact is expected to be limited.