← All articles
Aug 29, 2026

OpenAI's Jalapeño Chip Benchmark Numbers Beat Nvidia's Blackwell

OpenAI's first Jalapeño chip benchmark results beat Nvidia's Blackwell on inference efficiency, but the company's own data says Blackwell isn't the fair comparison.

On June 24, 2026, OpenAI and Broadcom announced Jalapeño, OpenAI's first custom silicon, an inference-only chip built from scratch as the start of a multi-generation compute platform meant to reduce OpenAI's dependence on Nvidia. That was the announcement. On August 25, the first real lab performance numbers landed, and the openai jalapeño chip benchmark results are the more interesting story: an AI lab's own chip beating the market leader on the metric that actually determines inference cost, with the honest caveat attached in the same breath.

The numbers come from SemiAnalysis's InferenceX benchmark suite, which put Jalapeño up against Nvidia's Blackwell running three real open models, DeepSeek R1, Kimi-K2.5, and GPT-OSS. On tokens generated per megawatt of power, the metric that determines what it actually costs to serve a model at scale, Jalapeño came out ahead of Blackwell across all three. On Kimi-K2.5 specifically, it reached close to 700 tokens per second per user at low concurrency. SemiAnalysis's own framing was blunt: "Jalapeño smokes every other chip" on that metric.

What makes the jalapeño chip benchmark result unusual

The detail worth sitting with is not just that Jalapeño won, it's how. Nvidia's competing setups leaned on Multi Token Prediction (MTP), a speculative-decoding technique that lets a chip guess several tokens ahead and verify them in a batch, trading extra compute for higher throughput. Jalapeño posted its numbers using single-token prediction only, no speculative decoding, no prefill-decode disaggregation. That means the chip beat systems that were given a real efficiency crutch while running without one itself. If OpenAI's team ever adds MTP-style tricks to a future Jalapeño generation, the gap likely widens rather than closes.

None of this happened by accident. SemiAnalysis notes the whole design, from initial silicon work to a chip actually running production-relevant benchmarks, moved fast by ASIC standards, and OpenAI has said it used its own models to help accelerate parts of that development. Broadcom handled the silicon partnership, TSMC manufactures it, and Celestica handles the board, rack, and system integration. That's a real supply chain standing behind a real result, not a slide deck.

Why blackwell isn't the fair fight

Here's the part a lot of the social chatter around this story skipped, and it's the part that matters most if you're actually trying to read the result correctly. SemiAnalysis raises its own caveat rather than burying it: Blackwell isn't really the comparable generation. Nvidia's upcoming Rubin chip uses the same HBM4 memory standard Jalapeño does, which is the real architectural peer. Against Rubin, the two land close to performance-per-dollar parity, not a Jalapeño blowout.

That distinction is the difference between an honest result and a hype cycle. Jalapeño beating Blackwell is genuinely notable, current-generation Nvidia silicon versus a first-generation OpenAI chip is not a fair fight OpenAI was supposed to win, and it won on the metric that costs real money at scale. But treating that as "OpenAI beat Nvidia" full stop, without the Rubin caveat, overstates what today's data actually supports. The more accurate read: OpenAI's first inference chip is already competitive with Nvidia's next real generation, not just its current one. That's a stronger claim than "beats Blackwell" once you see what it's actually being measured against.

What this means if you're building on inference costs

If your own AI systems run on someone else's inference pricing, whether that's an API bill from a foundation model provider or a self-hosted GPU cluster, the openai jalapeño chip benchmark result is a signal worth tracking, not a decision to act on yet. Jalapeño isn't available to rent. It's an internal chip OpenAI built to run its own models more cheaply, not a service you can point your workloads at. What it does tell you: the assumption that Nvidia pricing is the permanent floor for inference cost is getting tested seriously by at least one major lab, and the test is working better than a skeptic would have expected a year ago.

The practical takeaway for anyone tracking AI infrastructure costs is to watch for two things next. First, whether OpenAI's own API pricing moves in response to cheaper internal inference, since a lab running its own silicon at lower cost per token has room to pass some of that through, or simply improve its own margins while pricing stays flat. Second, whether Google, Anthropic, or Amazon follow with their own inference-silicon results this year, since a genuinely competitive Jalapeño result raises the pressure on every other lab still paying full Nvidia rates for the bulk of their serving costs.

The honest caveat is what makes this credible

The detail that separates this story from a typical vendor benchmark brag is that SemiAnalysis, an independent outside benchmarking outfit, published the Rubin caveat in the same piece as the headline number, unprompted. That's not how a marketing benchmark usually reads. It's closer to how a real technical result gets reported: here's what won, here's the fairer comparison, here's what that fairer comparison actually shows. For anyone deciding how much weight to put on any AI hardware claim going forward, that's the pattern worth expecting before trusting the headline number alone.

Sources: SemiAnalysis, "OpenAI Jalapeño: Better Than Nvidia Blackwell", August 25, 2026; OpenAI and Broadcom, original chip announcement, June 24, 2026.

Join the newsletter

AI workflows and systems, straight to your inbox.

No spam. Unsubscribe anytime.