Jalapeño: OpenAI's First In-House Chip Claims to Beat Nvidia Blackwell on Efficiency

By Carlos Montiel | Enterprise AI Specialist
Leer en español →
Published: 2026-08-26 | By: Carlos Montiel | Reading time: ~4 minutes

Nvidia still dominates frontier model training. But OpenAI just showed numbers suggesting that on inference — the part that bills every time someone uses ChatGPT — a custom-designed chip can beat it on its own turf.

What Jalapeño Is

OpenAI unveiled its first internally designed inference processor, built in collaboration with Broadcom, codenamed "Jalapeño." Unlike a general-purpose GPU like Nvidia's, this chip was designed specifically for the needs of OpenAI's inference systems — meaning, for serving responses from already-trained models, not for training them.

The Benchmark Numbers

# Results reported by OpenAI, compared to Nvidia Blackwell: # - 1.5x to 1.9x more AI work per watt # - 1.7x to 3.6x lower latency # - Especially strong in low-latency scenarios # with high throughput simultaneously

Richard Ho, OpenAI's VP of Hardware, summed it up this way: the chip "achieves high throughput and low latency simultaneously, something unprecedented in the industry" — normally optimizing for one of those two factors means sacrificing the other.

What Doesn't Change

Jalapeño is inference-only. Nvidia's dominance in frontier model training remains intact — this chip doesn't compete there. Rollout will also be gradual: it starts shipping in the remainder of 2026, with real volume production not until 2027. It's not an immediate replacement for OpenAI's Nvidia infrastructure, it's a second path that reduces dependency.

Why OpenAI needed this: every percentage point of energy efficiency on inference translates directly into margin, because unlike training (a one-time expense), inference is a cost that runs with every query from every user, every day, forever.

What It Means for Anyone Buying AI Compute

For teams that consume models via API instead of running their own infrastructure, the expected effect is indirect but real: more competition in the inference chip market historically pushes API prices down, and reduces the risk that a single provider's (Nvidia's) compute availability becomes the whole industry's bottleneck — the same diversification pattern we already saw with the AMD-Anthropic deal.

Carlos Montiel
Enterprise AI Solutions Architect
Specialist in LLMs, Agents, and Orchestration
guatemalia.com/en/#contact · info@guatemalia.com

Need to implement AI at your company?

Carlos Montiel is an enterprise AI solutions architect. He implements LLMs, Agents, RAG, and orchestrators for companies across Guatemala and Latin America. Reach out for a consultation.

Contact Carlos Montiel

info@guatemalia.com