Nvidia still dominates frontier model training. But OpenAI just showed numbers suggesting that on inference — the part that bills every time someone uses ChatGPT — a custom-designed chip can beat it on its own turf.
OpenAI unveiled its first internally designed inference processor, built in collaboration with Broadcom, codenamed "Jalapeño." Unlike a general-purpose GPU like Nvidia's, this chip was designed specifically for the needs of OpenAI's inference systems — meaning, for serving responses from already-trained models, not for training them.
Richard Ho, OpenAI's VP of Hardware, summed it up this way: the chip "achieves high throughput and low latency simultaneously, something unprecedented in the industry" — normally optimizing for one of those two factors means sacrificing the other.
Jalapeño is inference-only. Nvidia's dominance in frontier model training remains intact — this chip doesn't compete there. Rollout will also be gradual: it starts shipping in the remainder of 2026, with real volume production not until 2027. It's not an immediate replacement for OpenAI's Nvidia infrastructure, it's a second path that reduces dependency.
For teams that consume models via API instead of running their own infrastructure, the expected effect is indirect but real: more competition in the inference chip market historically pushes API prices down, and reduces the risk that a single provider's (Nvidia's) compute availability becomes the whole industry's bottleneck — the same diversification pattern we already saw with the AMD-Anthropic deal.
Carlos Montiel is an enterprise AI solutions architect. He implements LLMs, Agents, RAG, and orchestrators for companies across Guatemala and Latin America. Reach out for a consultation.
Contact Carlos Montiel