OpenAI's first custom chip "Jalapeño" reportedly beats Nvidia's Blackwell and Rubin in inference benchmarks
In shortOpenAI claims Jalapeño delivers 1.5x to 1.9x more AI work per watt at peak throughput across all three tested models, with 1.7x to 3.6x lower end-to-end latency than the best commercially available systems.
What happened
Even here, Jalapeño squeezes out more output tokens per megawatt than Vera Rubin, even though Nvidia's accelerator uses the multi-token prediction optimization that Jalapeño hasn't adopted yet. Nvidia and AMD have already published results with larger models like Deepseek V4 Pro and Kimi K3 that haven't been tested on Jalapeño yet. OpenAI CFO Sarah Friar says the chip fits into a broader compute strategy where data centers, chips, models, the developer platform, products, and devices all work as one integrated system.
Why it matters
Evidence about OpenAI's first custom chip "Jalapeño" reportedly beats Nvidia's can change model or system selection. Teams should reproduce the result on their own workload and compare failure cases before adopting it.
Who it affects
Researchers evaluating OpenAI's first custom chip "Jalapeño" reportedly beats Nvidia's
The bigger picture
OpenAI's first custom chip "Jalapeño" reportedly beats Nvidia's fits a broader move from headline benchmarks toward reproducible evidence, disclosed failure modes, and tests that resemble real operating conditions.
What happens next
- Watch for independent reproduction of OpenAI's first custom chip "Jalapeño" reportedly beats Nvidia's, including failure cases and results on workloads that were not selected by the authors.
This information is still under review.