kurt.news

Clean, fast AI news without the hype or doom.

Ai

OpenAI's Jalapeño Chip Beats Nvidia Blackwell on Inference Efficiency at Hot Chips

OpenAI's Jalapeño Chip Beats Nvidia Blackwell on Inference Efficiency at Hot Chips

OpenAI presented Jalapeño chip benchmark results at the Hot Chips conference on August 25, 2026. Testing on SemiAnalysis' InferenceX benchmark showed the chip outperformed an Nvidia Blackwell system on two metrics: tokens per user and throughput per kilowatt.

What the Benchmark Covers

The InferenceX benchmark targets inference workloads specifically. Jalapeño beat the Blackwell comparison system on capacity (tokens per user) and energy efficiency (throughput per kilowatt).

The chip is designed for both high throughput and low latency. Architecturally, it minimizes delays during the prefill and communication phases of inference. Explicit KV cache placement keeps model state local, reducing round-trip overhead within the inference pipeline.

Who Built It

Jalapeño is a collaboration between OpenAI and Broadcom. OpenAI's own AI models assisted in the development process. Richard Ho, OpenAI's head of hardware, leads the program.

OpenAI frames it as a multigenerational platform integrating AI products, models, chips, and memory. That is a longer-term roadmap claim, not a single product announcement.

Deployment Timeline

Jalapeño was first announced in October 2025. Small-volume deployment is planned for late 2026. Significant deployment is expected in 2027.

The benchmark results arrive roughly a year before meaningful scale. Hot Chips presentations typically precede production volumes by that margin.

What to Make of It

Beating Blackwell on inference efficiency is a meaningful claim. Blackwell is the current benchmark for high-throughput AI inference at scale. The comparison is specific enough to be falsifiable.

The caveat: controlled benchmarks and production inference are different environments. Workload variance, cooling constraints, and software stack differences rarely surface in conference presentations. Real-world performance data will not arrive until 2027 at the earliest.

The multigenerational framing is worth noting. OpenAI is describing Jalapeño as infrastructure with a roadmap, not a one-time chip. Whether subsequent generations ship on schedule is a separate question.

Source: Techcrunch