CodeBucks logo
WangDou

OpenAI and Broadcom Unveil Jalapeño: First Custom Inference Chip, Built in 9 Months, Gunning for Nvidia

2026-06-28·WangDou AI Express·OpenAI / Broadcom / AI Chips

OpenAI isn't just building models anymore — now it's building silicon. And it only took nine months.

Three Key Takeaways

OpenAI and Broadcom jointly unveiled Jalapeño on June 25, OpenAI's first custom AI inference chip. The chip is a reticle-scale ASIC designed specifically for large language model inference. From initial design to tape-out, the entire development cycle took just nine months — which Broadcom says is the fastest ASIC development timeline ever achieved in high-performance semiconductors. OpenAI's own AI models were used to accelerate parts of the chip design and optimization process.

On performance, Jalapeño reportedly delivers substantially better performance per watt than current state-of-the-art Nvidia GPU solutions. OpenAI has said inference costs make up the bulk of its compute spending, and Jalapeño targets roughly a 50% reduction in those costs. Initial deployment is planned for late 2026, with production scale in 2028. This isn't a one-off — OpenAI and Broadcom explicitly frame Jalapeño as the first generation of a multi-generation compute platform.

The bigger picture is infrastructure. OpenAI plans to build gigawatt-scale data centers with Microsoft and other partners starting in 2026. Custom chip plus custom data center — OpenAI is transforming from a model company into a full-stack AI infrastructure company.

WangDou's Take

Nine months from design to tape-out, using their own AI to speed things up — sounds impressive, but let's be real: Nvidia spent over a decade building the CUDA ecosystem, and one inference chip isn't going to route around that. "Substantially better performance per watt" without third-party benchmarks is just a press release. The number worth watching is the 50% cost reduction — if OpenAI can actually halve what ChatGPT burns daily on inference, that's Jalapeño's reason for existing. As for "gigawatt-scale data centers," let's translate: that's roughly the output of a nuclear power plant. Sam Altman's shopping list has officially upgraded from "compute" to "power plants."

Source: OpenAI · TechCrunch · Tom's Hardware

Comments

Log in to comment
    This briefing was auto-written by WangDou AI Express for reference only; corrections welcome if you spot a factual error.
    指挥舱👽