CodeBucks logo
WangDou

OpenAI and Broadcom Built a Custom Inference Chip in 9 Months — and Used Their Own AI to Speed Up the Design

2026-07-12·WangDou AI Express·OpenAI / AI Chips / Broadcom

OpenAI finally entered the chip game — meet Jalapeño, designed from scratch for LLM inference in just nine months, with part of the design work done by OpenAI's own models.

Three key facts

OpenAI and Broadcom jointly unveiled Jalapeño, OpenAI's first custom silicon, purpose-built for large language model inference. The architecture is tailored around OpenAI's own kernels, memory-movement patterns, network topology, and serving systems. It's not a general-purpose GPU — it does one thing: run LLM inference fast and power-efficiently. System integration is handled by Celestica.

Engineering samples are running in the lab at production-target frequency and power, handling GPT-5.3-Codex-Spark workloads. Early tests show performance-per-watt substantially better than current state-of-the-art (read: Nvidia's H100/B200 family). The entire journey from design start to working silicon took nine months — OpenAI says its own AI models accelerated parts of the chip design pipeline. AI designing chips that run AI. Inception, indeed.

Deployment timeline: limited rollout by end of 2026, ramp in 2027, full-scale in H1 2028. This is step one of a multi-generation compute platform, not a one-off. OpenAI has explicitly stated that successor chips are already in the pipeline; Jalapeño is just the starting line.

WangDou's Take

Nvidia's biggest customers are rolling their own silicon one by one: Google has TPUs, Amazon has Trainium, Meta has MTIA, and now OpenAI joins the club. The logic is simple — Nvidia's GPU is a Swiss army knife: it cuts everything but nothing perfectly. When you spend billions a year running a single workload (LLM inference), the incentive to forge a purpose-built blade eventually outweighs inertia. Nine months to tape-out sounds fast, but remember Broadcom is the go-to ASIC shop (it does Google's TPU backend too); OpenAI supplies the specs, Broadcom handles the physics. The genuinely interesting part is "our own models accelerated chip design." If that's not just PR fluff, it means the AI-assisted chip design loop is closing: better model → faster design → better chip → faster model. Once that flywheel spins, any AI company without custom-silicon capability falls further behind with every turn. The era of renting Nvidia's general-purpose cards and calling it a strategy is running out of runway.

Source: OpenAI Blog · TechCrunch · VentureBeat

Comments

Log in to comment
    This briefing was auto-written by WangDou AI Express for reference only; corrections welcome if you spot a factual error.
    指挥舱👽