OpenAI's Jalapeño Chip Went Concept-to-Silicon in 20 Months, Aided by Its Own LLMs
OpenAI has revealed Jalapeño, its first in-house AI accelerator, built in partnership with Broadcom and aimed at its inference fleet. The chip delivers up to 13.4 petaflops of 4-bit compute, taps 232 GB of leading-edge memory at 15.4 TB/s, and — per OpenAI’s own benchmarks — cuts end-to-end latency by as much as 3.6x versus the Nvidia GB300 it currently leans on, while drawing less power. Those figures are unproven at scale, but the more striking claim is speed of development: the design went from first architecture to first silicon in under 20 months, with just nine months between initial RTL and tape-out, and it was done by a team averaging fewer than 100 people (excluding Broadcom).
The acceleration came from pointing OpenAI’s own models at the parts of chip design that resemble software. The front-end workflow was built around XLS, an open-source high-level synthesis toolchain originally created at Google by an OpenAI staffer now on the project; because designers write in software-like languages (DSLX, C++) that XLS compiles to Verilog, LLMs could iterate on it far more readily than on raw hardware description. The payoff showed up after first silicon arrived in May, when internal models optimized a DeepSeek attention-kernel benchmark from 0.31% of theoretical peak to nearly 89% in roughly 40 hours — a repeatable result OpenAI says will let it compress the gap between first chips and production ramp. OpenAI split the work with Broadcom, keeping system architecture, memory hierarchy, and networking in-house while handing physical design ‘from the gates onward’ to its partner.
Outside experts call the timeline credible and ‘likely best in class,’ but temper the story: Verkor.io’s founders argue Broadcom’s involvement was essential and that a from-scratch effort couldn’t have matched it. The project also spanned a fast-moving model generation, starting with o3-class systems and finishing on precursors to GPT-6 Astra — newer models that can work directly in Verilog without XLS translation and are nearing autonomous use of proprietary design tools. LLMs proved far less useful on backend tasks like routing, clock, and power, which Broadcom largely owned. OpenAI, which used undisclosed internally fine-tuned models, says the lessons will feed its commercial LLMs, and that Astra and its successors will be ‘very good at chip design.’
Read the full article
Continue reading at Hacker News →This is an AI-generated summary. Read the original for the full story.