AivexaNewsSearch
AI news for builders and product teamsChecked every hour

How OpenAI Used Its Own LLMs to Design Its Jalapeño Chip

Collected Oct 1, 2026

OpenAI fully unveiled Jalapeño, its debut AI accelerator chip, on 25 August. According to benchmarks cited by OpenAI, Jalapeño delivers up to 13.4 petaflops of 4-bit compute, accesses 232 gigabytes of memory at 15.4 terabytes per second, and can reduce end-to-end latency by up to 3.6 times compared with Nvidia's GB300, a chip OpenAI currently relies on, while consuming less power. Whether those figures translate into real-world gains once Jalapeño enters widespread service in OpenAI's inference fleet remains to be seen.

Jalapeño moved from first architecture concept to first silicon in under 20 months, with only nine months between the first RTL code and tape-out. Richard Ho, OpenAI's vice president of hardware, said the design team averaged fewer than 100 people and stands at roughly 100 today, excluding Broadcom staff. OpenAI handled end-to-end system design, including the inference accelerator, memory hierarchy and networking; Broadcom handled physical design from the gates onward.

OpenAI built its front-end workflow around Accelerated Hardware Synthesis (XLS), an open-source high-level synthesis chain originally developed at Google, which converts languages such as DSLX and C++ into Verilog. Chris Leary, an OpenAI staff member who started XLS at Google, said the AI was better at software-like work. After first chips returned from the foundry in May, internal models designed software for benchmarks including SemiAnalysis's InferenceX; on DeepSeek's multi-head latent attention kernel benchmark, performance rose from 0.31 percent to 88.94 percent of the theoretical ceiling in roughly 40 hours. At IEEE Hot Chips 2026, Ho and Leary reported a 10 percent area reduction for matrix multiplication units versus an optimized human baseline.

Leary said the project began with models such as o3 and later used precursors to GPT-6 Astra, which can work directly in Verilog. Ho confirmed internal chip-design-tuned LLMs were used but declined to detail them. Verkor.io co-founders David Chin and Ravi Krishna called the schedule credible and the speed impressive, with Chin saying Broadcom's help was essential. Andrew Kahng of UC San Diego called it likely best in class today. Ho and Leary said chip design cannot be fully automated.

Read at IEEE Spectrum · AI

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

On 25 August, OpenAI fully unveiled Jalapeño, the company’s debut AI accelerator chip. Jalapeño delivers up to 13.4 petaflops of 4-bit compute and accesses 232 gigabytes of the most advanced memory available, linking to it at a blazing 15.4 terabytes per second. Benchmarks cited by OpenAI show that Jalapeño can reduce end-to-end latency (the time between prompt to last token) by up to 3.6 times when compared to Nvidia’s GB300 —a chip the company currently relies on—and do so while consuming less power. Whether thes