OpenAI and Broadcom Unveil "Jalapeno" - The LLM-Optimized Inference Chip That Could Reshape AI Infrastructure
OpenAI and Broadcom have officially unveiled "Jalapeno," a custom LLM-optimized inference chip at the heart of a program codenamed Nexus - marking OpenAI's most serious move yet to own its silicon stack and reduce dependence on Nvidia.
OpenAI and Broadcom Unveil "Jalapeno" - The LLM-Optimized Inference Chip That Could Reshape AI Infrastructure
OpenAI has pulled back the curtain on "Jalapeno," the custom inference chip it has been co-developing with Broadcom. The announcement marks the most significant step yet in the company's years-long push to build silicon specifically tuned to its own large language models - and it signals a broader inflection point in how the industry thinks about AI hardware.
From Rumor to Reality: The Jalapeno Chip
The chip's existence has been an open secret since Reuters first reported on the partnership in 2024. The two companies made it official on October 13, 2025, when they announced a formal strategic collaboration 1. As CNBC reported, OpenAI and Broadcom had already been collaborating for 18 months on co-designed chips before going public 2.
Now branded "Jalapeno" - a codename first reported by The Information and cited by The Decoder 3 - the first chip is part of a larger program internally codenamed "Nexus." The chip is designed to run OpenAI's models more efficiently than Nvidia's hardware, though mass production is not expected until 2027 3. The full Nexus program targets 10 gigawatts of data center capacity 1.
Per the official joint announcement: "OpenAI will design the accelerators and systems, which will be developed and deployed in partnership with Broadcom. The racks, scaled entirely with Ethernet and other connectivity solutions from Broadcom, will meet surging global demand for AI, with deployments across OpenAI's facilities and partner data centers." 4
Technical Architecture
The Jalapeno chip's precise specs have not been officially published by OpenAI or Broadcom. Based on reporting from SemiWiki and consistent with industry norms for this class of accelerator, the chip is expected to feature a systolic array architecture optimized for AI inference, high-bandwidth memory, and manufacturing on TSMC's 3nm process node 5. A second-generation chip is reportedly planned.
The inference-first orientation is the key design choice: unlike Nvidia's general-purpose GPUs, which are optimized across training and inference workloads, Jalapeno is narrowly scoped to the inference tasks that consume the lion's share of OpenAI's daily compute budget.
The economics behind this focus are compelling. CNBC reports that industry estimates peg the cost of a 1-gigawatt data center at roughly $50 billion, with $35 billion of that typically allocated to chips at Nvidia's current pricing 2. By baking model-serving assumptions directly into silicon - attention patterns, KV cache access, token generation pipelines - OpenAI aims to extract far better cost-per-token ratios than commodity hardware allows.
A $180 Billion Program With an $18 Billion First Phase
The scale of this program is difficult to overstate. The $18 billion figure refers specifically to the first phase, which covers approximately 1.3 gigawatts of data center capacity 3. According to The Decoder, which cited The Information, the full 10-gigawatt Nexus plan could cost up to $180 billion in chip production alone - not including data center construction, power, and other infrastructure 3.
Deployment of the first racks is targeted to begin in the second half of 2026, with full rollout expected by end of 2029 6.
Financing remains a live challenge. According to The Information as cited by The Decoder and Data Center Dynamics, Broadcom has required that Microsoft commit to purchasing roughly 40 percent of the chips before it will fund the first phase 3. Under the proposed arrangement, Microsoft would install the chips in its data centers and lease them back to OpenAI 3.
Sachin Katti, OpenAI's head of compute, reportedly told colleagues in an internal message that requiring Microsoft's purchase commitment makes the deal "financially unattractive" and that "this business structure is likely unworkable for subsequent generations of chips" 7. However, the company is pressing ahead for the strategic upside.
OpenAI Is Not Alone - But It Is Late
OpenAI is entering a club that hyperscalers joined years ago. The tie-up with Broadcom places OpenAI among cloud-computing giants such as Google and Amazon that are developing custom chips to reduce dependence on Nvidia's costly processors 8.
Google's seventh-generation Ironwood TPU - an inference-first design - now powers Gemini across Search, Workspace, and Cloud. Amazon's Trainium architecture is running at scale for Anthropic. And Microsoft has now deployed its own custom inference chip, Maia 200, unveiled on January 26, 2026 9.
Built on TSMC's 3nm process with over 140 billion transistors, Maia 200 delivers over 10 petaFLOPS of FP4 compute and features 216 GB of HBM3E memory at 7 TB/s bandwidth 10. Microsoft's official blog confirmed that Maia 200 "will serve multiple models, including the latest GPT-5.2 models from OpenAI" via Azure 10. It is already deployed in Microsoft's US Central datacenter near Des Moines, Iowa, with US West 3 in Phoenix next 10.
Analysts caution that custom chips don't pose a near-term threat to Nvidia's dominance. As NBC News noted, similar efforts by Microsoft and Meta have run into delays or failed to match Nvidia chip performance, and custom chips do not pose a threat to Nvidia's dominance in the short term 8. The software gap is also real: Nvidia's CUDA platform remains the default target for nearly every AI framework in use today, and migrating off it means rewriting core libraries, retraining engineers, and adapting models to new hardware.
Why It Matters
The Jalapeno program is a declaration that the era of generic GPU procurement is over for frontier AI labs. As OpenAI stated in its official announcement: "By designing its own chips and systems, OpenAI can embed what it's learned from developing frontier models and products directly into the hardware, unlocking new levels of capability and intelligence." 4 That's not just marketing language - it reflects a genuine architectural opportunity. When you control both the model and the chip, you can co-design attention mechanisms, memory hierarchies, and numerical formats in ways that third-party silicon simply cannot accommodate.
For engineers building on top of OpenAI's APIs, the practical implication is lower inference costs and higher throughput over time - assuming the program stays on schedule. For the broader AI hardware ecosystem, Jalapeno is further evidence that the next competitive frontier is not just model quality, but the silicon stack underneath it. OpenAI CEO Sam Altman announced at the company's Dev Day in October 2025 that ChatGPT had surpassed 800 million weekly active users 11; at that scale, even marginal improvements in per-token efficiency translate to billions of dollars annually. That alone justifies the bet.
Sources
- 1. OpenAI and Broadcom announce strategic collaboration — OpenAI
- 2. Broadcom stock pops 9% on OpenAI custom chip deal — CNBC
- 3. Broadcom reportedly won't build OpenAI's custom chip unless Microsoft buys 40% — The Decoder
- 4. OpenAI and Broadcom announce strategic collaboration — Broadcom Investors
- 5. OpenAI taps Broadcom to build its first AI processor — SemiWiki
- 6. OpenAI and Broadcom to develop and deploy 10GW of custom AI accelerators — Data Center Dynamics
- 7. OpenAI and Broadcom in discussions over financing of $18bn custom chip project — Data Center Dynamics
- 8. OpenAI taps Broadcom to build its first AI processor — NBC News
- 9. Microsoft Unveils Maia 200 AI Chip on TSMC 3nm — TrendForce
- 10. Maia 200: The AI accelerator built for inference — Official Microsoft Blog
- 11. Sam Altman says ChatGPT has hit 800M weekly active users — TechCrunch
This article was researched and drafted by an AI writer agent (claude-sonnet-4-6) and reviewed by an editor agent before publishing.