OpenAI Unveils Custom Inference Chip 'Jalapeño' to Reduce GPU Costs
OpenAI and Broadcom have announced a co-developed custom silicon chip manufactured by TSMC. Titled 'Jalapeño', it is engineered to run large-scale inference workloads at 50% of the cost of current-generation GPUs.
In a significant shift for the artificial intelligence hardware market, OpenAI and Broadcom have officially announced the completion of their co-developed inference silicon, codenamed Jalapeño. Manufactured on TSMC's cutting-edge process node, the custom chip is designed to run live model executions at half the financial cost of legacy GPU infrastructure.
Challenging NVIDIA's Server Dominance
For the past three years, the expansion of Generative AI has been bottlenecked by the availability and cost of high-end graphics processors. By co-designing its own Application-Specific Integrated Circuit (ASIC) with Broadcom, OpenAI is attempting to build a custom runtime platform optimized exclusively for transformer operations. Jalapeño drops generic computing graphics logic in favor of dedicated tensor arithmetic and high-bandwidth interconnects.
Key Specifications and Architecture
While full architectural documentation remains confidential, industry disclosures outline several key technical specs for the Jalapeño silicon:
- TSMC Packaging: Leverages TSMC's advanced CoWoS (Chip-on-Wafer-on-Substrate) packaging to stack high-bandwidth memory directly adjacent to the compute core.
- Broadcom Interconnects: Integrates Broadcom's custom PCIe and Ethernet switching fabrics to allow low-latency multi-node scaling across data center desks.
- Transformer-Specific Core: Built specifically to accelerate matrix multiplications and KV-cache lookups, reducing energy consumption during inference.
"ASIC optimization is the natural evolution of AI scaling. You do not build a general-purpose gaming card to run a pipeline that only executes static matrix addition and token selection."
What It Means for AI Cost Structures
Currently, running millions of queries through models like GPT-4o represents an immense operational expense. OpenAI projects that migrating its primary production nodes to Jalapeño clusters will lower inference power draw and execution latency by 50%. This cost reduction could enable developers to deploy larger, more complex agentic loops without exponential infrastructure costs.
Initial deployments of the Jalapeño chip are scheduled for early autumn 2026 inside Microsoft Azure datacenters.
Subscribe for Updates
Get official press announcements and version releases sent directly to your email.