OpenAI and Broadcom unveiled Jalapeño on June 24, 2026, OpenAI’s first purpose-built AI inference chip. The announcement confirmed months of industry speculation: the world’s leading AI lab is no longer satisfied renting compute exclusively from Nvidia. It is now building its own.
Jalapeño is a large custom ASIC designed from the ground up around the specific demands of large language model inference. OpenAI led the architecture and system-level design, while Broadcom handled silicon implementation, high-speed connectivity, and networking (including its Tomahawk silicon). Celestica is supporting board, rack, and system integration. Engineering samples have already been delivered to OpenAI and are running production workloads in the lab. Initial deployment is targeted for late 2026, with broader rollout planned at gigawatt scale alongside data center partners.
Why Inference, Why Now
The distinction between training and inference is critical. Training is a periodic, capital-intensive process. Inference, running models for every user prompt, is continuous and, at OpenAI’s scale with hundreds of millions of ChatGPT users, represents a massive and growing cost center.
Nvidia’s H100 and B200 GPUs are versatile general-purpose accelerators capable of training, inference, graphics, and scientific computing. That flexibility comes at a premium. An ASIC optimized solely for transformer-based LLM inference can be tailored to the exact data movement patterns, memory hierarchies, and precision requirements of these workloads. Jalapeño specifically targets the data-movement bottleneck between compute and memory, an approach similar to those pursued by Cerebras and Groq, but executed at OpenAI’s model and scale.
OpenAI President Greg Brockman, appearing on CNBC with Broadcom CEO Hock Tan, described the chip’s early results directly: it delivers real improvements in performance per watt and performance per dollar. On X, he noted simply that “Perf per watt [is] looking incredible.”
The chip is designed to run OpenAI’s most important inference workloads close to theoretical hardware limits. While optimized using insights from OpenAI’s own models, it is built with flexibility across modern LLM inference workloads, not locked exclusively to OpenAI’s models. This suggests potential future use beyond internal deployment.
The entire development cycle, from initial design to tape-out, took just nine months. For context, complex high-performance ASICs typically require two to four years. OpenAI leveraged its own AI models to accelerate portions of the design process, creating a notable feedback loop between software and hardware development.
The Custom Silicon Landscape
- Google is on track for millions of Ironwood TPU shipments in 2026.
- Amazon’s Trainium program has reached substantial internal scale and is growing rapidly.
- Microsoft’s Maia accelerators are already deployed in Azure.
- Meta’s MTIA chips power internal recommendation and inference workloads.
ASIC shipments are projected to grow significantly faster than general-purpose GPU shipments in 2026. The common thread is a willingness to trade generality for efficiency. These chips cannot run arbitrary workloads like video games or physics simulations, but for high-volume LLM inference they can deliver better performance per watt, lower power draw, and reduced cost at scale.
Implications for Nvidia
The picture for Nvidia is nuanced. The company remains the dominant supplier of AI infrastructure, with unmatched software (CUDA), developer ecosystem, and multi-year relationships. Hyperscalers and OpenAI continue to purchase large volumes of Nvidia GPUs for training and general-purpose workloads even as they develop custom chips for specific inference use cases.
However, the long-term pressure is real. As inference increasingly dominates total AI compute spend, because it runs continuously rather than periodically, custom silicon can capture a meaningful share of that demand. Every token served on a Jalapeño or Trainium-based system is one that does not run on an H100 or B200.
Two important risks remain. First, ASICs are inherently less flexible than GPUs. If model architectures evolve significantly (as they have every 18–24 months historically), purpose-built inference chips risk becoming stranded assets. Nvidia’s GPUs can be reprogrammed for new paradigms. OpenAI’s claim that Jalapeño supports a range of LLM workloads is a hedge against this risk, though it remains to be proven at full production scale.
Second, advanced chip production depends on constrained TSMC capacity at leading nodes, subject to export controls and geopolitical factors. OpenAI’s ambitious gigawatt-scale plans will require substantial wafer supply. While Nvidia faces similar constraints, its established customer relationships and long-term agreements give it priority access that a newer entrant must still secure.
Jalapeño represents a credible and impressive first step rather than an immediate disruption. OpenAI has shown it can deliver a production-viable inference ASIC in record time with strong partners. The harder challenges, scaling deployment, iterating through future chip generations, and building the software and operational stack needed for broad adoption, now begin.
The AI hardware landscape has a new serious player. How quickly and how far the shift toward custom silicon progresses will depend on execution at this new scale.
Follow @UTXOMacro for daily Bitcoin, macro, and AI breakdowns.