Home Companies OpenAI
Captive silicon Private Deploying late 2026

OpenAI (Jalapeño)

OpenAI designed its own inference chip because it is one of the few organisations whose workload is large enough, and well enough understood, to justify it. Jalapeño is not for sale — it exists to serve OpenAI's own traffic, which is exactly what makes it significant for everyone selling inference silicon.

Key facts

Chip
Jalapeño
Announced
24 June 2026
Design partner
Broadcom
Foundry
TSMC, 3nm
Memory
Eight HBM stacks, reticle-limited package
Networking
Broadcom Tomahawk Ultra Ethernet
Systems integrator
Celestica
Sold externally
No — serves OpenAI's own traffic
Deployment
Gigawatt-scale sites with Microsoft and partners, from late 2026
Last reviewed
21 August 2026

Overview

OpenAI is not a chip company, and Jalapeño is not a product. That is the most important thing to understand about it. The chip is captive silicon: there is no SKU, no cloud instance type, no rental market and no API that exposes it directly. It exists to serve OpenAI's own inference traffic more cheaply than buying equivalent capacity from Nvidia.

Announced on 24 June 2026 alongside Broadcom, Jalapeño is the clearest statement yet that the largest consumers of inference compute would rather own the economics than rent them. OpenAI joins Google, Amazon, Microsoft and Meta in designing its own inference accelerator — but unlike those four it does not operate a public cloud, so it has no route to monetise the chip beyond its own products.

The development timeline drew most of the attention: OpenAI and Broadcom took the design from start to tape-out in roughly nine months, an unusually short cycle for an ASIC of this complexity. OpenAI has said its own models were used to accelerate parts of the design process.

The chip

AttributeDetailWhy it matters
ArchitectureSystolic array paired with eight HBM stacks in a single reticle-limited packageA large compute die with substantial high-bandwidth memory, rather than the SRAM-heavy or LPDDR-based approaches other inference designs take.
Memory choiceHBM, not commodity DRAMA deliberate divergence from cost-optimised inference accelerators. It signals OpenAI is targeting throughput and latency together rather than lowest unit cost.
ProcessTSMC 3nmThe same generation as Nvidia's Blackwell. OpenAI is not economising on the node.
NetworkingBroadcom Tomahawk Ultra Ethernet siliconEthernet-based scale-out rather than a proprietary fabric — consistent with Broadcom's wider position against NVLink and InfiniBand.
Division of labourOpenAI designed the core architecture; Broadcom did silicon engineering and implementation; TSMC fabricates; Celestica builds racks and boardsOpenAI owns the part that encodes its knowledge of how its own models behave, and outsources everything else.
DeploymentGigawatt-scale data centres with Microsoft and other partners, from late 2026OpenAI does not operate its own data centres at this scale, so the chip reaches production through partners.

Transistor count, die size, HBM generation and power figures have not been publicly disclosed. Performance claims relative to GPU alternatives have not been independently verified, and because the chip is not sold, third parties cannot benchmark it.

Role in inference

Jalapeño is inference-only. It does not target training, which OpenAI continues to run on purchased GPU capacity. The rationale is straightforward arithmetic: a model is trained a finite number of times but served continuously, so at OpenAI's traffic volumes even a modest improvement in cost per token compounds into a very large number.

What makes the design interesting is that OpenAI had information no merchant vendor has — precise knowledge of how its own models behave in production, which operations dominate, where data movement is wasteful, and what the real balance between compute and memory should be. The company has described designing the architecture around those observed bottlenecks rather than around a general notion of what an AI accelerator should do. That is the strongest available argument for captive silicon, and it is unavailable to anyone selling into a broad market.

The choice of HBM over cheaper memory is the most revealing decision. Many inference ASICs economise on memory to reach a lower price point. OpenAI, serving latency-sensitive interactive products at enormous scale, evidently concluded that memory bandwidth was not the place to save money.

Strengths and key questions

Strengths

What the approach has going for it

  • Perfect workload knowledge. OpenAI is designing for models it wrote, serving traffic it can measure exactly.
  • Guaranteed demand. There is no customer acquisition problem. Every chip has a workload waiting for it.
  • Execution speed. Nine months to tape-out suggests the Broadcom partnership is working well.
  • Negotiating leverage. Even partial success improves OpenAI's position with every merchant supplier it buys from.
  • Lower specialisation risk. Betting on a narrow workload is far safer when you also control the models the chip must run.
Key questions

What is unresolved

  • No external verification. Because it is not sold, no independent party can benchmark it or test the efficiency claims.
  • Volume. Custom silicon economics depend on amortising design cost over enough units. A single captive generation is a narrower base than a merchant product line.
  • Architecture drift. Nine months is fast, but model architectures can move faster. The chip must stay well-matched to models OpenAI has not built yet.
  • Partner dependency. Deployment runs through Microsoft and others, so OpenAI controls the design but not the data centre.
  • Displacement, not replacement. Nothing published suggests Jalapeño ends OpenAI's GPU purchasing. The realistic question is what share of inference it takes.

Why this matters for the wider market

The most consequential effect of Jalapeño may not be on OpenAI's costs at all. Every merchant inference vendor — Etched, d-Matrix, OLIX and the rest — is selling into a market whose largest potential customers are progressively deciding to build rather than buy. OpenAI was, in principle, the single most attractive logo any of them could have won.

It also sharpens the argument about Nvidia. The threat has never been mainly that a startup builds a better chip; it is that Nvidia's largest customers each build an adequate one. Google, Amazon, Microsoft, Meta and now OpenAI have all done so. What none of them has done is displace Nvidia for the workloads where flexibility and day-one model support matter, which remains the larger part of the market.

Sources

  1. OpenAI — OpenAI and Broadcom unveil LLM-optimized inference chip. Primary announcement.
  2. Broadcom Investor Relations — OpenAI and Broadcom unveil LLM-optimized intelligence processor. Counterparty statement.
  3. CNBC — OpenAI and Broadcom reveal Jalapeño (24 June 2026). Announcement date and deployment plans.
  4. Tom's Hardware — Reticle-sized ASIC built in a nine-month cycle. Package and design-cycle detail.
  5. VentureBeat — Development accelerated with OpenAI's own models. Architecture rationale and design process.

Because Jalapeño is captive silicon, almost all available information originates with OpenAI or Broadcom. This profile separates disclosed technical facts from company positioning and marks undisclosed specifications as such. See the editorial methodology.