Home Companies
Company landscape

The builders of AI inference infrastructure.

A research view of the public GPU vendors, hyperscaler silicon teams and specialist companies shaping AI inference. Entries with a full profile carry their own sourced page; the rest summarise the company's role and the evidence currently available.

Public GPU and accelerator platforms

The merchant vendors selling general-purpose accelerators that any buyer can purchase and deploy.

GPU incumbent

Nvidia →

The default platform for AI inference at scale. The advantage is the combination of accelerator, interconnect, rack-scale systems and the CUDA software stack that most production workloads were written against.

Financial signal: Fiscal 2026 revenue of $215.9B, with data centre revenue of $197.3B.

Key questions: Rubin timing, hyperscaler mix shift to internal silicon, power availability, export controls, and regulatory treatment of the Groq licensing structure.

GPU challenger

AMD →

The only merchant vendor shipping general-purpose AI accelerators at a scale that makes it a credible second source. Competes on memory capacity, open software and supply diversity rather than architectural novelty.

Financial signal: First-half 2026 data centre revenue of $12.5B, up 81% year over year.

Key questions: Helios rack-scale execution, ROCm parity for new architectures, the OpenAI warrant, and HBM4 supply.

Enterprise alternative

Intel →

The only company here that both designs accelerators and owns leading-edge fabs. Its inference strategy now centres on Crescent Island, an inference-only GPU using 160GB of LPDDR5X rather than HBM — cheaper, air-coolable and outside the HBM queue.

Financial signal: Q2 2026 revenue of $16.1B, up 25%, with Data Center and AI up 59% to $6.3B.

Key questions: OEM availability, software traction, roadmap continuity, and customer proof beyond early adopters.

Edge leader

Qualcomm

Central to on-device inference across phones, PCs, vehicles and embedded devices, where power draw and privacy usually matter more than data-centre throughput.

Strength: Low-power deployment footprint across billions of endpoints.

Key questions: On-device model support, AI PC adoption, automotive inference and developer tooling. Full profile in preparation.

Hyperscaler custom silicon

Chips designed by cloud providers primarily to improve the economics of their own inference workloads. These are consumed through the provider's cloud rather than sold as components, which changes how buyers should evaluate them.

Captive silicon

OpenAI →

Jalapeño, announced with Broadcom in June 2026, is OpenAI's first custom inference chip: a reticle-sized ASIC on TSMC 3nm with eight HBM stacks, taped out in roughly nine months. It is not sold — it serves OpenAI's own traffic.

Product signal: Deploying at gigawatt scale with Microsoft and other partners from late 2026.

Key questions: No independent benchmarking is possible, and it is unclear what share of OpenAI inference it displaces from GPUs.

TPU platform

Google TPU →

The longest-running custom AI silicon programme, combining chips, compiler work, runtime software and Google Cloud distribution. Recent generations are positioned around large-scale inference and reasoning workloads.

Specification signal: Google documentation lists TPU7x with 192 GiB HBM per chip and up to 9,216 chips per pod.

Key questions: Availability, JAX and PyTorch support depth, and usage by customers outside Google. Full profile in preparation.

AWS silicon

Amazon Web Services →

Trainium and Inferentia exist to improve the cost structure of AI training and inference inside AWS, and to reduce dependence on a single external supplier.

Positioning: AWS describes Trainium as purpose-built for training and inference economics at scale.

Key questions: Neuron SDK maturity, instance availability and the scale of large customer workloads. Full profile in preparation.

Azure silicon

Microsoft Maia →

Azure Maia is Microsoft's custom accelerator family, aimed at Azure, Copilot and partner model-serving economics.

Product signal: Microsoft introduced Maia 200 as an inference accelerator with FP8 and FP4 tensor cores and HBM3e.

Key questions: SDK maturity, deployment regions, and how broadly Maia is exposed to customers rather than used internally. Full profile in preparation.

Internal inference

Meta

MTIA addresses very large internal inference workloads in ranking, recommendation, advertising and generative AI. Meta is a buyer of merchant silicon and a builder of its own simultaneously.

Constraint: Public technical and financial detail is thinner than for merchant vendors, which limits what can be responsibly stated.

Key questions: Deployment scale, workload mix, and whether Meta ever exposes silicon externally. Full profile in preparation.

Inference specialists

Companies whose entire proposition is that general-purpose GPUs are the wrong shape for production inference. This is where the architectural diversity — and the commercial risk — is concentrated.

Photonic computing

OLIX →

London-based, computing with light rather than electricity. Its optical tensor processing units pair photonics with SRAM and avoid HBM entirely; the first product is the DX-1 decode accelerator.

Funding: $312M Series B at a $3.3B valuation in August 2026 — reported as Europe's largest semiconductor round — with Arm and the UK Sovereign AI fund.

Key questions: Photonic manufacturability at yield, numerical precision, and an entirely new software stack.

Wafer scale

Cerebras →

Builds processors the size of an entire silicon wafer. The most architecturally radical company shipping inference at commercial scale, and since May 2026 the only inference specialist whose figures can be checked against filings.

Financial signal: Q2 2026 core revenue of $209.9M, up 103%; remaining performance obligations of $25.4B at 30 June 2026.

Key questions: Customer concentration, delivering 750 MW of contracted capacity, and the gap between GAAP and core revenue.

Inference cloud

Groq →

Built the most commercially visible non-GPU inference architecture of the last decade, then licensed it to Nvidia in a transaction reported at approximately $20 billion. Now operates as an inference cloud and neocloud rather than a chip designer.

Financial signal: Valued at $3.5B in August 2026, down from $6.9B in September 2025.

Key questions: Differentiation as a neocloud, the future of the LPU, and leadership continuity.

Digital in-memory

d-Matrix →

Places computation inside the memory arrays holding model weights, attacking data movement directly. One of the few specialists to have reached volume production.

Financial signal: $275M Series C at a $2B valuation in November 2025; $450M raised in total. Revenue undisclosed.

Key questions: Named production customers, the single-rack model size ceiling, and independent benchmarks.

Specialised ASIC

Etched →

The most aggressive specialisation bet in the market, now broadened from a transformer-only chip to a split prefill-compute and decode-memory architecture.

Financial signal: $300M Series C at a $21B valuation in July 2026; more than $1B in pre-booked orders. Revenue undisclosed.

Key questions: Production at scale, conversion of contracts to revenue, and independent verification of model-compatibility claims.

RDU systems

SambaNova →

Reconfigurable Dataflow Unit platforms sold as full-stack systems to enterprises and governments running frontier models on infrastructure they control. The fifth-generation SN50, aimed at agentic workloads, ships from H2 2026 with SoftBank as first customer.

Funding signal: $1B Series F first close at an $11B valuation in July 2026 — eight months after being reported to be exploring a sale, with Intel's ~$1.6B approach having lapsed.

Key questions: A $5.1B–$1.6B–$11B valuation round trip in five years, undisclosed revenue, and lumpy systems sales.

Sovereign infrastructure

Rebellions →

South Korea's national champion, formed by the December 2024 merger with SK Telecom's SAPEON. The Rebel100 pairs four UCIe chiplets on Samsung 4nm with 144GB of HBM3e; RebelRack and RebelPOD have shipped since March 2026.

Funding signal: $400M pre-IPO at a $2.34B valuation in March 2026, including ~$166M directly from the South Korean government. Cumulative funding ~$850M.

Key questions: Whether the sovereignty argument that wins at home travels, and what the IPO prospectus reveals about revenue.

UK inference startup

Fractile →

Founded in 2022 by Walter Goodwin, an Oxford roboticist. Performs the matrix multiplications that dominate transformer inference inside SRAM cells sitting alongside the compute logic — an in-memory approach that uses no HBM and avoids off-chip DRAM movement entirely.

Funding signal: $220M Series B in May 2026 at roughly $1B post-money, led by Accel, Factorial Funds and Founders Fund, with Conviction, Felicis, 8VC and existing backers Kindred Capital, the NATO Innovation Fund and Oxford Science Enterprises.

Reported, not confirmed: an initial agreement to supply Anthropic with roughly $250M of chips, and advanced talks to raise about $600M at a $6.5B pre-money valuation — reportedly co-led by Redpoint and Lightspeed. Both companies declined to comment, and the round has not closed.

Key questions: Silicon is not expected to deploy until 2027, so a six-fold reprice rests on a contract for hardware that does not yet exist. Full profile in preparation.

Large-model chips

MatX

Designs chips for frontier labs and very large model workloads, founded by former Google TPU contributors Mike Gunter and Reiner Pope.

Funding signal: Private. Investors listed by the company include Jane Street, Spark Capital and Situational Awareness LP.

Key questions: Public deployment evidence remains limited. Full profile in preparation.

AI + RISC-V

Tenstorrent

Open, programmable AI compute paired with a RISC-V CPU strategy and an IP licensing business model that differs from its peers.

Products: Wormhole, Blackhole and successor roadmaps.

Key questions: Converting developer interest into large production deployments. Full profile in preparation.

How to read this landscape

Evidence tiers

Not all figures are equal

Public companies file audited results. Private companies announce rounds and publish their own benchmarks. Where a figure is a vendor claim or a forward-looking commitment rather than delivered revenue, profiles on this site say so explicitly.

Undisclosed

Silence is recorded, not filled

Private-company revenue is marked undisclosed rather than estimated. Where public sources conflict — as they currently do on Groq's leadership — the conflict is recorded rather than resolved by preference.

Coverage

Profiles in preparation

Companies without a linked profile are tracked here and scheduled for full treatment. If you have primary-source material on any of them, the contact page is the fastest route.

Landscape last reviewed: 21 August 2026 Six full profiles published; further profiles in preparation Editorial methodology Submit a correction or update