Home Companies AWS
Cloud-captive silicon Part of Amazon — NASDAQ: AMZN Trainium3 rolling out

AWS (Inferentia and Trainium)

AWS runs one of the largest deployments of non-Nvidia AI silicon anywhere — close to a million Trainium2 chips serving a single customer's models. It sells none of it. The chips exist to lower the cost of AWS, which is both their strength and their limit.

Key facts

Design organisation
Annapurna Labs, acquired by Amazon in 2015
Current generation
Trainium3, announced at re:Invent 2025
Trainium3 performance
2.52 PFLOPS per chip, TSMC 3nm, ~40% faster than Trainium2
Deployed scale
Close to 1M Trainium2 chips serving Anthropic models
Project Rainier
~500,000 Trainium2 chips, 1,200-acre Indiana site, live since October 2025
Inference-specific line
Inferentia and Inferentia2, from 2018; now secondary to Trainium
Software
Neuron SDK, with native PyTorch support
Sold externally
No — available only as AWS instances
Programme revenue
Not disclosed
Last reviewed
21 August 2026

Overview

AWS designs its AI accelerators through Annapurna Labs, the Israeli chip company it acquired in 2015 and which also produced the Nitro system and Graviton processors. Annapurna is the least publicised and arguably most successful custom-silicon organisation in the cloud industry, and its AI work now runs at a scale that makes AWS one of the largest deployers of non-Nvidia accelerators in the world.

The portfolio has two names, and the relationship between them has quietly changed. Inferentia, launched in 2018, was AWS's inference accelerator; Trainium arrived later for training. In practice Trainium has become the strategic line and now serves both workloads, with Trainium3 announced at re:Invent 2025 delivering 2.52 petaFLOPS per chip on TSMC 3nm and roughly a 40% performance improvement over its predecessor. Readers encountering the Inferentia name should treat it as the older half of a portfolio whose centre of gravity has moved.

What distinguishes AWS from the other hyperscaler programmes is the sheer deployed volume. Anthropic models are served on close to one million Trainium2 chips, and Project Rainier — activated in October 2025 across a 1,200-acre site in Indiana — deploys nearly 500,000 Trainium2 chips dedicated to training Anthropic's Claude models.

Chip and product families

ProductWhat it isInference relevanceStatus
Trainium3Announced at re:Invent 2025. 2.52 PFLOPS per chip on TSMC 3nm, deployed in UltraServer configurationsRoughly 40% faster than Trainium2, developed with direct Anthropic input on training speed, latency and energy efficiency. Serves inference as well as training.Rolling out through 2026
Trainium2The volume generation, deployed at very large scaleClose to one million chips serving Anthropic models, plus general AWS customer inference.Deployed at scale
Inferentia / Inferentia2The original inference-specific line, from 2018Still available, but strategically superseded by Trainium as the primary accelerator for both workloads.Available; legacy emphasis
Neuron SDKThe software layer: compiler, runtime and framework integrationThe deciding factor in adoption. Recent releases added native PyTorch support, which had been the main friction for customers.Shipping

Project Rainier and the Anthropic relationship

Project Rainier is the clearest illustration of what hyperscaler silicon at scale actually looks like. Activated in October 2025, it occupies a 1,200-acre facility in Indiana and runs nearly 500,000 Trainium2 chips dedicated to Anthropic, providing more than five times the compute used to train earlier Claude models.

The relationship runs deeper than a supply agreement. Anthropic contributed direct input to the Trainium3 design on training speed, latency and energy efficiency — a co-design arrangement rather than a purchase. AWS gets a demanding frontier customer shaping its roadmap; Anthropic gets silicon tuned to its workloads and a supply position independent of GPU allocation.

It is worth noting that Anthropic runs the same strategy across multiple suppliers. It is the anchor customer for Google's TPU programme at up to a million chips, it has an agreement with Fractile, and it buys from Nvidia and AMD. Diversification across accelerator vendors is deliberate, and it means no single one of these relationships should be read as exclusive.

Role in inference

AWS's objective is not to sell chips but to lower the cost of serving AI on its own platform, and to reduce how much of its capital expenditure flows to Nvidia. Customers do not buy Trainium; they buy instances, and AWS captures the margin difference between what a Trainium instance costs to run and what a comparable GPU instance would.

For a customer, the calculation is straightforward and entirely about software. If a workload runs well through Neuron on a mainstream framework, Trainium instances are typically cheaper than equivalent GPU capacity. If it depends on custom CUDA kernels or an unusual model architecture, the porting cost may exceed the saving. Native PyTorch support in Neuron has narrowed that gap considerably, but it has not closed it for the hardest cases.

Reporting in March 2026 indicated that Trainium had attracted OpenAI and Apple alongside Anthropic. If accurate, that is significant: it suggests Trainium is winning frontier workloads on merit rather than only serving AWS-native customers taking the cheapest available option.

Strengths and key questions

Strengths

What is working

  • Deployed scale. Close to a million Trainium2 chips serving one customer's models is a volume no merchant specialist approaches.
  • Annapurna's track record. Nitro and Graviton both succeeded; this is an organisation that has shipped custom silicon repeatedly.
  • Frontier co-design. Anthropic shaping the Trainium3 specification is a stronger feedback loop than customer surveys.
  • Distribution. Every AWS customer can try Trainium by changing an instance type, with no procurement cycle.
  • Neuron maturity. Native PyTorch support removes the most-cited barrier to adoption.
Key questions

What is unresolved

  • Disclosure. AWS publishes no Trainium revenue, unit volumes or margin data. External assessment relies on announcements and reporting.
  • Customer concentration. Anthropic dominates the visible deployment. How much Trainium serves everyone else is unclear.
  • Cloud lock-in. Trainium exists only on AWS. Building on it is a commitment to one provider in a way buying GPUs is not.
  • Inferentia's future. AWS has not clearly articulated whether the inference-specific line continues or is absorbed into Trainium.
  • Specialisation pressure. With Google splitting training and inference silicon, a single line serving both may face the same efficiency argument AWS makes against GPUs.

What to watch

  • Trainium3 deployment scale. How quickly it reaches the volumes Trainium2 achieved, and whether a Rainier-equivalent is built around it.
  • Named non-Anthropic customers. Confirmation of OpenAI or Apple usage at scale would materially change the picture.
  • Trainium4. Whether AWS follows Google in splitting training and inference designs, or continues with a converged part.
  • Neuron time-to-support. How quickly new frontier model architectures run well on Trainium relative to GPU.
  • Instance pricing. The real measure of whether custom silicon is delivering cost advantage is what AWS charges for it.

Sources

  1. AWS — Trainium. Product documentation and positioning.
  2. AWS — AI chips at re:Invent 2025. Trainium3 announcement and Neuron SDK updates.
  3. Technology Magazine — How 500,000 Trainium2 chips power Project Rainier. Scale, location and Anthropic dedication.
  4. SemiAnalysis — AWS Trainium3 deep dive. Independent technical analysis and per-chip specifications.
  5. TechCrunch — Inside Amazon's Trainium lab (22 March 2026). Reported customer adoption beyond Anthropic.

AWS does not disclose Trainium or Inferentia revenue, volumes or margins, and the chips are not sold outside AWS, so no independent benchmarking of purchased hardware is possible. Reported customer names beyond Anthropic have not been confirmed by the companies involved. See the editorial methodology.