Overview
Google has been building AI accelerators longer than anyone else in this market. The first Tensor Processing Unit was deployed internally in 2015, years before the current cycle began, and the programme has run continuously since. That accumulated experience — across silicon, compilers, interconnect, data-centre design and the models themselves — is the closest thing to a genuine peer of Nvidia's full-stack position.
The seventh generation, Ironwood, reached general availability for cloud customers at Google Cloud Next in April 2026. Google positioned it explicitly as “the first Google TPU for the age of inference”, a notable framing from a company whose earlier TPU generations were justified largely by training economics. Ironwood delivers 4.6 petaFLOPS per chip, and a 9,216-chip superpod reaches 42.5 exaFLOPS.
Two things have changed the strategic picture in 2026. Google has begun selling TPU systems externally rather than only renting capacity through Google Cloud, and it has committed to splitting future generations into separate training and inference designs.