Home Companies Groq
Inference cloud Private Architecture licensed to Nvidia

Groq

Groq built the most commercially visible non-GPU inference architecture of the last decade, then licensed it to Nvidia in a transaction reported at approximately $20 billion. What remains is an inference cloud operator, and its story is now a case study in how hard it is to sustain a merchant alternative to the GPU.

Key facts

Founded
2016
Headquarters
Mountain View, California
Founder
Jonathan Ross — departed to Nvidia, December 2025
Original architecture
LPU — Language Processing Unit
Current business
GroqCloud inference platform and neocloud capacity
Most recent valuation
$3.5B, August 2026 (from $6.9B, September 2025)
Stated scale
13 data centres; targeting 200+ MW by end of 2027
Last reviewed
18 August 2026

Overview

Groq was founded in 2016 by Jonathan Ross, who had previously worked on Google's original Tensor Processing Unit effort. The company's thesis was that large language model inference — specifically the sequential decode stage — was a fundamentally different problem from training, and that a chip designed around determinism and on-chip memory could serve it far faster than a GPU.

For several years that thesis produced the most striking demonstrations in the market. GroqCloud served open-weight models at token rates that GPU-based providers could not match, and it did so through a conventional developer API, which made the advantage legible to people who had no interest in chip architecture.

In December 2025 Nvidia agreed to license Groq's inference technology on a non-exclusive basis in a deal reported at roughly $20 billion in cash — Nvidia's largest transaction on record. Ross, president Sunny Madra and other senior leaders joined Nvidia. Groq continued as an independent company, and GroqCloud continued to operate, but the company that exists today is a materially different business from the one that raised money as a chip designer.

The LPU architecture

The Language Processing Unit is worth understanding on its own terms, because its design choices explain both the performance results and the commercial difficulty.

Design choiceWhat it meansConsequence
SRAM instead of HBM Model weights are held in fast on-chip memory rather than off-chip high-bandwidth memory. Removes the memory-bandwidth wall that dominates GPU decode — but on-chip memory is small, so a large model must be spread across many chips.
Deterministic execution The compiler schedules every operation and data movement ahead of time; there is no dynamic scheduling or cache hierarchy to introduce variance. Latency is highly predictable, which is valuable for interactive and agentic workloads. It also means the compiler carries enormous responsibility.
Software-scheduled interconnect Chip-to-chip communication is planned at compile time rather than negotiated at runtime. Very low and predictable communication overhead across a large array of chips.
Inference only No attempt to serve training workloads. A cleaner design, but it forgoes the training revenue that funds competitors' inference roadmaps.

The economic difficulty follows directly from the first row. Because SRAM capacity per chip is small relative to HBM, serving a large model requires a substantial number of LPUs working as one unit. That produces excellent latency, but it means the capital cost of a deployment scales with model size in a way that GPU deployments do not. As frontier models grew and reasoning workloads pushed context lengths and generated output lengths up, that arithmetic became progressively harder.

The Nvidia transaction

On 24 December 2025 it was reported that Nvidia would pay approximately $20 billion in cash for rights to Groq's inference technology under a non-exclusive licensing agreement. Groq's own announcement framed it as a licensing arrangement intended to accelerate AI inference at global scale. Jonathan Ross, Sunny Madra and other senior personnel joined Nvidia to advance the licensed technology.

The structure attracted immediate comment. Because it was framed as a licence plus a hiring arrangement rather than an acquisition, Groq continued to exist as a nominally independent competitor. At least one analyst quoted by CNBC described the structure as keeping the "fiction of competition alive," and the transaction drew attention from US lawmakers. Whether licence-and-hire structures of this kind receive sustained antitrust scrutiny is an open question with implications well beyond Groq.

The effect on Groq's standalone valuation was significant. The company had raised $750 million in September 2025 at approximately $6.9 billion post-money. By August 2026 its valuation was reported at $3.5 billion — roughly half the peak — even though the licensing payment itself was a substantial cash inflow. Investors repriced the business once its founding technical leadership had moved to the dominant chipmaker.

What Groq is now

Groq today is best classified as an inference cloud and neocloud operator rather than a chip designer. Its June 2026 funding announcement described a business serving more than five million developers and thousands of AI-native companies, processing trillions of tokens per week across 13 data centres in North America, Europe, the Middle East and APAC, with a plan to scale toward 200 megawatts of capacity by the end of 2027.

Reporting on the August 2026 round describes a further shift: Groq operating medium and large clusters of Nvidia accelerated computing for training and inference, alongside its existing platform. On that account the company has moved from selling an alternative to the GPU to selling capacity built on it — a considerable change in what the business is, and one that places Groq in a crowded neocloud market rather than a differentiated silicon one.

This shift is the single most important thing to understand about Groq's current position, and it is why this profile is filed under inference cloud rather than inference silicon. Readers encountering older coverage that describes Groq primarily as an LPU chip company should treat that framing as out of date.

Strengths and key questions

Strengths

What Groq retains

  • Developer base. Five to six million developers and an OpenAI-compatible API represent genuine distribution that is expensive to rebuild.
  • Operating footprint. Thirteen data centres across four regions, with a stated path from roughly 54 MW toward 200+ MW.
  • Capital. The licensing proceeds plus $650 million in June 2026 and $350 million in August 2026 leave the company well funded relative to its new market position.
  • Latency reputation. The brand is strongly associated with fast inference, which remains commercially useful.
  • Nvidia relationship. Nvidia's planned participation in the August round suggests supply access on favourable terms.
Key questions

What is unresolved

  • Differentiation. As a neocloud running Nvidia hardware, what distinguishes Groq from the many other operators doing the same thing?
  • The LPU's future. The licence was non-exclusive, so Groq retains rights. Whether it continues to develop and deploy its own silicon is not clearly established.
  • Leadership continuity. The chief executive has changed more than once since December 2025. Public sources are not consistent on who currently holds the role.
  • Margin structure. Neocloud economics are thinner and more capital-intensive than proprietary-silicon economics.
  • Valuation trajectory. A halved valuation with a large cash inflow is an unusual combination and reflects real uncertainty about the forward business.

Leadership

Jonathan Ross founded Groq in 2016 and led it until December 2025, when he joined Nvidia as part of the licensing arrangement, along with president Sunny Madra. Ross's earlier work on Google's first-generation TPU is the direct intellectual lineage of the LPU.

Succession has been less clearly documented. CNBC reported at the time of the Nvidia deal that finance chief Simon Edwards would become chief executive of the continuing company. Groq's own June 2026 funding announcement identifies Adam Winter as chief executive. Subsequent reporting on the August 2026 round has associated the company's direction closely with lead investor Disruptive. This publication has not been able to reconcile these accounts from primary sources, and the position is recorded here as unresolved rather than stated with false confidence.

Funding history

DateEventTerms
August 2024Series D$640M at a $2.8B valuation
September 2025Growth round$750M at approximately $6.9B post-money
December 2025Nvidia licensing agreementReported at approximately $20B in cash; non-exclusive; founding leadership joins Nvidia
June 2026Growth capital$650M led by Disruptive and Infinitum; valuation not disclosed
August 2026Series A (recapitalisation)$350M led by Disruptive with planned Nvidia participation, at $3.5B

The August 2026 round is labelled a Series A despite following a Series D. Reporting indicates this reflects a deliberate reset establishing a new valuation for the post-licensing entity, rather than a conventional early-stage round. The relationship between the June and August raises is not fully explained in public sources.

Why this profile matters beyond Groq

Groq is the clearest test case so far for the central question in this market: can a merchant inference-silicon company sustain itself against a vertically integrated incumbent? Groq had the strongest possible version of the argument — a real architectural advantage, demonstrable performance, a developer following and substantial capital — and the outcome was a licensing transaction that transferred the technology and the team to the incumbent.

That does not prove the thesis wrong. It does establish that the exit path for a successful inference architecture may run through Nvidia rather than around it, which is a material consideration for anyone assessing Etched, d-Matrix, or the other specialists tracked on this site.

Sources

  1. Groq — Non-exclusive inference technology licensing agreement with Nvidia. Company statement on the December 2025 transaction.
  2. CNBC — Nvidia buying Groq's assets for about $20 billion (24 December 2025). Deal value, leadership moves and CEO succession.
  3. CNBC — Analyst commentary on the deal structure (26 December 2025). Critique of the licence-and-hire structure.
  4. Groq — $650M raise to scale its AI inference cloud (June 2026). Developer count, token volume, data centre footprint, capacity targets and named chief executive.
  5. TechCrunch — Groq raises $350M to fuel its pivot from AI chips to neocloud (17 August 2026). Valuation, round framing and strategy shift.
  6. Groq — $640M Series D announcement (August 2024). Historical funding reference.
  7. Data Center Dynamics — Nvidia to license technology from Groq and hire its leadership. Independent account of the arrangement.

Where public sources conflict — particularly on the identity of the current chief executive and on the relationship between the June and August 2026 raises — this profile records the conflict rather than selecting one account. Corrections with primary-source citations are welcome via the contact page.