Home Companies d-Matrix
Digital in-memory compute Private Full production since June 2026

d-Matrix

d-Matrix attacks the problem most inference architectures agree on — that generating tokens is limited by moving data, not by arithmetic — by putting the computation inside the memory. It is one of the few specialists to have moved from announcement to volume production.

Key facts

Founded
2019
Headquarters
Santa Clara, California
Co-founder and chief executive
Sid Sheth
Co-founder and chief technology officer
Sudeep Bhoja
Inference architecture
Digital in-memory compute (DIMC), chiplet-based
Products
Corsair accelerator, JetStream networking, Aviator software
Total disclosed funding
$450M
Last disclosed valuation
$2B, November 2025
Last reviewed
18 August 2026

Overview

d-Matrix was founded in 2019 by Sid Sheth and Sudeep Bhoja, who had previously worked together in high-speed connectivity silicon at Inphi and Broadcom. That background is relevant: the company's founding insight was a data-movement problem, and its founders came from the part of the industry that spends its time moving data efficiently.

The company committed to inference specifically, and early — before the current market existed in its present form. Sheth has framed the origin as a prediction that once trained models had to run continuously at scale, the available infrastructure would not be the right shape for the job. That timing looks prescient now; it required several years of patience at the time.

d-Matrix is at a stage that distinguishes it from most of its peer group. In June 2026 it announced that the Corsair platform had entered full production and was shipping in volume to priority customers. In a field where a great deal of coverage concerns pre-production silicon and forward-looking contracts, shipping product in volume is a meaningful differentiator.

The architecture: digital in-memory compute

In a conventional accelerator, weights live in memory and are fetched to compute units for every operation. During decode, when a model generates one token at a time, that fetch dominates: the arithmetic per weight is tiny, so the processor spends most of its time waiting on memory and most of its energy on moving bytes rather than computing with them.

In-memory compute inverts this by performing the multiply-accumulate operations within or immediately adjacent to the memory arrays holding the weights. The data does not travel; the computation goes to it.

The word digital in d-Matrix's approach is the important qualifier. Analog in-memory compute has been researched for years and offers extraordinary theoretical efficiency, but it introduces noise, drift, calibration burden and precision limits that make production deployment difficult. d-Matrix implements the same locality principle in the digital domain, accepting a lower theoretical ceiling in exchange for deterministic, reproducible numerical behaviour — the kind of engineering trade-off that separates research demonstrations from shipping products.

Chip and product families

ProductWhat it isRole
Corsair Chiplet-based PCIe inference accelerator implementing digital in-memory compute. The compute product. In full production and shipping in volume since June 2026.
JetStream Networking accelerator for scale-out between Corsair nodes. Because a single card cannot hold a large model, the interconnect determines whether the in-memory advantage survives at rack scale.
Aviator Software stack for orchestration, model compilation and serving. The adoption bottleneck for every non-GPU architecture, and the part buyers should scrutinise most closely.

Company performance claims

d-Matrix publishes the following comparisons against GPU-based systems. These are vendor claims and have not been independently verified by this publication:

  • Up to 10x faster performance than GPU-based systems
  • 3x lower cost than GPU alternatives
  • 3–5x better energy efficiency
  • Up to 30,000 tokens per second at 2ms per token on a Llama 70B-class model
  • Support for models up to 100 billion parameters in a single rack

The 100-billion-parameter-per-rack figure is the most useful number in this list for assessing fit, because it defines the ceiling of what a single rack can serve. Buyers evaluating larger frontier models — particularly large mixture-of-experts architectures — should establish how many racks a given deployment would require and what that does to the cost comparison.

Role in inference

d-Matrix targets a specific and commercially large segment: high-volume serving of models in roughly the 7-to-100-billion-parameter range, where latency matters and the workload is stable enough to justify specialised infrastructure. This covers a great deal of production AI — enterprise assistants, retrieval-augmented generation, classification and extraction pipelines, coding tools and the increasingly important category of agentic systems that issue many sequential model calls.

The company has explicitly connected demand to the growth of agentic tooling, noting that these workloads push inference volumes beyond what GPU-only infrastructure was designed to handle. An agent that makes dozens of sequential model calls to complete one user request multiplies decode-stage load in exactly the way that favours a decode-optimised architecture.

What d-Matrix is not positioned for is training, or serving the largest frontier models within a single rack. It has chosen a segment rather than trying to displace the GPU everywhere — which is a more defensible strategy than it sounds, provided the segment holds.

Strengths and key questions

Strengths

What is working

  • Shipping product. Full production since June 2026 puts d-Matrix ahead of most specialists, which remain pre-production.
  • Right bottleneck. Memory movement is the correct thing to optimise for decode, and this is now broadly agreed across the industry.
  • Digital, not analog. The engineering choice that makes in-memory compute deployable rather than experimental.
  • Complete stack. Compute, networking and software sold together, which reduces integration risk for buyers.
  • Sovereign-friendly investor base. Temasek, QIA and EDBI participation aligns with national AI infrastructure programmes.
Key questions

What has yet to be proven

  • Named customers. "Priority customers" is not a disclosed reference list. Public, named production deployments are the missing evidence.
  • Model size ceiling. If production workloads shift decisively to very large mixture-of-experts models, the single-rack limit becomes a harder constraint.
  • Independent benchmarks. All published performance figures are the company's own, measured on its own configurations.
  • Software depth. Aviator must keep pace with new model architectures on release, not months later.
  • Capital scale. $450 million raised is substantial but an order of magnitude below what several competitors have available.

Funding

d-Matrix is privately held. Revenue is not publicly disclosed.

DateRoundTerms and investors
November 2025 Series C $275M at a $2B valuation, bringing total raised to $450M. Led by BullhoundCapital, Triatomic Capital and Temasek, with participation from Qatar Investment Authority, EDBI, M12 (Microsoft's venture fund), Nautilus Venture Partners, Industry Ventures and Mirae Asset.
Prior rounds Seed through Series B Approximately $175M cumulative, implied by the $450M total. Microsoft's M12 and Temasek were among earlier backers.
Revenue Undisclosed. As a private company d-Matrix does not report financial results.

The investor list is worth reading closely. M12's presence indicates Microsoft has been tracking the technology since well before the current round. Temasek, the Qatar Investment Authority and EDBI are sovereign or state-linked funds, consistent with the pattern of national investment vehicles taking positions in inference infrastructure.

Leadership

Sid Sheth is co-founder and chief executive. Sudeep Bhoja is co-founder and chief technology officer. Both came from senior roles in high-speed connectivity and networking silicon at Inphi and Broadcom, which is an unusual founding background for an AI accelerator company and shows in the product: d-Matrix treats interconnect as a first-class product rather than an afterthought, which is why JetStream exists alongside Corsair.

What to watch

  • Named production customers. The single most valuable disclosure d-Matrix could make, and the clearest signal that volume shipping is translating into durable revenue.
  • Independent benchmark results. Third-party measurement across models, context lengths and concurrency levels.
  • Next-generation roadmap. Whether the follow-on to Corsair raises the per-rack parameter ceiling, which would directly address the main architectural limitation.
  • Cloud availability. Whether Corsair capacity becomes purchasable through a cloud or neocloud provider rather than only through direct systems sales.
  • Sovereign deployments. Given the investor base, national AI infrastructure programmes are a plausible early market.

Sources

  1. d-Matrix — $275M Series C announcement (12 November 2025). Amount, valuation, total raised, full investor list, executive titles and performance claims.
  2. d-Matrix — Corsair enters full production (June 2026). Production status and volume shipping to priority customers.
  3. PR Newswire — d-Matrix Series C release. Wire distribution of the funding announcement.
  4. Data Center Dynamics — d-Matrix raises $275m at $2bn valuation. Independent reporting on the round.
  5. The AI Insider — d-Matrix funding coverage. Secondary reporting and context.

All performance figures on this page originate from d-Matrix and are labelled as vendor claims. No independent benchmark of Corsair is available to this publication at the time of review. See the editorial methodology for how vendor claims are handled.