Chips

Groq vs Cerebras: Speed, Price, Status (2026)

Independently measured, Cerebras runs gpt-oss-120b at 1,737 tokens per second against Groq's 477 — about 3.6x faster, at higher input cost. Nvidia did not acquire Groq; it licensed the technology for a reported $20B and the DOJ is investigating. Cerebras IPO'd in May 2026.

Cerebras Wafer-Scale Engine 3 Turbo processor, a single wafer-sized AI chip
Image: Cerebras

TL;DR

  • Cerebras is about 3.6x faster than Groq on the one model both serve at scale: 1,737 against 477 output tokens per second on gpt-oss-120b, independently measured by Artificial Analysis.
  • Groq is cheaper on that model — $0.15/$0.60 per million input/output tokens against Cerebras at $0.35/$0.75.
  • Nvidia did not acquire Groq. It paid a reported $20 billion in December 2025 for a non-exclusive licence to Groq's inference technology, plus most of its senior team. The DOJ opened a formal antitrust probe in September 2026. Unresolved.
  • Cerebras is publicly traded. The IPO priced on 13 May 2026 at $185 a share, raising $5.55B; it trades on Nasdaq as CBRS.

Groq vs Cerebras comes down to one measured gap. Cerebras delivers 1,737 tokens per second on gpt-oss-120b against Groq's 477 — roughly 3.6x faster — while Groq charges less than half as much per million input tokens. Choose Cerebras for latency-critical decode. Choose Groq for cost-sensitive volume and audio. Neither is a pure silicon startup anymore.

Did Nvidia acquire Groq?

No. Nvidia bought a licence and a team, not the company.

In December 2025, Groq announced a non-exclusive inference technology licensing agreement with Nvidia. Founder Jonathan Ross and president Sunny Madra left for Nvidia along with much of the engineering organisation. Groq stated it would continue to operate as an independent company under new CEO Simon Edwards, with GroqCloud running without interruption.

Neither party disclosed the price. Reporting puts it at $20 billion, though some outlets have cited $17 billion. Senators Elizabeth Warren and Richard Blumenthal, writing to Nvidia in March 2026, argued the company had "effectively acquired Groq in all but name."

The DOJ is now investigating. In September 2026 the Justice Department opened a formal antitrust probe into whether the licence-plus-hiring structure was built to avoid merger review — specifically whether it sidestepped Hart-Scott-Rodino filing thresholds that a straight acquisition would have triggered. No resolution has been announced. This is the first real legal test of the reverse-acquihire pattern, and the outcome is genuinely unknown. We covered the contract mechanics in our read of the DOJ probe.

Meanwhile Groq refounded itself as a cloud. In August 2026 it closed a $350M Series A led by Disruptive — with planned participation from Nvidia — valuing it at $3.5B, following $650M in June for $1B of recent funding. It now runs 13 data centers, serves more than six million developers, and is a certified Nvidia Cloud Partner.

Read that carefully. Groq today is an Nvidia-aligned neocloud. That is a different company from the LPU challenger the older comparison articles still describe.

Groq vs Cerebras: specs, speed and pricing

Groq (GroqCloud) Cerebras
Measured output speed, gpt-oss-120b 477 tok/s 1,737 tok/s
Vendor-claimed speed, same model ~500 tok/s ~3,000 tok/s
Measured time to first token 0.67s 0.47s
Measured end-to-end response 5.92s 1.91s
gpt-oss-120b price, in/out per 1M $0.15 / $0.60 $0.35 / $0.75
Text model families on public API 6+ 2
Speech-to-text / text-to-speech Yes No
Current hardware LPU; lineage now Nvidia Groq 3 LPX CS-4 / WSE-3 Turbo
Training workloads Via Nvidia GPU clusters Yes, on-premises systems
Corporate status Private, $3.5B valuation Public — Nasdaq: CBRS

Speeds measured by Artificial Analysis. Prices from Groq's model documentation and Cerebras pricing.

Which is faster, Groq or Cerebras?

Cerebras, by about 3.6x on output tokens and 3.1x end-to-end. On gpt-oss-120b — the only model both serve where a clean comparison exists — Artificial Analysis measures Cerebras at 1,737 tok/s and Groq at 477 tok/s. Full response time is 1.91s against 5.92s.

Now check those against the vendors' own claims.

Groq advertises roughly 500 tok/s for gpt-oss-120b. Measured: 477. Within 5%.

Cerebras advertises roughly 3,000 tok/s for the same model. Measured: 1,737. About 42% below the claim.

Cerebras still wins decisively. It just wins by less than its own marketing says. Anyone quoting the 3,000 figure — including the vendor comparison page ranking near the top for this query — is quoting a number no independent benchmark reproduces.

Cerebras chart comparing CS-4 inference speed against production GPU systems across a model set
Chart: Cerebras — vendor-published, sourced by Cerebras to "Artificial analysis and internal benchmarking (August 2026)".

LPU vs wafer-scale engine: the architectural difference

Both bet on SRAM instead of HBM, and both have now converged on the same deployment model.

Cerebras builds one chip the size of a wafer. CS-4, announced in August 2026, packs three WSE-3 Turbo processors per system. Cerebras claims tokens generated up to 30x faster than production GPU systems, up to 10x more throughput per watt than CS-3, wafer-to-wafer interconnect latency as low as 2 microseconds, and 1,000+ tok/s on models of 10 trillion parameters and beyond. Those are vendor figures, and Cerebras sources its own benchmark chart to a mix of Artificial Analysis data and internal testing.

Groq's LPU took the opposite route: many small deterministic chips, scheduled by the compiler rather than by hardware. That design now ships inside Nvidia as Nvidia Groq 3 LPX, a rack-scale inference accelerator for the Vera Rubin platform, launched at GTC in March 2026.

The convergence is the interesting part. Both camps now advocate disaggregated inference — run prefill on a GPU or ASIC, run latency-critical decode on the specialised engine. Neither argues its architecture should run the whole pipeline alone. That is a meaningful retreat from the 2024-era pitch, and it is why the case for custom silicon now rests on economics rather than raw capability.

Is Cerebras publicly traded?

Yes. Cerebras priced its IPO on 13 May 2026: 30 million Class A shares at $185.00, raising $5.55B. Trading opened on the Nasdaq Global Select Market on 14 May 2026 under CBRS and closed the first day up 68% at $311.07, the largest US technology IPO of the year at that point.

That means the financials are public, and they carry a warning. Per its Q2 2026 quarterly filing, revenue grew from $290.3M in 2024 to $510.0M in 2025, reaching $373.5M in the first half of 2026 — up 84% year over year. Remaining performance obligations stand at $25.4 billion, driven largely by an OpenAI agreement.

The concentration risk is real. MBZUAI accounted for 34% of Q2 2026 revenue and 49% of the first half; G42 added 9% and 10%. Two customers represented 76% of accounts receivable. Gross margin compressed as lower-margin cloud services grew as a share of revenue.

Groq vs Cerebras vs SambaNova

SambaNova sits between them. On gpt-oss-120b, Artificial Analysis measures SambaNova at 703 tok/s — faster than Groq, slower than Cerebras — at $0.14 input / $0.95 output per million tokens.

So the ordering on that model is straightforward: Cerebras for speed, Groq for cheapest output, SambaNova for cheapest input. All three serve open-weight models only. None serves the frontier proprietary models.

What this means for you

If you're building a real-time voice agent or interactive coding assistant: Cerebras. The 0.47s time to first token and 1.91s full response are the numbers that decide whether an interaction feels live. Accept the input-token premium and the two-model catalogue.

If you're batch-processing documents overnight: Groq, or neither. Speed you cannot perceive is not worth paying for. Price Cerebras' batch option against a conventional GPU provider before committing, and consider whether an open-source serving stack on rented GPUs is cheaper still.

If you need speech: Groq, by default. Its speech-to-text and text-to-speech models have no Cerebras equivalent — Cerebras handles text and image input only.

If procurement cares about vendor risk: neither is comfortable, for opposite reasons. Groq's technology roadmap now lives inside Nvidia, and the DOJ probe could unwind assumptions in either direction. Cerebras is public and audited but leans on a small number of customers and one enormous contract.

Practical hedge: dual-source behind an OpenAI-compatible gateway. Both expose the same API surface, so switching cost is low today — keep it that way. If you are sizing your own GPUs instead, start with H200 against B200 economics.

Frequently Asked Questions

Did Nvidia acquire Groq?

No. In December 2025 Nvidia paid a reported $20 billion for a non-exclusive licence to Groq's inference technology and hired founder Jonathan Ross plus much of the senior team. Groq remains a separate company under a new CEO and raised $350M in August 2026. The DOJ opened an antitrust probe into the structure in September 2026.

Which is faster, Groq or Cerebras?

Cerebras. On gpt-oss-120b, Artificial Analysis measures Cerebras at 1,737 output tokens per second against Groq's 477 — roughly 3.6x. End-to-end response time is 1.91s against 5.92s. Note that Cerebras' own marketing claims around 3,000 tok/s, about 42% above the independently measured figure.

If Cerebras is faster, why isn't everyone switching?

Three reasons. Cerebras serves only two model families on its public API against six or more on Groq, with no speech models at all. It costs more per million input tokens. And most workloads — batch jobs, background summarisation, embeddings — gain nothing a user can perceive from 1,700 tokens per second over 477.

How do Groq and Cerebras compare to Nvidia GPUs?

Both trade flexibility for decode latency. Cerebras claims up to 30x faster inference than production GPU systems; Nvidia makes comparable per-megawatt throughput claims for its Groq-derived LPX accelerator. But GPUs still handle training and prefill better, which is why both camps now recommend disaggregated inference — prefill on GPU, decode on the specialised engine.

Is Cerebras publicly traded?

Yes. Cerebras listed on Nasdaq as CBRS on 14 May 2026 after pricing its IPO at $185 a share and raising $5.55B. Shares closed the first session up 68% at $311.07. Quarterly filings are public: first-half 2026 revenue was $373.5M, up 84%, with remaining performance obligations of $25.4B.


Editor's note — sources: Additional references consulted but not linked above: Cerebras' CS-4 announcement (linked in the chart credit) for the CS-4 and WSE-3 Turbo specifications; Nvidia's developer blog on the Groq 3 LPX accelerator; Groq's August 2026 Series A release for funding, data-center and developer figures; Cerebras' Q2 2026 Form 10-Q filed with the SEC for revenue, customer concentration and remaining performance obligations; and the Warren-Blumenthal letter to Nvidia for the $20 billion figure and the "in all but name" characterisation. Vendor-claimed throughput is labeled as such throughout and separated from the Artificial Analysis measurements.

Get Edgewisely in your inbox

Business stories that matter, free. Enter your email — no password, no account to set up.
jamie@example.com
Subscribe