Chips

H200 vs B200: Specs, Real Benchmarks and Cost

B200's real advantage over H200 is 1.6x to 2.3x, not the 4x NVIDIA's FP4 headline implies. Verified specs, MLPerf results at matched precision, and current rental prices.

NVIDIA DGX B200 system containing eight Blackwell B200 GPUs
Image: NVIDIA

TL;DR

  • B200 beats H200 by 2.27x on FP8 throughput and 1.67x on memory bandwidth — not the 4-5x implied by NVIDIA's FP4 headline numbers. Hopper has no FP4 hardware at all.
  • In MLPerf Inference v5.0, the only like-for-like test (Stable Diffusion XL, both at FP8) showed B200 at 1.60x H200 per GPU. The LLM tests showing 2.8-3.0x ran B200 at FP4 against H200 at FP8.
  • Rental, 24 September 2026: B200 ~$6.69-$8.60/GPU-hr, H200 ~$4.50-$6.31, H100 ~$3.49-$6.16. B200 carries roughly a 48% premium over H200 at Runpod list.
  • Verdict: take B200 for training, 100B+ models, and FP4-quantised inference. H200 is still the better dollar for FP8/BF16 inference on models under ~70B.

H200 vs B200: the short answer

The B200 beats the H200 on every axis, but by less than the marketing implies. At matched precision it delivers 2.27x the FP8 throughput and 1.67x the memory bandwidth, with 1.28x the capacity. NVIDIA's larger claims compare Blackwell at FP4 against Hopper at FP8 — a precision Hopper cannot run.

H200 vs B200: full specification comparison

Per GPU, SXM form factor. Tensor Core figures are dense unless marked; NVIDIA's H200 product page publishes Hopper figures only with sparsity, so Hopper dense values are derived by halving.

Spec H100 SXM H200 SXM B200
Architecture Hopper Hopper Blackwell
Transistors 80B 80B 208B
Die configuration 1 die 1 die 2 reticle-limited dies, 10 TB/s NV-HBI
Memory 80 GB HBM3 141 GB HBM3e 180 GB HBM3e
Memory bandwidth 3.35 TB/s 4.8 TB/s 8 TB/s
FP4 Tensor Core not supported not supported 9 PFLOPS (18 sparse)
FP8 Tensor Core 1,979 TF (3,958 sparse) 1,979 TF (3,958 sparse) 4,500 TF (9,000 sparse)
BF16/FP16 Tensor Core 990 TF (1,979 sparse) 990 TF (1,979 sparse) 2,200 TF (4,500 sparse)
FP64 34 TFLOPS 34 TFLOPS 37 TFLOPS
NVLink per GPU 900 GB/s 900 GB/s 1.8 TB/s
Max TDP 700 W 700 W 1,000 W

Two things stand out. The FP8 and BF16 ratio is identical at 2.27x — Blackwell's compute gain is uniform across precisions Hopper also supports. And memory bandwidth only improves 1.67x, well below the compute gain. That gap is the whole story.

Why the FP4 comparison is misleading

Hopper has no FP4 hardware. NVIDIA's fourth-generation Tensor Core supports FP8, FP16, BF16, TF32, FP64 and INT8. FP4 and FP6 arrive with Blackwell's fifth-generation Tensor Core and second-gen Transformer Engine, per NVIDIA's Blackwell architecture page.

So when a page quotes B200 at 18 PFLOPS FP4 against H200 at 3,958 TFLOPS FP8 and calls it 4.5x, it is comparing two different number formats. The honest matched-precision figure is 2.27x.

FP4 is a real advantage — halving weight bytes doubles effective parameter bandwidth and model capacity per GPU. But it is a quantisation gain, available only if your model tolerates 4-bit weights. It is not a like-for-like silicon comparison.

What MLPerf actually shows

MLPerf Inference v5.0 is the last round where NVIDIA submitted B200, H200 and H100 together. Per-GPU figures below are derived from MLCommons' published results by dividing throughput by accelerator count.

Benchmark Scenario B200 (per GPU) H200 (per GPU) Ratio Precision
Llama 2 70B Server 12,305 tok/s 4,134 tok/s 2.98x FP4 vs FP8
Llama 2 70B Offline 12,357 tok/s 4,374 tok/s 2.83x FP4 vs FP8
Llama 3.1 405B Offline 190.8 tok/s 70.1 tok/s 2.72x FP4 vs FP8
Mixtral 8x7B Offline 16,019 tok/s 7,829 tok/s 2.05x FP4 vs FP8
Stable Diffusion XL Offline 3.80 samples/s 2.37 samples/s 1.60x FP8 vs FP8

The last row is the control. Stable Diffusion XL is the only benchmark in the round where both systems ran the same precision — and the advantage collapses from ~2.8-3.0x to 1.60x, almost exactly what the matched-precision spec ratio and bandwidth ratio predict.

Note also: NVIDIA stopped submitting Hopper to MLPerf Inference from v5.1 onward. Every current B200-versus-Hopper comparison traces back to this single v5.0 round.

Why measured B200 gains fall short of the marketing multiple

Three reasons, all verifiable.

1. Decode is memory-bandwidth bound, not compute bound. LLM token generation at modest batch sizes is limited by how fast weights stream from HBM. B200 gives you 1.67x the bandwidth, not 2.27x the compute. If your workload is decode-heavy and small-batch, 1.67x is your realistic ceiling regardless of the FLOPS on the box.

2. The headline multiples are best-case, cherry-picked configurations. SemiAnalysis modelled NVIDIA's claimed 30x GB200-over-H200 figure and found that removing the FP4-vs-FP8 mismatch leaves "only an ~18x performance gain." Change the benchmark from 32K-input/1K-output to 512-input/2K-output and the gain drops to "less than 8x." They also note NVIDIA benchmarked the H200 at TP64 — the worst possible parallelism scheme for an 8-GPU NVLink domain.

3. Software maturity. Blackwell's advantage is unlocked by FP4 kernels that took over a year to mature. vLLM's GB200 optimisation work — NVFP4 GEMM via FlashInfer, FP8 GEMM for MLA, NVFP4 MoE dispatch, RoPE+quant fusion — reached 26.2K prefill and 10.1K decode tokens per GPU-second, which the team describes as a "3-5x improvement over H200." Those gains came from kernel work, not new silicon.

Per silicon area the picture is starker still: Blackwell is roughly 1,600 mm² against Hopper's ~800 mm². SemiAnalysis calculates the air-cooled B200 delivers only a 14% FP16 FLOPS improvement per unit silicon area.

Chart comparing vLLM prefill and decode throughput per GPU-second on GB200 against H200
Chart: vLLM

How much faster is B200 than H100?

About 3.16x on Llama 2 70B in MLPerf v5.0 — but that is FP4 against FP8. At matched precision the honest figure is the same 2.27x compute ratio the H200 shows, plus a larger 2.39x memory bandwidth gain (8 TB/s vs 3.35 TB/s) and 2.25x the capacity.

The H100's real problem versus both is 80 GB. A 70B model at FP16 does not fit on one. The H200's jump to 141 GB, not its compute, is what made Hopper viable for larger models.

How much does a B200 cost?

NVIDIA does not publish list prices for B200, H200 or HGX systems. What is public is rental. All figures below are on-demand per-GPU-hour list prices observed 24 September 2026.

Provider B200 H200 H100 SXM
Lambda $6.69 not offered $3.99
Runpod $6.79 $4.59 $3.49
Nebius $7.15 $4.50 $3.85
Together AI $8.19 $5.99 $3.99
CoreWeave $8.60 $6.31 $6.16
AWS (p6-b200.48xlarge ÷8) $14.24 $7.91 (p5en ÷8) $6.88 (p5 ÷8)

Two caveats. CoreWeave and AWS figures are 8-GPU instance prices divided by eight. And Nebius has published an increase effective 1 October 2026: B200 to $8.50, H200 to $5.40, H100 to $4.50.

On purchase cost, be sceptical of any specific number. Jensen Huang told CNBC in March 2024 that Blackwell GPUs would cost $30,000-$40,000 each, then clarified that pricing varies because NVIDIA sells systems rather than bare chips. The widely-cited ~$3M GB200 NVL72 rack figure originates from an analyst estimate, not NVIDIA. The ~$515,000 DGX B200 figure comes from a reseller listing. Current per-GPU transaction prices for B200 and H200 are not reliably public.

Those margins are precisely why every large AI company is now designing its own accelerator, and why challengers are attacking the memory side of the inference problem rather than the FLOPS side.

GB200 vs GB300: what actually changed

Neither is a GPU. GB200 is a superchip: one Grace CPU plus two Blackwell GPUs. GB200 NVL72 racks 36 of them into 72 GPUs. GB300 NVL72 uses Blackwell Ultra GPUs at the same 72-GPU count.

Spec GB200 NVL72 GB300 NVL72
GPUs 72 Blackwell 72 Blackwell Ultra
FP4 dense 720 PFLOPS 1,080 PFLOPS
FP4 sparse 1,440 PFLOPS 1,440 PFLOPS — unchanged
FP8 sparse 720 PFLOPS 720 PFLOPS — unchanged
HBM3e per GPU 186 GB 279 GB
Total HBM3e 13.4 TB 20 TB
NVLink aggregate 130 TB/s 130 TB/s — unchanged
Max TDP per GPU 1,200 W 1,400 W

The GB300's advertised 1.5x is a dense FP4 FLOPS ratio. Sparse FP4, FP8 and NVLink bandwidth are all identical. The genuine upgrade is 50% more HBM per GPU — which matters for long-context and reasoning workloads. NVIDIA's "50x AI factory output" claim is against Hopper, and its own footnote reveals GB300 measured at FP4 with Dynamo disaggregation against H100 at FP8 with in-flight batching.

Note that Blackwell Ultra cuts FP64 hard: 2,880 → 100 TFLOPS per rack. If you run FP64 HPC, GB300 is a downgrade.

What this means for you

If you're renting by the hour: do the arithmetic. At Runpod list, B200 costs 1.48x an H200. If your workload runs FP4 and hits ~2.9x, you get roughly 2x better throughput per dollar — take the B200. If you run FP8 or BF16 and see ~1.6x, you land near 1.08x — effectively break-even, and H200 is the safer pick on software maturity.

If you're specifying a cluster: buy B200 or GB200/GB300 for training and for inference on 100B+ models, where the 180 GB capacity and 1.8 TB/s NVLink change what fits. The NVLink doubling matters more than the FLOPS at tensor-parallel scale.

If you're serving models under ~70B at FP8: H200 remains the value pick through 2026. Supply is good, prices have softened, and the Hopper software stack needs no re-tuning.

If you're on H100 and deciding whether to move: go to H200 first if you're capacity-constrained, B200 if you're bandwidth-constrained and willing to invest in FP4 quantisation.

If you're evaluating non-NVIDIA silicon: the software stack, not the spec sheet, is the deciding variable — which is the whole substance of AMD's long campaign against CUDA.

Availability as of September 2026: B200 is broadly available on-demand across neoclouds and AWS. Azure has no 8-GPU B200 SKU — its Blackwell VM is a 4-GPU GB200 instance. GCP's A4 (B200) has no standard on-demand rate; it requires reservation, Spot or Flex-start. GB300 NVL72 is listed "Available Now" and deployed at scale by Microsoft, CoreWeave and Oracle.

Frequently Asked Questions

How much does a B200 cost?

NVIDIA does not publish a list price. Renting costs $6.69-$8.60 per GPU-hour on-demand at neoclouds and $14.24 on AWS as of 24 September 2026. Jensen Huang said in March 2024 Blackwell GPUs would run $30,000-$40,000 each, then qualified it. Current transaction prices are not reliably public.

Is the B200 worth it over the H200?

For FP4 inference and large-model training, yes — roughly 2x better throughput per dollar at current rental prices. For FP8 or BF16 inference, the measured advantage is closer to 1.6x against a 1.48x price premium, which is near break-even. Check your precision before paying the premium.

How much faster is B200 than H100?

3.16x on Llama 2 70B in MLPerf v5.0, but that compares B200 at FP4 to H100 at FP8. At matched precision it is 2.27x on compute and 2.39x on memory bandwidth. The bigger practical gain is capacity: 180 GB versus 80 GB.

What is the difference between GB200 and B200?

B200 is a single Blackwell GPU. GB200 is a superchip pairing one Grace CPU with two Blackwell GPUs over NVLink-C2C. The GB200 NVL72 rack combines 36 superchips into a 72-GPU NVLink domain at 130 TB/s — far beyond the 8-GPU domain of an HGX B200 server.

Why does my B200 benchmark slower than expected?

Most likely three causes: your workload is memory-bandwidth bound (only 1.67x better than H200), you are not running FP4 so you get 2.27x at best, or your inference stack lacks tuned Blackwell kernels. vLLM's own Blackwell gains came largely from kernel work landed through early 2026.


Editor's note — sources: Specifications verified against NVIDIA's H200 product page, NVIDIA's DGX B200 page and NVIDIA's Blackwell architecture page. Per-GPU throughput figures are derived from MLCommons' MLPerf Inference v5.0 results, including the per-submission precision field that shows B200 LLM entries running FP4 against H200 at FP8. Analysis of the marketing multiples draws on SemiAnalysis's Blackwell performance and TCO analysis, and the software-maturity section on vLLM's GB200 optimisation write-up. Rental prices were read from provider pricing pages on 24 September 2026 and will move; treat them as a snapshot. Purchase prices for B200 and H200 are not published by NVIDIA and no third-party figure is quoted here as fact.

Get Edgewisely in your inbox

Business stories that matter, free. Enter your email — no password, no account to set up.
jamie@example.com
Subscribe