Chips

A100 vs H100: Specs, Price and Which to Pick

Every A100 vs H100 comparison ranking on Google is published by a company that rents you GPUs. Here is a neutral one, built from NVIDIA datasheets and MLPerf results: the real dense-precision specs, the 1.7-4.3x speedup range, verified hourly pricing, and a straight verdict.

NVIDIA A100 80GB Tensor Core GPU
Image: NVIDIA

TL;DR

  • Pick the H100 for almost any new workload. NVIDIA's own Hopper whitepaper puts its FP16 Tensor Cores at 3x the A100's, and its FP8 path at 6x the A100's FP16 rate. Third-party MLPerf Inference v3.0 results land at 1.7x–4.3x depending on the model.
  • The A100 has no FP8 and no Transformer Engine. If you run modern quantised inference, that is the whole decision. There is no software fix.
  • The price gap is narrower than the performance gap. Verified published rates: $2.79/hr A100 vs $3.99/hr H100 at Lambda — 1.43x for 2–4x the work. The median premium across providers is roughly 2.1x.
  • Both are legacy parts. NVIDIA has shipped H200, B200/GB200, GB300 and Vera Rubin since. Neither chip appears on NVIDIA's current data center line card. Rent, don't buy.

A100 vs H100: which should you use?

In the a100 vs h100 decision, choose the H100. It delivers 1.7x to 4.3x more inference throughput than the A100 in MLPerf Inference v3.0 while renting for only 1.4x to 2.3x the hourly price, and its FP8 Tensor Cores have no A100 equivalent at all. Pick the A100 only when hourly cost dominates and FP16 is sufficient.

Why you should distrust most A100 vs H100 comparisons

Every commercial result currently ranking for this query is published by a company that rents you GPUs by the hour. Gcore, OpenMetal, Hyperstack, Voltage Park, JarvisLabs, Northflank and Vast.ai are all GPU cloud providers — seven of the top ten results.

That is not an accusation of dishonesty. It is a structural conflict. Each of those pages has a commercial interest in which SKU you pick, and several of them have stopped stocking one of the two chips.

The only non-vendor results are a Reddit thread and a PyTorch forum post, both from 2023. They predate Blackwell entirely.

Two things follow. Specs below come only from NVIDIA datasheets, the Hopper architecture whitepaper and CUDA documentation. Benchmark numbers come from MLCommons, not from NVIDIA's marketing slides.

What is the difference between the A100 and H100?

Different process node, a bigger chip, faster memory and — decisively — a new numeric format. The A100 is TSMC 7nm, 54.2 billion transistors. The H100 is TSMC 4N, 80 billion transistors.

The gap that matters most is FP8. The H100 adds FP8 Tensor Cores and a Transformer Engine that switches between FP8 and 16-bit precision per layer automatically. The A100 has no FP8 path whatsoever.

Spec A100 80GB SXM A100 80GB PCIe H100 SXM5 H100 PCIe
Architecture Ampere Ampere Hopper Hopper
Memory 80GB HBM2e 80GB HBM2e 80GB HBM3 80GB HBM2e
Memory bandwidth 2,039 GB/s 1,935 GB/s 3,352 GB/s 2,039 GB/s
FP8 Tensor none none 1,978.9 (3,957.8*) 1,513 (3,026*)
FP16/BF16 Tensor 312 (624*) 312 (624*) 989.4 (1,978.9*) 756 (1,513*)
TF32 Tensor 156 (312*) 156 (312*) 494.7 (989.4*) 378 (756*)
INT8 Tensor (TOPS) 624 (1,248*) 624 (1,248*) 1,978.9 (3,957.8*) 1,513 (3,026*)
FP32 19.5 19.5 66.9 51.2
FP64 / FP64 Tensor 9.7 / 19.5 9.7 / 19.5 33.5 / 66.9 25.6 / 51.2
NVLink 600 GB/s 600 GB/s (bridge) 900 GB/s —
PCIe Gen4 Gen4 Gen5 Gen5
TDP 400W 300W 700W 350W
Process / transistors 7nm / 54.2B 7nm / 54.2B 4N / 80B 4N / 80B
Launched May 2020 May 2020 Oct 2022 Oct 2022

TFLOPS unless noted. *Asterisked figures use structural sparsity — the inflated number vendor blogs usually quote. The dense figures are the realistic ones. Sources: the NVIDIA A100 datasheet and the NVIDIA H100 architecture whitepaper.

Note that the H100 PCIe carries HBM2e at 2,039 GB/s — identical bandwidth to the A100 80GB SXM. If a provider sells you "an H100" at a discount, check the form factor.

NVIDIA H100 Tensor Core GPU in SXM5 form factor
Image: NVIDIA

How much faster is the H100 than the A100?

Between 1.7x and 4.3x, depending entirely on the workload. Anyone quoting a single number is selling something.

NVIDIA's whitepaper claims "up to 9x faster AI training and up to 30x faster AI inference" on large language models. Treat those carefully: they are Transformer Engine best cases, and the whitepaper body states no model, batch size or sequence length behind them.

NVIDIA's own architectural figures in the same document are more modest and more useful. Figure 8 puts the H100 FP16 Tensor Core at 3x the A100's. Figure 10 puts H100 FP8 at 6x the A100's FP16 rate. Both match the dense spec table almost exactly.

For a third-party number, MLPerf Inference v3.0 is the cleanest comparison available. NVIDIA submitted 8x A100-SXM-80GB and 8x H100-SXM-80GB in the same round, same division, same software stack. Equal GPU counts, so the system ratio is the per-GPU ratio.

MLPerf Inference v3.0, Offline 8x A100 8x H100 Speedup
BERT 99.9% 14,620.9 62,398.7 4.27x
BERT 99% 28,275.7 73,107.6 2.59x
ResNet-50 324,612 727,437 2.24x
3D U-Net 30.62 55.11 1.80x
RNN-T 106,221 179,738 1.69x

Samples per second, from the MLCommons MLPerf Inference v3.0 results.

The pattern is clear. Vision and speech models land near 1.7x–2.2x. Transformer workloads at high accuracy targets, where FP8 and the larger memory bandwidth both bite, reach 4x and above.

One honest caveat: host CPUs differ between the paired systems, so these are system-level results attributed to the GPU.

What does each GPU cost to rent?

Published on-demand rates, verified 1 October 2026:

Provider A100 80GB H100 80GB Ratio
DataCrunch $1.16 $1.99 1.72x
Jarvislabs $1.49 $2.69 1.81x
Hyperstack (SXM) $1.60 $3.20 2.00x
Lambda (SXM) $2.79 $3.99 1.43x
CoreWeave $2.70 $6.16 2.28x
Azure (ND v4 / v5) $4.10 $12.29 3.00x

Lambda's published pricing is the clearest per-GPU list. CoreWeave and Azure publish per-instance rates; those two rows are divided by eight.

The median premium is around 2.1x. Set that against 2–4x the throughput and the H100 wins on cost-per-token at most providers — decisively at Lambda's 1.43x.

Azure's 3.00x is the one place the A100 still looks genuinely cheap per unit of work.

Also telling: Nebius, Together.ai, DigitalOcean and Voltage Park no longer list the A100 at all. The H100 is now the oldest NVIDIA part on several price lists. The same economics drive the newer accelerator market we covered in Groq vs Cerebras.

Should you still buy or rent an A100 in 2026?

Rent, never buy. And check what you are actually renting it for.

The A100 is a legacy part commercially. It has been removed from NVIDIA's current data center GPU lineup page and from the official line card, which now runs from GB300 NVL72 down to HGX H200. The A100 product page is still live, but it is delisted.

The software picture is the opposite, and this is where most comparisons get it wrong. The A100 is not out of support. CUDA 13 dropped Maxwell, Pascal and Volta — everything before Turing — but the CUDA Toolkit release notes confirm Ampere was retained. NVIDIA AI Enterprise Infra 8.x still lists A100 40GB and 80GB as supported GPUs on its end-of-life notices page, last updated 9 September 2026.

So the A100 is dead to NVIDIA's sales catalogue and alive in its software stack. Driver and CUDA support are not your risk. Spare parts and hourly rates are.

For a new buy, both chips are several generations back. NVIDIA has since shipped the H200 (141GB HBM3e, 4.8 TB/s), B200/GB200, GB300 NVL72 and Vera Rubin NVL72. Hopper has no FP4; Blackwell introduced it. We compared the two most relevant current parts in H200 vs B200, and the strategic reason every large buyer now wants custom silicon in why every AI giant now wants its own chip.

What this means for you

Fine-tuning a 7B model as a solo researcher. The A100 80GB is fine, and it is the cheaper experiment. A 7B model in LoRA or QLoRA fits comfortably in 80GB and you are rarely bandwidth-bound. At DataCrunch's $1.16/hr you save real money on long, interruptible runs.

Serving production inference. H100, without hesitation. FP8 halves your weight footprint and roughly doubles throughput against FP16, and the A100 cannot do it at any price. If you are choosing a managed endpoint instead of raw GPUs, see our AI inference provider comparison.

Training from scratch. H100 or newer. You want the 900 GB/s NVLink and 3,352 GB/s HBM3 far more than the cheaper hourly rate — interconnect dominates at multi-node scale. If your run is large enough to justify the booking, price H200 and B200 too.

FP64 and HPC. H100, by a wide margin: 66.9 vs 19.5 TFLOPS FP64 Tensor. The A100's historical HPC strength does not hold against Hopper.

Buying hardware rather than renting. Don't buy either. Both are end-of-life for new production, and we could not find published, verifiable secondhand A100 80GB pricing. Rent, or buy current-generation silicon.

Frequently Asked Questions

Why choose an H100 over an A100 for LLM work?

FP8. The H100's Transformer Engine switches between FP8 and 16-bit per layer, halving memory footprint and roughly doubling throughput. The A100 has no FP8 path at all. In MLPerf Inference v3.0, BERT 99.9% ran 4.27x faster on H100 at equal GPU counts — the largest gap of any tested workload.

Is there an H100 vs A100 memory usage difference?

Capacity is identical at 80GB, but usage differs in practice. FP8 weights on H100 occupy half the space of FP16, so the same model leaves more room for KV cache and larger batches. The H100 also has 50MB of L2 cache against the A100's 40MB, plus 228KB shared memory per SM versus 164KB.

Is the A100 still worth it in 2026?

For cheap fine-tuning and FP16 inference, yes. NVIDIA AI Enterprise Infra 8.x still lists it as supported and CUDA 13 retained Ampere, so software is not your risk. But it is off NVIDIA's line card, several providers have dropped it, and the H100 usually costs under 2.1x for 2–4x the work.

How much faster is an H100 than an A100?

Between 1.7x and 4.3x in MLPerf Inference v3.0 at equal GPU counts — 1.69x on RNN-T, 2.24x on ResNet-50, 4.27x on BERT 99.9%. NVIDIA's own whitepaper figures cite 3x on FP16 Tensor Cores and 6x for FP8 versus A100 FP16. Ignore the "up to 30x" marketing number.


Editor's note — sources. Specifications are taken from the NVIDIA A100 datasheet and the NVIDIA H100 Tensor Core GPU architecture whitepaper (Table 3), both linked above. Throughput comparisons use MLCommons MLPerf Inference v3.0 closed-division results for NVIDIA's own 8x A100-SXM-80GB and 8x H100-SXM-80GB submissions. Lifecycle and software-support claims come from the NVIDIA CUDA Toolkit release notes and the NVIDIA AI Enterprise end-of-life notices. Rental pricing was read from each provider's own published pricing page on 1 October 2026; only Lambda is linked, to stay within our source cap. Pricing changes frequently — DataCrunch, Jarvislabs, Hyperstack, CoreWeave and Azure figures were current on that date.

We could not verify the following and have left it out: published secondhand A100 80GB pricing; the measured basis for NVIDIA's "9x training / 30x inference" claim, which the whitepaper does not state; and any A100-versus-H100 GPT-3 175B training ratio, because no A100 GPT-3 training submission exists in MLPerf. A widely circulated claim that an 8x8 A100 cluster trains GPT-3 in roughly 28 minutes against 7 minutes on H100 could not be traced to any NVIDIA or MLCommons source. AWS was excluded from the pricing table because only reserved Capacity Block rates were readable, not on-demand.

Get Edgewisely in your inbox

Business stories that matter, free. Enter your email — no password, no account to set up.
jamie@example.com
Subscribe