> ## Content Index
> Fetch the complete content index at: https://www.edgewisely.com/llms.txt
> Use this file to discover other available public pages before exploring further.

# L40S vs A100: Which GPU for Inference in 2026?
- URL: https://www.edgewisely.com/l40s-vs-a100/
- Published: 2026-10-02T06:00:52.000Z
- Updated: 2026-10-02T06:00:52.000Z
- Description: L40S vs A100 compared on bandwidth, FP8, VRAM and 2026 rental rates. Which wins for LLM inference, and which for video and rendering.
- Author: John Karpentar
- Tags: Chips, Deep Tech

**TL;DR**

- **Verdict:** The **L40S** wins on mixed AI-plus-visual work and FP8-quantised models. The **A100 80GB** wins on memory-bound LLM serving and anything needing multi-GPU scale-out.
- **The number that decides it:** A100 80GB moves **1,935–2,039 GB/s**. The L40S moves **864 GB/s** — roughly 2.2–2.4× less, and token generation is memory-bound.
- **Rental, checked 2 October 2026:** Runpod lists L40S at **$1.09/hr** and A100 80GB at **$1.59/hr**. CoreWeave works out to **$2.25** and **$2.70** per GPU-hour. Per dollar of memory bandwidth, the A100 is the better buy on both.
- **Neither card is end-of-life.** NVIDIA still lists A100 40/80GB and L40S as supported hardware as of September 2026.

Pick the **L40S** if your workload mixes AI with graphics, video or rendering, or runs quantised models in **FP8**. Pick the **A100** for memory-bound LLM serving and multi-GPU jobs: its **80GB HBM2e** moves **2,039 GB/s** against the L40S's **864 GB/s**, and NVLink lets it scale where the L40S cannot.

One thing to flag before the detail. Most of what ranks for this comparison is published by companies that rent or sell these two GPUs — GPU clouds, hosting firms and hardware resellers. Several of them no longer stock one of the two parts. We sell neither.

## What is the NVIDIA L40S?

The L40S is NVIDIA's **Ada Lovelace** data centre GPU built for workloads that mix AI with graphics. Per [NVIDIA's L40S product page](https://www.nvidia.com/en-us/data-center/l40s/?ref=edgewisely.com), it has **48GB of GDDR6 with ECC**, 18,176 CUDA cores, 142 third-generation RT cores and 568 fourth-generation Tensor cores, in a 350W dual-slot PCIe card.

The part that gets missed: it is a *graphics* card as well as an AI card. It carries **three NVENC and three NVDEC engines** with AV1 encode and decode, and RT cores rated at **212 TFLOPS**. The A100 has none of that.

It also supports **FP8** through Ada's Transformer Engine. The A100 does not — NVIDIA's A100 spec table lists no FP8 rate, because Ampere's Tensor cores don't implement it.

What it gives up: **no NVLink** and **no MIG**. The L40S is a single-card part that talks to its neighbours over PCIe Gen4 at 64 GB/s, and cannot be partitioned into isolated instances.

## L40S vs A100: the spec differences that matter

All figures from NVIDIA's own product specifications. Where a Tensor figure has two values, the first is **dense** and the second is **with sparsity** — NVIDIA's marketing usually quotes the sparse number.

|                  | **NVIDIA L40S**         | **NVIDIA A100 80GB**                         |
| ---------------- | ----------------------- | -------------------------------------------- |
| Architecture     | Ada Lovelace            | Ampere                                       |
| GPU memory       | 48GB GDDR6 with ECC     | 80GB HBM2e                                   |
| Memory bandwidth | **864 GB/s**            | **1,935 GB/s** (PCIe) · **2,039 GB/s** (SXM) |
| FP8 Tensor       | **733** \| 1,466 TFLOPS | Not supported                                |
| FP16/BF16 Tensor | 362 \| 733 TFLOPS       | **312** \| 624 TFLOPS                        |
| TF32 Tensor      | 183 \| 366 TFLOPS       | 156 \| 312 TFLOPS                            |
| INT8 Tensor      | 733 \| 1,466 TOPS       | 624 \| 1,248 TOPS                            |
| FP32             | **91.6 TFLOPS**         | 19.5 TFLOPS                                  |
| FP64             | Not published           | **9.7 TFLOPS** (19.5 Tensor)                 |
| RT cores         | 142, **212 TFLOPS**     | None                                         |
| NVENC / NVDEC    | **3× / 3×**, AV1        | None                                         |
| NVLink           | **No**                  | **Yes, 600 GB/s**                            |
| MIG              | No                      | Up to 7 instances @ 10GB                     |
| Form factor      | PCIe only               | PCIe or SXM                                  |
| Max power        | 350W                    | 300W (PCIe) · 400W (SXM)                     |

Four of those rows decide almost every real deployment.

**Memory bandwidth.** The A100 moves 2.2–2.4× more. For LLM decode this is the number that matters, because generating each token requires reading the whole model's weights out of memory.

**FP8.** A well-quantised model on the L40S runs at **733 dense TFLOPS**, against the A100's best dense rate of **312 TFLOPS** in FP16/BF16\. That is 2.35× more compute, and it claws back much of the bandwidth gap on the prefill-heavy and compute-bound parts of a request.

**Capacity.** 48GB versus 80GB decides what fits on one card. A 70B model in 4-bit fits on a single A100 with room for KV cache. On a 48GB L40S it is tight to impossible without aggressive quantisation.

**NVLink.** The A100 scales to multi-GPU model parallelism at 600 GB/s, per [NVIDIA's A100 specifications](https://www.nvidia.com/en-us/data-center/a100/?ref=edgewisely.com). The L40S falls back to PCIe at 64 GB/s, which is roughly 9× slower. If your model needs more than one card, this is disqualifying for the L40S.

For the Ampere background and how the A100 compares upward, see our [A100 vs H100 breakdown](https://www.edgewisely.com/a100-vs-h100/). For the current generation above both, see [H200 vs B200](https://www.edgewisely.com/h200-vs-b200/).

## Which is faster for LLM inference?

Depends on which half of the request you mean, and nobody on page one says this clearly.

**Token generation (decode) is memory-bound.** Throughput scales with memory bandwidth, so the A100 leads by roughly the bandwidth ratio — about **2.2–2.4×** per card, before quantisation.

**Prompt processing (prefill) is compute-bound.** Here the L40S's FP8 path gives it **733 dense TFLOPS** against the A100's **312**, and the L40S can come out ahead on long prompts.

So the honest answer: for long-context, short-output work — classification, RAG retrieval scoring, embedding, reranking — the L40S competes well. For long-output chat and agent loops, the A100 pulls ahead.

NVIDIA does not publish a head-to-head L40S-versus-A100 inference benchmark. Its own L40S performance charts compare against the A40, not the A100\. We could not verify a first-party direct comparison, so treat any single "X% faster" claim on this topic with suspicion.

Which serving stack you run matters as much as the card. Our [vLLM vs SGLang comparison](https://www.edgewisely.com/vllm-vs-sglang/) covers the two engines most teams choose between.

## What do the L40S and A100 cost to rent in 2026?

Rates below were read off each provider's published pricing page on **2 October 2026**.

| Provider                       | L40S (48GB)            | A100 80GB                   |
| ------------------------------ | ---------------------- | --------------------------- |
| **Runpod**                     | **$1.09**/hr           | **$1.59**/hr (PCIe and SXM) |
| **CoreWeave** (8-GPU node ÷ 8) | **$2.25**/GPU-hr       | **$2.70**/GPU-hr            |
| **Nebius**                     | from **$1.55**/hr      | Not offered                 |
| **Lambda**                     | Not offered            | **$2.79**/hr (SXM)          |
| **Hyperstack**                 | Not offered (L40 only) | from **$1.35**/hr           |

Note the gaps. Nebius lists the L40S but no A100\. Lambda and Hyperstack list A100s but no L40S. Availability is genuinely split, and that will constrain your choice as much as the specs do. The [GPU cloud provider roundup](https://www.edgewisely.com/top-7-gpu-cloud-providers-for-ai-training-and-inference-2026/) covers who stocks what.

Now the number that actually governs cost-per-token on memory-bound serving — bandwidth bought per dollar per hour, using the conservative **1,935 GB/s** PCIe figure for the A100:

|                                                                       | L40S                | A100 80GB                 | A100 advantage |
| --------------------------------------------------------------------- | ------------------- | ------------------------- | -------------- |
| **[Runpod](https://www.runpod.io/pricing?ref=edgewisely.com)**        | \~790 GB/s per $/hr | \~**1,220** GB/s per $/hr | **1.5×**       |
| **[CoreWeave](https://www.coreweave.com/pricing?ref=edgewisely.com)** | \~380 GB/s per $/hr | \~**715** GB/s per $/hr   | **1.9×**       |

On both clouds, the A100 delivers more memory bandwidth per dollar — despite costing more per hour. The L40S being the cheaper card does not make it the cheaper way to generate tokens. That is Edgewisely's own calculation from the two published rate cards, not a vendor figure.

Where the L40S reverses this: it bills one card at a time with no node minimum, and it does video transcode and rendering that the A100 physically cannot. If your pipeline decodes video *and* runs a model, one L40S replaces a CPU transcode farm plus a GPU.

## Is the A100 end-of-life?

No. As of [NVIDIA's AI Enterprise lifecycle notices](https://docs.nvidia.com/ai-enterprise/lifecycle/latest/eol-notices.html?ref=edgewisely.com), last updated 9 September 2026, the deprecated data centre hardware is the **V100** and five workstation-class cards, removed from Infra 8.0 onward.

The **A100 40GB and 80GB** are listed as *currently supported* GPUs on Infra 8.x — explicitly named as a migration target for customers leaving V100\. The **L40S** is listed as supported in the same table.

The A100 is six years old and has been pushed down the rental stack by Hopper and Blackwell, but it is not deprecated. [Tom's Hardware reported in August 2026](https://www.tomshardware.com/tech-industry/coreweave-ceo-mike-intrator-says-it-has-signed-an-a100-contract-running-into-2029?ref=edgewisely.com) that CoreWeave has signed A100 capacity contracts running into **2029**.

## What this means for you

**If you're serving a 7B–13B model for chat or agents:** take the **A100 80GB**. Decode is memory-bound, the A100 gives you 1.5–1.9× more bandwidth per dollar, and 80GB leaves real room for KV cache at high concurrency.

**If you're running a quantised model with heavy prompts and short outputs:** take the **L40S**. FP8 at 733 dense TFLOPS plus a lower hourly rate wins on RAG scoring, reranking, classification and embedding.

**If your workload touches video, rendering, Omniverse or virtual workstations:** take the **L40S**, and stop comparing. NVENC, NVDEC and RT cores are not optional extras the A100 happens to lack — the A100 cannot do this work at all.

**If your model needs more than one GPU:** take the **A100**. NVLink at 600 GB/s versus PCIe at 64 GB/s is not a margin you optimise around.

**If you need FP64 for HPC or simulation:** take the **A100**. 9.7 TFLOPS FP64, rising to 19.5 on Tensor cores. NVIDIA publishes no FP64 figure for the L40S.

**If you need to slice one GPU across tenants:** take the **A100**. MIG partitions it into up to seven isolated instances at 10GB each. The L40S does not support MIG.

**If you're buying in 2026 rather than renting:** check whether an **RTX PRO 6000 Blackwell** (96GB) belongs in your shortlist. Runpod lists it at $2.09/hr and CoreWeave at $2.50/GPU-hr, and it supersedes much of what the L40S was built for.

## Frequently Asked Questions

### Is the L40S better than the A100?

Not generally — they are different tools. The L40S beats the A100 on FP8 compute, FP32, rendering and video, at a lower hourly rate. The A100 beats it on memory bandwidth, VRAM capacity, FP64 and multi-GPU scaling. Choose on your bottleneck, not on a ranking.

### Is the L40S good for LLM inference?

Yes, with conditions. It handles quantised models up to roughly 30B well, and its **FP8** support makes it strong on prompt-heavy, short-output work like RAG and reranking. Its **864 GB/s** bandwidth limits token generation speed, and 48GB caps model size on a single card.

### How much VRAM does the L40S have compared to the A100?

The L40S has **48GB of GDDR6 with ECC**. The A100 ships in 40GB and 80GB HBM2e versions, with 80GB the common rental option. So the A100 80GB gives you **67% more capacity** and, more importantly, far higher bandwidth — 1,935 GB/s versus 864 GB/s.

### Should I rent an L40S or an A100?

Rent the **A100 80GB** for memory-bound LLM serving, multi-GPU jobs or FP64 work. Rent the **L40S** for FP8-quantised inference, video transcode, rendering or mixed AI-plus-graphics pipelines. Check availability first — several major clouds stock only one of the two.

---

**Editor's note — sources:** Every spec figure was read from NVIDIA's own L40S and A100 product specification tables on 2 October 2026, not from reseller summaries. Tensor throughput is given as dense first, sparse second; NVIDIA's headline marketing figures are the sparse numbers. Rental rates were read the same day from Runpod and CoreWeave's published pricing pages, plus nebius.com/prices, lambda.ai/instances and hyperstack.cloud/gpu-pricing for availability; CoreWeave publishes 8-GPU node rates, divided by eight here. The bandwidth-per-dollar table is Edgewisely's calculation from those published rates. Lifecycle status comes from NVIDIA's AI Enterprise end-of-life notices. No first-party NVIDIA benchmark comparing the L40S directly against the A100 could be located, and none is cited.