Compute Comparison
NVIDIAAda Lovelace2023

L40S

Best Ada Lovelace data center GPU. High FP32, 48GB GDDR6. Strong inference-per-dollar.

VRAM
48GB
GDDR6
FP16
183.0
TFLOPS
Bandwidth
864.0
GB/s
TDP
350W
power
Best for:Inference servingGenerative AIMixed compute + graphics

L40S Overview

The NVIDIA L40S is the best-value data center GPU for inference serving in 2025, offering 48GB of GDDR6 memory and 183 TFLOPS of FP16 throughput at a significantly lower price point than the H100 or A100. Built on the Ada Lovelace architecture (AD102 die, TSMC 4N process), it was designed specifically for inference workloads — its 91.6 TFLOPS FP32 figure is unusually high for a data center GPU, making it well-suited for mixed compute and graphics workloads alongside AI inference.

The 48GB GDDR6 configuration fits models up to ~24B parameters at FP16 or ~48B at INT8. Memory bandwidth is 864 GB/s — substantially lower than HBM-based alternatives like the A100 (2,039 GB/s) or H100 (3,350 GB/s). This bandwidth gap means the L40S is less competitive for memory-bound autoregressive inference of large models, but for batch inference, image generation, and smaller model serving, the lower cost per hour often makes it the better choice. The L40S uses PCIe 4.0 with no NVLink, so VRAM cannot be pooled across cards.

The L40S is the right choice for inference serving of models up to 30B parameters, generative AI workloads (Stable Diffusion, image generation), and mixed compute environments where both AI inference and graphics rendering are needed. It is not well-suited for training large models (limited VRAM, no NVLink) or for memory-bound inference of 70B+ models where HBM bandwidth matters. At $0.80–$1.80/hr on-demand, it typically delivers better cost-per-token than A100 for models that fit in 48GB.

Memory

VRAM48 GB
Memory TypeGDDR6
Bandwidth864 GB/s

Compute Performance

FP3291.6 TFLOPS
FP16183 TFLOPS
BF16183 TFLOPS
INT8366 TOPS

Hardware

ArchitectureAD102
GenerationAda Lovelace
Process NodeTSMC 4N
Transistors76.3B
TDP350 W
InterconnectPCIe 4.0
Release Year2023

Relative Performance

FP16 Compute2%
VRAM Capacity17%
Mem Bandwidth5%

Relative to highest-spec GPU in database

Limitations

GDDR6 memory bandwidth (864 GB/s) far below HBM alternatives
No NVLink — cannot pool VRAM across cards
Not ideal for training — optimized for inference

Live Cloud PricingOn-demand hourly rates

Loading live prices…

Compare L40S vs…

Use Case Guidance

Inference serving
Generative AI
Mixed compute + graphics

LLM Model Size Guidance

Max model (FP16)~24Bparameters at FP16 precision
Max model (INT8)~48Bparameters at INT8 precision
Max model (INT4)~96Bparameters at INT4/GGUF

Estimates only. Actual capacity depends on context length, KV cache, and framework overhead.

Related Guides

LLM APIs Running on This GPU Class

Providers that serve frontier LLM inference on Ada Lovelace-class hardware.

Browse all 42 LLM models Self-hosting guide: best LLM models for this GPU

Related GPUs

Frequently Asked Questions

How much VRAM does the L40S have?

The L40S has 48GB of GDDR6 memory with 864 GB/s bandwidth. This enables running models up to approximately 96B parameters at INT4 precision, 48B at INT8, or 24B at FP16.

What is the FP16 performance of the L40S?

The L40S delivers 183 TFLOPS of FP16 performance and 183 TFLOPS BF16. INT8 throughput is 366 TOPS. For transformer inference, memory bandwidth (864 GB/s) is often the binding constraint rather than raw TFLOPS.

What is the L40S best used for?

The L40S is best suited for: Inference serving, Generative AI, Mixed compute + graphics. Best Ada Lovelace data center GPU. High FP32, 48GB GDDR6. Strong inference-per-dollar.

What interconnect does the L40S use?

The L40S uses PCIe 4.0. Without NVLink, VRAM cannot be pooled across multiple cards — the single-card capacity is the hard ceiling for model size.

What LLM model sizes can the L40S run?

With 48GB of GDDR6, the L40S can run models up to approximately 24B parameters at FP16 (2 bytes/param), 48B at INT8 (1 byte/param), or 96B at INT4/GGUF (0.5 bytes/param). These are estimates — actual capacity depends on context length, KV cache size, and framework overhead. Longer context windows require more KV cache memory, reducing the effective model size that fits.

How does the L40S compare to the A100 for LLM inference?

The L40S has 183 TFLOPS FP16 vs the A100 80GB's 312 TFLOPS, and 864 GB/s memory bandwidth vs the A100's 2,039 GB/s. For memory-bound autoregressive LLM inference, bandwidth is the primary determinant of tokens-per-second. The A100's higher bandwidth gives it a throughput advantage for large model inference, despite the L40S's lower cost.

What is the power consumption of the L40S?

The L40S has a TDP (Thermal Design Power) of 350W. This is the maximum sustained power draw under full load. For data center deployments, total rack power consumption is typically 1.2–1.5× the GPU TDP when accounting for CPU, memory, networking, and cooling overhead. At 350W, the L40S is in the mid-range tier — compatible with standard data center power infrastructure.

Ready to rent?

Compare L40S prices across 97+ providers

Live on-demand & spot rates · monthly cost estimates · availability status

Compare rental prices