Compute Comparison
NVIDIAAda Lovelace2023

L4

Extremely low TDP (72W). Ideal for high-density inference racks. Best performance-per-watt in its class.

VRAM
24GB
GDDR6
FP16
60.6
TFLOPS
Bandwidth
300.0
GB/s
TDP
72W
power
Best for:Efficient inferenceEdge deploymentsLow-power serving

L4 Overview

The NVIDIA L4 is an Ada Lovelace data-center GPU optimized for efficiency rather than maximum model size. It provides 24GB of GDDR6 and 60.6 TFLOPS of FP16/BF16 performance in an exceptionally low 72W envelope, which makes dense, power-conscious inference deployments practical.

The 300 GB/s memory bandwidth is modest, so autoregressive models that repeatedly stream large weight sets will become memory-bound sooner than on HBM GPUs. Its 24GB capacity is a useful fit for compact models and quantized serving, but it does not leave much headroom for large FP16 weights, long contexts, or oversized batches.

The L4 excels at efficient inference, edge-adjacent deployment, video and vision pipelines, and high-density serving of modest models. It is not designed for large-model training, and its lack of NVLink means each card remains an independent 24GB resource.

Memory

VRAM24 GB
Memory TypeGDDR6
Bandwidth300 GB/s

Compute Performance

FP3230.3 TFLOPS
FP1660.6 TFLOPS
BF1660.6 TFLOPS
INT8242 TOPS

Hardware

ArchitectureAD104
GenerationAda Lovelace
Process NodeTSMC 4N
Transistors35.8B
TDP72 W
InterconnectPCIe 4.0
Release Year2023

Relative Performance

FP16 Compute1%
VRAM Capacity8%
Mem Bandwidth2%

Relative to highest-spec GPU in database

Limitations

Only 24GB VRAM — limits to ~13B models at FP16
Low memory bandwidth (300 GB/s) — bottleneck for large batches
No NVLink — single-card only

Live Cloud PricingOn-demand hourly rates

Loading live prices…

Compare L4 vs…

Use Case Guidance

Efficient inference
Edge deployments
Low-power serving

LLM Model Size Guidance

Max model (FP16)~12Bparameters at FP16 precision
Max model (INT8)~24Bparameters at INT8 precision
Max model (INT4)~48Bparameters at INT4/GGUF

Estimates only. Actual capacity depends on context length, KV cache, and framework overhead.

Related Guides

LLM APIs Running on This GPU Class

Providers that serve frontier LLM inference on Ada Lovelace-class hardware.

Browse all 42 LLM models Self-hosting guide: best LLM models for this GPU

Related GPUs

Frequently Asked Questions

How much VRAM does the L4 have?

The L4 has 24GB of GDDR6 memory with 300 GB/s bandwidth. This enables running models up to approximately 48B parameters at INT4 precision, 24B at INT8, or 12B at FP16.

What is the FP16 performance of the L4?

The L4 delivers 60.6 TFLOPS of FP16 performance and 60.6 TFLOPS BF16. INT8 throughput is 242 TOPS. For transformer inference, memory bandwidth (300 GB/s) is often the binding constraint rather than raw TFLOPS.

What is the L4 best used for?

The L4 is best suited for: Efficient inference, Edge deployments, Low-power serving. Extremely low TDP (72W). Ideal for high-density inference racks. Best performance-per-watt in its class.

What interconnect does the L4 use?

The L4 uses PCIe 4.0. Without NVLink, VRAM cannot be pooled across multiple cards — the single-card capacity is the hard ceiling for model size.

What LLM model sizes can the L4 run?

With 24GB of GDDR6, the L4 can run models up to approximately 12B parameters at FP16 (2 bytes/param), 24B at INT8 (1 byte/param), or 48B at INT4/GGUF (0.5 bytes/param). These are estimates — actual capacity depends on context length, KV cache size, and framework overhead. Longer context windows require more KV cache memory, reducing the effective model size that fits.

How does the L4 compare to the A100 for LLM inference?

The L4 has 60.6 TFLOPS FP16 vs the A100 80GB's 312 TFLOPS, and 300 GB/s memory bandwidth vs the A100's 2,039 GB/s. For memory-bound autoregressive LLM inference, bandwidth is the primary determinant of tokens-per-second. The A100's higher bandwidth gives it a throughput advantage for large model inference, despite the L4's lower cost.

What is the power consumption of the L4?

The L4 has a TDP (Thermal Design Power) of 72W. This is the maximum sustained power draw under full load. For data center deployments, total rack power consumption is typically 1.2–1.5× the GPU TDP when accounting for CPU, memory, networking, and cooling overhead. At 72W, the L4 is in the low-power tier — enables high-density deployments with standard rack power.

Ready to rent?

Compare L4 prices across 97+ providers

Live on-demand & spot rates · monthly cost estimates · availability status

Compare rental prices