Compute Comparison
NVIDIATuring2018

T4

Ultra-low TDP (70W). Extremely cheap per hour. Ideal for high-density inference racks. Widely available on GCP and AWS.

VRAM
16GB
GDDR6
FP16
65.0
TFLOPS
Bandwidth
320.0
GB/s
TDP
70W
power
Best for:Cost-efficient inferenceHigh-density servingLegacy workloads

T4 Overview

The NVIDIA T4 is a Turing-generation data center GPU designed for high-density inference serving. At just 70W TDP — the lowest of any data center GPU in widespread use — it enables extremely dense rack deployments: a single 1U server can host multiple T4 cards. Released in 2018, it delivers 65 TFLOPS of FP16 throughput and 130 TOPS of INT8, making it well-suited for serving quantized models at scale. It is the most widely available low-cost GPU on Google Cloud Platform and AWS.

The 16GB GDDR6 configuration provides 320 GB/s of memory bandwidth — the lowest of any current data center GPU. This limits the T4 to models up to approximately 7B parameters at FP16, or 16B at INT8. The low bandwidth makes it unsuitable for memory-bound autoregressive inference of larger models, but for batch inference of 7B models and smaller, the low cost per hour often compensates. The T4 uses PCIe 3.0 with no NVLink and lacks BF16 hardware acceleration.

The T4 is the right choice for high-volume, cost-sensitive inference of models up to 7B parameters — particularly when deployed in large quantities. It is the most cost-effective GPU for serving quantized BERT, DistilBERT, and similar encoder models at scale. It is not suitable for training, for models above 7B at FP16, or for workloads requiring BF16 precision. At $0.10–$0.35/hr on GCP, AWS, and other providers, it remains the cheapest data center GPU for inference despite its age.

Memory

VRAM16 GB
Memory TypeGDDR6
Bandwidth320 GB/s

Compute Performance

FP328.1 TFLOPS
FP1665 TFLOPS
BF1665 TFLOPS
INT8130 TOPS

Hardware

ArchitectureTU104
GenerationTuring
Process NodeTSMC 12nm
Transistors13.6B
TDP70 W
InterconnectPCIe 3.0
Release Year2018

Relative Performance

FP16 Compute1%
VRAM Capacity6%
Mem Bandwidth2%

Relative to highest-spec GPU in database

Limitations

Only 16GB GDDR6 — limits to ~7B models at FP16
Very low FP32 throughput (8.1 TFLOPS)
Older Turing architecture — no BF16 hardware support

Live Cloud PricingOn-demand hourly rates

Loading live prices…

Compare T4 vs…

Use Case Guidance

Cost-efficient inference
High-density serving
Legacy workloads

LLM Model Size Guidance

Max model (FP16)~8Bparameters at FP16 precision
Max model (INT8)~16Bparameters at INT8 precision
Max model (INT4)~32Bparameters at INT4/GGUF

Estimates only. Actual capacity depends on context length, KV cache, and framework overhead.

Related Guides

LLM APIs Running on This GPU Class

Providers that serve frontier LLM inference on Turing-class hardware.

Browse all 42 LLM models

Related GPUs

Frequently Asked Questions

How much VRAM does the T4 have?

The T4 has 16GB of GDDR6 memory with 320 GB/s bandwidth. This enables running models up to approximately 32B parameters at INT4 precision, 16B at INT8, or 8B at FP16.

What is the FP16 performance of the T4?

The T4 delivers 65 TFLOPS of FP16 performance and 65 TFLOPS BF16. INT8 throughput is 130 TOPS. For transformer inference, memory bandwidth (320 GB/s) is often the binding constraint rather than raw TFLOPS.

What is the T4 best used for?

The T4 is best suited for: Cost-efficient inference, High-density serving, Legacy workloads. Ultra-low TDP (70W). Extremely cheap per hour. Ideal for high-density inference racks. Widely available on GCP and AWS.

What interconnect does the T4 use?

The T4 uses PCIe 3.0. Without NVLink, VRAM cannot be pooled across multiple cards — the single-card capacity is the hard ceiling for model size.

What LLM model sizes can the T4 run?

With 16GB of GDDR6, the T4 can run models up to approximately 8B parameters at FP16 (2 bytes/param), 16B at INT8 (1 byte/param), or 32B at INT4/GGUF (0.5 bytes/param). These are estimates — actual capacity depends on context length, KV cache size, and framework overhead. Longer context windows require more KV cache memory, reducing the effective model size that fits.

How does the T4 compare to the A100 for LLM inference?

The T4 has 65 TFLOPS FP16 vs the A100 80GB's 312 TFLOPS, and 320 GB/s memory bandwidth vs the A100's 2,039 GB/s. For memory-bound autoregressive LLM inference, bandwidth is the primary determinant of tokens-per-second. The A100's higher bandwidth gives it a throughput advantage for large model inference, despite the T4's lower cost.

What is the power consumption of the T4?

The T4 has a TDP (Thermal Design Power) of 70W. This is the maximum sustained power draw under full load. For data center deployments, total rack power consumption is typically 1.2–1.5× the GPU TDP when accounting for CPU, memory, networking, and cooling overhead. At 70W, the T4 is in the low-power tier — enables high-density deployments with standard rack power.

Ready to rent?

Compare T4 prices across 97+ providers

Live on-demand & spot rates · monthly cost estimates · availability status

Compare rental prices