Compute Comparison
NVIDIAVolta2018

V100 32GB

Previous-gen workhorse. 32GB HBM2 at low cost. Still widely available. No FP8 or BF16 hardware support.

VRAM
32GB
HBM2
FP16
112.0
TFLOPS
Bandwidth
900.0
GB/s
TDP
300W
power
Best for:Legacy ML trainingHPC workloadsBudget multi-GPU setups

V100 32GB Overview

The NVIDIA V100 32GB is a Volta-generation data center GPU that was the flagship AI accelerator from 2018 to 2020, before the A100 superseded it. Built on the GV100 die (TSMC 12nm), it delivers 112 TFLOPS of FP16 throughput with 32GB of HBM2 at 900 GB/s bandwidth. It was the first GPU to introduce Tensor Cores for accelerated matrix multiplication, establishing the hardware foundation for modern deep learning training.

The 32GB HBM2 configuration fits models up to approximately 16B parameters at FP16. NVLink 2.0 at 300 GB/s bidirectional enables multi-GPU tensor parallelism, though at lower bandwidth than the A100's NVLink 3.0 (600 GB/s) or H100's NVLink 4.0 (900 GB/s). The V100 lacks BF16 and FP8 hardware support — both introduced in later architectures — which limits its efficiency for modern transformer training compared to A100 and H100.

The V100 32GB is best suited for legacy ML training workloads, HPC applications, and budget-conscious teams that need more than 16GB of VRAM at the lowest possible cost. It remains widely available on AWS (p3 instances) and other providers at significantly lower rates than A100. For new workloads, the A100 is almost always the better choice — but for teams with existing V100-optimized code or tight budgets, it remains a viable option.

Memory

VRAM32 GB
Memory TypeHBM2
Bandwidth900 GB/s
NVLink BW300 GB/s

Compute Performance

FP3214 TFLOPS
FP16112 TFLOPS
BF16112 TFLOPS
INT8224 TOPS

Hardware

ArchitectureGV100
GenerationVolta
Process NodeTSMC 12nm
Transistors21.1B
TDP300 W
InterconnectNVLink 2.0 / PCIe 3.0
Release Year2018

Relative Performance

FP16 Compute1%
VRAM Capacity11%
Mem Bandwidth6%

Relative to highest-spec GPU in database

Limitations

No BF16 or FP8 hardware support
Older Volta architecture — significantly slower than A100/H100
PCIe 3.0 — lower host-to-device bandwidth than newer GPUs

Live Cloud PricingOn-demand hourly rates

Loading live prices…

Compare V100 32GB vs…

Use Case Guidance

Legacy ML training
HPC workloads
Budget multi-GPU setups

LLM Model Size Guidance

Max model (FP16)~16Bparameters at FP16 precision
Max model (INT8)~32Bparameters at INT8 precision
Max model (INT4)~64Bparameters at INT4/GGUF

Estimates only. Actual capacity depends on context length, KV cache, and framework overhead.

Related Guides

LLM APIs Running on This GPU Class

Providers that serve frontier LLM inference on Volta-class hardware.

Browse all 42 LLM models

Related GPUs

Frequently Asked Questions

How much VRAM does the V100 32GB have?

The V100 32GB has 32GB of HBM2 memory with 900 GB/s bandwidth. This enables running models up to approximately 64B parameters at INT4 precision, 32B at INT8, or 16B at FP16.

What is the FP16 performance of the V100 32GB?

The V100 32GB delivers 112 TFLOPS of FP16 performance and 112 TFLOPS BF16. INT8 throughput is 224 TOPS. For transformer inference, memory bandwidth (900 GB/s) is often the binding constraint rather than raw TFLOPS.

What is the V100 32GB best used for?

The V100 32GB is best suited for: Legacy ML training, HPC workloads, Budget multi-GPU setups. Previous-gen workhorse. 32GB HBM2 at low cost. Still widely available. No FP8 or BF16 hardware support.

What interconnect does the V100 32GB use?

The V100 32GB uses NVLink 2.0 / PCIe 3.0 with 300 GB/s NVLink bandwidth for multi-GPU configurations. NVLink enables near-linear tensor-parallel scaling across multiple cards for models that exceed single-card VRAM.

What LLM model sizes can the V100 32GB run?

With 32GB of HBM2, the V100 32GB can run models up to approximately 16B parameters at FP16 (2 bytes/param), 32B at INT8 (1 byte/param), or 64B at INT4/GGUF (0.5 bytes/param). These are estimates — actual capacity depends on context length, KV cache size, and framework overhead. Longer context windows require more KV cache memory, reducing the effective model size that fits.

How does the V100 32GB compare to the A100 for LLM inference?

The V100 32GB has 112 TFLOPS FP16 vs the A100 80GB's 312 TFLOPS, and 900 GB/s memory bandwidth vs the A100's 2,039 GB/s. For memory-bound autoregressive LLM inference, bandwidth is the primary determinant of tokens-per-second. The A100's higher bandwidth gives it a throughput advantage for large model inference, despite the V100 32GB's lower cost.

What is the power consumption of the V100 32GB?

The V100 32GB has a TDP (Thermal Design Power) of 300W. This is the maximum sustained power draw under full load. For data center deployments, total rack power consumption is typically 1.2–1.5× the GPU TDP when accounting for CPU, memory, networking, and cooling overhead. At 300W, the V100 32GB is in the mid-range tier — compatible with standard data center power infrastructure.

Ready to rent?

Compare V100 32GB prices across 97+ providers

Live on-demand & spot rates · monthly cost estimates · availability status

Compare rental prices