GPU Specifications / NVIDIA

NVIDIA L4

Ada Lovelace datacenter GPU with 24 GB of GDDR6 memory, 300 GB/s of bandwidth and up to 242 TFLOPS of FP16 tensor compute.

Key Specifications

Memory

24 GB GDDR6

Memory Bandwidth

300 GB/s

TDP

72 W

Architecture

Ada Lovelace

Interconnect

PCIe 4.0 · 64 GB/s

Est. On-demand Price

~$1.00/h

Hourly rates are indicative on-demand estimates; actual pricing varies by provider and commitment.

Compute Performance

PrecisionPeak throughput
FP64 (double precision)0.5 TFLOPS
FP32 (single precision)30.3 TFLOPS
FP2460.6 TFLOPS
FP16 (tensor)242 TFLOPS
INT8 (tensor)485 TFLOPS
INT4 (tensor)970 TFLOPS

Tensor figures use the vendor's peak numbers (with structured sparsity where supported).

System Requirements

Recommended CPU

AMD EPYC 7313 or Intel Xeon Silver 4310

Max VRAM per node (8 GPUs)

192 GB

System RAM (min / recommended)

64 / 128 GB

Minimum PSU

500 W

LLMs on the L4

Number of L4 GPUs needed to serve popular models at 8-bit quantization (including 20% overhead for activations and KV cache).

ModelParamsVRAM (8-bit)GPUs needed
GPT-5.6 Sol2400B2682 GB112x L4
GPT-5 Flagship2100B2347 GB98x L4
GPT-5.6 Luna1400B1565 GB66x L4
Kimi K3 (1.2T)1200B1341 GB56x L4
Kimi K2.6 (1T)1000B1118 GB47x L4
GPT-5.6 Terra800B894 GB38x L4
DeepSeek V4 Pro (671B)671B750 GB32x L4
Llama 4 Behemoth (500B)500B559 GB24x L4
Claude 5 Fable (480B)480B536 GB23x L4
GLM 5.2 (400B)400B447 GB19x L4
Grok 4.5350B391 GB17x L4
Claude 4.8 Opus (300B)300B335 GB14x L4
Muse Spark 1.1300B335 GB14x L4
Grok 4270B302 GB13x L4
Gemini 3.1 Pro250B279 GB12x L4
Qwen 3.7 Max (235B)235B263 GB11x L4
Mistral Large 3 (200B)200B224 GB10x L4
Grok 3 Mini190B212 GB9x L4
Claude 5 Sonnet (175B)175B196 GB9x L4
Gemini 3.5 Flash150B168 GB7x L4
Gemini 2.5 Flash140B156 GB7x L4
Llama 4 Maverick (128B)128B143 GB6x L4
DeepSeek V4 Flash (120B)120B134 GB6x L4
Qwen 3.6 Plus (110B)110B123 GB6x L4
Nova Premier (80B)80B89 GB4x L4
Qwen 3 Coder-Next (80B)80B89 GB4x L4
Claude 4.5 Haiku (70B)70B78 GB4x L4
Llama 3.3 Instruct (70B)70B78 GB4x L4
Mistral Medium 3.5 (70B)70B78 GB4x L4
Yi 1.5 (40B)40B45 GB2x L4
Nova Core (34B)34B38 GB2x L4
DeepSeek V3.1 (32B)32B36 GB2x L4
Gemma 3 (27B)27B30 GB2x L4
Mistral Small 4 (24B)24B27 GB2x L4
Yi 1.5 (15B)15B17 GB1x L4
Phi 4 (14B)14B16 GB1x L4
Nova Lite (12B)12B13 GB1x L4
Llama 3.2 Instruct (11B)11B12 GB1x L4
Gemma 3 (9B)9B10 GB1x L4
Yi 1.5 Lite (9B)9B10 GB1x L4
Phi 4 Mini (7B)7B8 GB1x L4
Phi 3.5 (3.8B)3.8B4 GB1x L4

Frequently Asked Questions

How much VRAM does the NVIDIA L4 have?

The NVIDIA L4 has 24 GB of GDDR6 memory with 300 GB/s of memory bandwidth.

Which LLMs can run on a single L4?

At 8-bit quantization, a single L4 (24 GB) can serve models up to roughly 15B parameters, such as Yi 1.5 (15B). Larger models require multiple GPUs or more aggressive quantization.

How much does it cost to rent a NVIDIA L4?

On-demand cloud pricing for the L4 is around $1.00/hour, i.e. about $730/month running 24/7. Actual prices vary by provider, region, and commitment.

What are the power and system requirements of the NVIDIA L4?

The L4 has a TDP of 72W. A power supply of at least 500W per GPU is recommended. Recommended host CPUs: AMD EPYC 7313 or Intel Xeon Silver 4310.