GPU Specifications / NVIDIA
NVIDIA L4
Ada Lovelace datacenter GPU with 24 GB of GDDR6 memory, 300 GB/s of bandwidth and up to 242 TFLOPS of FP16 tensor compute.
Key Specifications
Memory
24 GB GDDR6
Memory Bandwidth
300 GB/s
TDP
72 W
Architecture
Ada Lovelace
Interconnect
PCIe 4.0 · 64 GB/s
Est. On-demand Price
~$1.00/h
Hourly rates are indicative on-demand estimates; actual pricing varies by provider and commitment.
Compute Performance
| Precision | Peak throughput |
|---|---|
| FP64 (double precision) | 0.5 TFLOPS |
| FP32 (single precision) | 30.3 TFLOPS |
| FP24 | 60.6 TFLOPS |
| FP16 (tensor) | 242 TFLOPS |
| INT8 (tensor) | 485 TFLOPS |
| INT4 (tensor) | 970 TFLOPS |
Tensor figures use the vendor's peak numbers (with structured sparsity where supported).
System Requirements
Recommended CPU
AMD EPYC 7313 or Intel Xeon Silver 4310
Max VRAM per node (8 GPUs)
192 GB
System RAM (min / recommended)
64 / 128 GB
Minimum PSU
500 W
LLMs on the L4
Number of L4 GPUs needed to serve popular models at 8-bit quantization (including 20% overhead for activations and KV cache).
Frequently Asked Questions
How much VRAM does the NVIDIA L4 have?
The NVIDIA L4 has 24 GB of GDDR6 memory with 300 GB/s of memory bandwidth.
Which LLMs can run on a single L4?
At 8-bit quantization, a single L4 (24 GB) can serve models up to roughly 15B parameters, such as Yi 1.5 (15B). Larger models require multiple GPUs or more aggressive quantization.
How much does it cost to rent a NVIDIA L4?
On-demand cloud pricing for the L4 is around $1.00/hour, i.e. about $730/month running 24/7. Actual prices vary by provider, region, and commitment.
What are the power and system requirements of the NVIDIA L4?
The L4 has a TDP of 72W. A power supply of at least 500W per GPU is recommended. Recommended host CPUs: AMD EPYC 7313 or Intel Xeon Silver 4310.
Deploy on a GPU cloud
Rent the NVIDIA L4 by the hour instead of buying hardware.