GPU Specifications / NVIDIA
NVIDIA L40S
Ada Lovelace datacenter GPU with 48 GB of GDDR6 memory, 864 GB/s of bandwidth and up to 733 TFLOPS of FP16 tensor compute.
Key Specifications
Memory
48 GB GDDR6
Memory Bandwidth
864 GB/s
TDP
350 W
Architecture
Ada Lovelace
Interconnect
PCIe 4.0 · 64 GB/s
Est. On-demand Price
~$3.50/h
Hourly rates are indicative on-demand estimates; actual pricing varies by provider and commitment.
Compute Performance
| Precision | Peak throughput |
|---|---|
| FP64 (double precision) | 1.4 TFLOPS |
| FP32 (single precision) | 91.6 TFLOPS |
| FP24 | 183.2 TFLOPS |
| FP16 (tensor) | 733 TFLOPS |
| INT8 (tensor) | 1,466 TFLOPS |
| INT4 (tensor) | 2,932 TFLOPS |
Tensor figures use the vendor's peak numbers (with structured sparsity where supported).
System Requirements
Recommended CPU
AMD EPYC 7443 or Intel Xeon Gold 6348
Max VRAM per node (8 GPUs)
384 GB
System RAM (min / recommended)
128 / 256 GB
Minimum PSU
800 W
LLMs on the L40S
Number of L40S GPUs needed to serve popular models at 8-bit quantization (including 20% overhead for activations and KV cache).
| Model | Params | VRAM (8-bit) | GPUs needed |
|---|---|---|---|
| GPT-5.6 Sol | 2400B | 2682 GB | 56x L40S |
| GPT-5 Flagship | 2100B | 2347 GB | 49x L40S |
| GPT-5.6 Luna | 1400B | 1565 GB | 33x L40S |
| Kimi K3 (1.2T) | 1200B | 1341 GB | 28x L40S |
| Kimi K2.6 (1T) | 1000B | 1118 GB | 24x L40S |
| GPT-5.6 Terra | 800B | 894 GB | 19x L40S |
| DeepSeek V4 Pro (671B) | 671B | 750 GB | 16x L40S |
| Llama 4 Behemoth (500B) | 500B | 559 GB | 12x L40S |
| Claude 5 Fable (480B) | 480B | 536 GB | 12x L40S |
| GLM 5.2 (400B) | 400B | 447 GB | 10x L40S |
| Grok 4.5 | 350B | 391 GB | 9x L40S |
| Claude 4.8 Opus (300B) | 300B | 335 GB | 7x L40S |
| Muse Spark 1.1 | 300B | 335 GB | 7x L40S |
| Grok 4 | 270B | 302 GB | 7x L40S |
| Gemini 3.1 Pro | 250B | 279 GB | 6x L40S |
| Qwen 3.7 Max (235B) | 235B | 263 GB | 6x L40S |
| Mistral Large 3 (200B) | 200B | 224 GB | 5x L40S |
| Grok 3 Mini | 190B | 212 GB | 5x L40S |
| Claude 5 Sonnet (175B) | 175B | 196 GB | 5x L40S |
| Gemini 3.5 Flash | 150B | 168 GB | 4x L40S |
| Gemini 2.5 Flash | 140B | 156 GB | 4x L40S |
| Llama 4 Maverick (128B) | 128B | 143 GB | 3x L40S |
| DeepSeek V4 Flash (120B) | 120B | 134 GB | 3x L40S |
| Qwen 3.6 Plus (110B) | 110B | 123 GB | 3x L40S |
| Nova Premier (80B) | 80B | 89 GB | 2x L40S |
| Qwen 3 Coder-Next (80B) | 80B | 89 GB | 2x L40S |
| Claude 4.5 Haiku (70B) | 70B | 78 GB | 2x L40S |
| Llama 3.3 Instruct (70B) | 70B | 78 GB | 2x L40S |
| Mistral Medium 3.5 (70B) | 70B | 78 GB | 2x L40S |
| Yi 1.5 (40B) | 40B | 45 GB | 1x L40S |
| Nova Core (34B) | 34B | 38 GB | 1x L40S |
| DeepSeek V3.1 (32B) | 32B | 36 GB | 1x L40S |
| Gemma 3 (27B) | 27B | 30 GB | 1x L40S |
| Mistral Small 4 (24B) | 24B | 27 GB | 1x L40S |
| Yi 1.5 (15B) | 15B | 17 GB | 1x L40S |
| Phi 4 (14B) | 14B | 16 GB | 1x L40S |
| Nova Lite (12B) | 12B | 13 GB | 1x L40S |
| Llama 3.2 Instruct (11B) | 11B | 12 GB | 1x L40S |
| Gemma 3 (9B) | 9B | 10 GB | 1x L40S |
| Yi 1.5 Lite (9B) | 9B | 10 GB | 1x L40S |
| Phi 4 Mini (7B) | 7B | 8 GB | 1x L40S |
| Phi 3.5 (3.8B) | 3.8B | 4 GB | 1x L40S |
Frequently Asked Questions
How much VRAM does the NVIDIA L40S have?
The NVIDIA L40S has 48 GB of GDDR6 memory with 864 GB/s of memory bandwidth.
Which LLMs can run on a single L40S?
At 8-bit quantization, a single L40S (48 GB) can serve models up to roughly 40B parameters, such as Yi 1.5 (40B). Larger models require multiple GPUs or more aggressive quantization.
How much does it cost to rent a NVIDIA L40S?
On-demand cloud pricing for the L40S is around $3.50/hour, i.e. about $2,555/month running 24/7. Actual prices vary by provider, region, and commitment.
What are the power and system requirements of the NVIDIA L40S?
The L40S has a TDP of 350W. A power supply of at least 800W per GPU is recommended. Recommended host CPUs: AMD EPYC 7443 or Intel Xeon Gold 6348.
Deploy on a GPU cloud
Rent the NVIDIA L40S by the hour instead of buying hardware.