NVIDIA A40
Ampere datacenter GPU with 48 GB of GDDR6 memory, 696 GB/s of bandwidth and up to 299.4 TFLOPS of FP16 tensor compute.
Memory
48 GB
GDDR6
Bandwidth
696
GB/s
TDP
300 W
Released
Oct 2020
Ampere
Interconnect
NVLink 3.0
112.5 GB/s
Est. on‑demand
~$1.80/hr
varies by provider
Compute Performance
| Precision | Peak throughput |
|---|---|
| FP64 (double precision) | 0.6 TFLOPS |
| FP32 (single precision) | 37.4 TFLOPS |
| FP24 | 74.8 TFLOPS |
| FP16 (tensor) | 299.4 TFLOPS |
| INT8 (tensor) | 598.7 TFLOPS |
| INT4 (tensor) | 1,197.4 TFLOPS |
Tensor figures use the vendor's peak numbers (with structured sparsity where supported).
System Requirements
Recommended CPU
AMD EPYC 7543 or Intel Xeon Gold 6342
Max VRAM per node (8 GPUs)
384 GB
System RAM (min / recommended)
256 / 512 GB
Minimum PSU
850 W
LLMs on the A40
Number of A40 GPUs needed to serve popular models at 8-bit quantization (including 20% overhead for activations and KV cache).
| Model | Params | VRAM (8-bit) | GPUs needed |
|---|---|---|---|
| GPT-5.6 Sol | 2400B | 2682 GB | 56x A40 |
| Qwen 3.8 Max (2.4T) | 2400B | 2682 GB | 56x A40 |
| GPT-5 Flagship | 2100B | 2347 GB | 49x A40 |
| GPT-5.6 Luna | 1400B | 1565 GB | 33x A40 |
| Kimi K3 (1.2T) | 1200B | 1341 GB | 28x A40 |
| Kimi K2.6 (1T) | 1000B | 1118 GB | 24x A40 |
| GPT-5.6 Terra | 800B | 894 GB | 19x A40 |
| GLM 5.3 (743B) | 743B | 830 GB | 18x A40 |
| DeepSeek V4 Pro (671B) | 671B | 750 GB | 16x A40 |
| Llama 4 Behemoth (500B) | 500B | 559 GB | 12x A40 |
| Claude 5 Fable (480B) | 480B | 536 GB | 12x A40 |
| Gemini 3.5 Pro | 400B | 447 GB | 10x A40 |
| GLM 5.2 (400B) | 400B | 447 GB | 10x A40 |
| Grok 4.5 | 350B | 391 GB | 9x A40 |
| Claude 4.8 Opus (300B) | 300B | 335 GB | 7x A40 |
| Claude 5 Opus (300B) | 300B | 335 GB | 7x A40 |
| Muse Spark 1.1 | 300B | 335 GB | 7x A40 |
| Grok 4 | 270B | 302 GB | 7x A40 |
| Gemini 3.1 Pro | 250B | 279 GB | 6x A40 |
| Qwen 3.7 Max (235B) | 235B | 263 GB | 6x A40 |
| Mistral Large 3 (200B) | 200B | 224 GB | 5x A40 |
| Grok 3 Mini | 190B | 212 GB | 5x A40 |
| Claude 5 Sonnet (175B) | 175B | 196 GB | 5x A40 |
| Gemini 3.7 Flash | 160B | 179 GB | 4x A40 |
| Gemini 3.5 Flash | 150B | 168 GB | 4x A40 |
| Gemini 2.5 Flash | 140B | 156 GB | 4x A40 |
| Llama 4 Maverick (128B) | 128B | 143 GB | 3x A40 |
| DeepSeek V4 Flash (120B) | 120B | 134 GB | 3x A40 |
| Qwen 3.6 Plus (110B) | 110B | 123 GB | 3x A40 |
| Nova Premier (80B) | 80B | 89 GB | 2x A40 |
| Qwen 3 Coder-Next (80B) | 80B | 89 GB | 2x A40 |
| Claude 4.5 Haiku (70B) | 70B | 78 GB | 2x A40 |
| Llama 3.3 Instruct (70B) | 70B | 78 GB | 2x A40 |
| Mistral Medium 3.5 (70B) | 70B | 78 GB | 2x A40 |
| Yi 1.5 (40B) | 40B | 45 GB | 1x A40 |
| Nova Core (34B) | 34B | 38 GB | 1x A40 |
| DeepSeek V3.1 (32B) | 32B | 36 GB | 1x A40 |
| Fara 1.5 27B | 27.4B | 31 GB | 1x A40 |
| Gemma 3 (27B) | 27B | 30 GB | 1x A40 |
| Qwen 3.8 (27B) | 27B | 30 GB | 1x A40 |
| Mistral Small 4 (24B) | 24B | 27 GB | 1x A40 |
| Yi 1.5 (15B) | 15B | 17 GB | 1x A40 |
| Phi 4 (14B) | 14B | 16 GB | 1x A40 |
| Gemma 4 (12B) | 12B | 13 GB | 1x A40 |
| Nova Lite (12B) | 12B | 13 GB | 1x A40 |
| Llama 3.2 Instruct (11B) | 11B | 12 GB | 1x A40 |
| Gemma 3 (9B) | 9B | 10 GB | 1x A40 |
| Yi 1.5 Lite (9B) | 9B | 10 GB | 1x A40 |
| Phi 4 Mini (7B) | 7B | 8 GB | 1x A40 |
| Mage VL | 4.7B | 5 GB | 1x A40 |
| Fara 1.5 4B | 4.5B | 5 GB | 1x A40 |
| Phi 3.5 (3.8B) | 3.8B | 4 GB | 1x A40 |
| LFM 2.5 3B | 3.1B | 3 GB | 1x A40 |
| LFM 2.5 2.6B | 2.7B | 3 GB | 1x A40 |
Frequently Asked Questions
How much VRAM does the NVIDIA A40 have?
The NVIDIA A40 has 48 GB of GDDR6 memory with 696 GB/s of memory bandwidth.
Which LLMs can run on a single A40?
At 8-bit quantization, a single A40 (48 GB) can serve models up to roughly 40B parameters, such as Yi 1.5 (40B). Larger models require multiple GPUs or more aggressive quantization.
How much does it cost to rent a NVIDIA A40?
On‑demand cloud pricing for the A40 is around $1.80/hour, i.e. about $1,314/month running 24/7. Actual prices vary by provider, region, and commitment.
What are the power and system requirements of the NVIDIA A40?
The A40 has a TDP of 300W. A power supply of at least 850W per GPU is recommended. Recommended host CPUs: AMD EPYC 7543 or Intel Xeon Gold 6342.
Rent the A40 in the cloud
Compare live A40 rental prices across GPU cloud providers.
Deploy on a GPU cloud
Rent the NVIDIA A40 by the hour instead of buying hardware.