NVIDIA A30
Ampere datacenter GPU with 24 GB of HBM2 memory, 933 GB/s of bandwidth and up to 330 TFLOPS of FP16 tensor compute.
Memory
24 GB
HBM2
Bandwidth
933
GB/s
TDP
165 W
Released
Apr 2021
Ampere
Interconnect
NVLink 3.0
200 GB/s
Est. on‑demand
~$1.20/hr
varies by provider
Compute Performance
| Precision | Peak throughput |
|---|---|
| FP64 (double precision) | 5.2 TFLOPS |
| FP32 (single precision) | 10.3 TFLOPS |
| FP24 | 20.6 TFLOPS |
| FP16 (tensor) | 330 TFLOPS |
| INT8 (tensor) | 661 TFLOPS |
| INT4 (tensor) | 1,321 TFLOPS |
Tensor figures use the vendor's peak numbers (with structured sparsity where supported).
System Requirements
Recommended CPU
AMD EPYC 7413 or Intel Xeon Gold 6338
Max VRAM per node (8 GPUs)
192 GB
System RAM (min / recommended)
128 / 256 GB
Minimum PSU
600 W
LLMs on the A30
Number of A30 GPUs needed to serve popular models at 8-bit quantization (including 20% overhead for activations and KV cache).
| Model | Params | VRAM (8-bit) | GPUs needed |
|---|---|---|---|
| GPT-5.6 Sol | 2400B | 2682 GB | 112x A30 |
| Qwen 3.8 Max (2.4T) | 2400B | 2682 GB | 112x A30 |
| GPT-5 Flagship | 2100B | 2347 GB | 98x A30 |
| GPT-5.6 Luna | 1400B | 1565 GB | 66x A30 |
| Kimi K3 (1.2T) | 1200B | 1341 GB | 56x A30 |
| Kimi K2.6 (1T) | 1000B | 1118 GB | 47x A30 |
| GPT-5.6 Terra | 800B | 894 GB | 38x A30 |
| GLM 5.3 (743B) | 743B | 830 GB | 35x A30 |
| DeepSeek V4 Pro (671B) | 671B | 750 GB | 32x A30 |
| Llama 4 Behemoth (500B) | 500B | 559 GB | 24x A30 |
| Claude 5 Fable (480B) | 480B | 536 GB | 23x A30 |
| Gemini 3.5 Pro | 400B | 447 GB | 19x A30 |
| GLM 5.2 (400B) | 400B | 447 GB | 19x A30 |
| Grok 4.5 | 350B | 391 GB | 17x A30 |
| Claude 4.8 Opus (300B) | 300B | 335 GB | 14x A30 |
| Claude 5 Opus (300B) | 300B | 335 GB | 14x A30 |
| Muse Spark 1.1 | 300B | 335 GB | 14x A30 |
| Grok 4 | 270B | 302 GB | 13x A30 |
| Gemini 3.1 Pro | 250B | 279 GB | 12x A30 |
| Qwen 3.7 Max (235B) | 235B | 263 GB | 11x A30 |
| Mistral Large 3 (200B) | 200B | 224 GB | 10x A30 |
| Grok 3 Mini | 190B | 212 GB | 9x A30 |
| Claude 5 Sonnet (175B) | 175B | 196 GB | 9x A30 |
| Gemini 3.7 Flash | 160B | 179 GB | 8x A30 |
| Gemini 3.5 Flash | 150B | 168 GB | 7x A30 |
| Gemini 2.5 Flash | 140B | 156 GB | 7x A30 |
| Llama 4 Maverick (128B) | 128B | 143 GB | 6x A30 |
| DeepSeek V4 Flash (120B) | 120B | 134 GB | 6x A30 |
| Qwen 3.6 Plus (110B) | 110B | 123 GB | 6x A30 |
| Nova Premier (80B) | 80B | 89 GB | 4x A30 |
| Qwen 3 Coder-Next (80B) | 80B | 89 GB | 4x A30 |
| Claude 4.5 Haiku (70B) | 70B | 78 GB | 4x A30 |
| Llama 3.3 Instruct (70B) | 70B | 78 GB | 4x A30 |
| Mistral Medium 3.5 (70B) | 70B | 78 GB | 4x A30 |
| Yi 1.5 (40B) | 40B | 45 GB | 2x A30 |
| Nova Core (34B) | 34B | 38 GB | 2x A30 |
| DeepSeek V3.1 (32B) | 32B | 36 GB | 2x A30 |
| Fara 1.5 27B | 27.4B | 31 GB | 2x A30 |
| Gemma 3 (27B) | 27B | 30 GB | 2x A30 |
| Qwen 3.8 (27B) | 27B | 30 GB | 2x A30 |
| Mistral Small 4 (24B) | 24B | 27 GB | 2x A30 |
| Yi 1.5 (15B) | 15B | 17 GB | 1x A30 |
| Phi 4 (14B) | 14B | 16 GB | 1x A30 |
| Gemma 4 (12B) | 12B | 13 GB | 1x A30 |
| Nova Lite (12B) | 12B | 13 GB | 1x A30 |
| Llama 3.2 Instruct (11B) | 11B | 12 GB | 1x A30 |
| Gemma 3 (9B) | 9B | 10 GB | 1x A30 |
| Yi 1.5 Lite (9B) | 9B | 10 GB | 1x A30 |
| Phi 4 Mini (7B) | 7B | 8 GB | 1x A30 |
| Mage VL | 4.7B | 5 GB | 1x A30 |
| Fara 1.5 4B | 4.5B | 5 GB | 1x A30 |
| Phi 3.5 (3.8B) | 3.8B | 4 GB | 1x A30 |
| LFM 2.5 3B | 3.1B | 3 GB | 1x A30 |
| LFM 2.5 2.6B | 2.7B | 3 GB | 1x A30 |
Frequently Asked Questions
How much VRAM does the NVIDIA A30 have?
The NVIDIA A30 has 24 GB of HBM2 memory with 933 GB/s of memory bandwidth.
Which LLMs can run on a single A30?
At 8-bit quantization, a single A30 (24 GB) can serve models up to roughly 15B parameters, such as Yi 1.5 (15B). Larger models require multiple GPUs or more aggressive quantization.
How much does it cost to rent a NVIDIA A30?
On‑demand cloud pricing for the A30 is around $1.20/hour, i.e. about $876/month running 24/7. Actual prices vary by provider, region, and commitment.
What are the power and system requirements of the NVIDIA A30?
The A30 has a TDP of 165W. A power supply of at least 600W per GPU is recommended. Recommended host CPUs: AMD EPYC 7413 or Intel Xeon Gold 6338.
Deploy on a GPU cloud
Rent the NVIDIA A30 by the hour instead of buying hardware.