NVIDIA A10
Ampere datacenter GPU with 24 GB of GDDR6 memory, 600 GB/s of bandwidth and up to 250 TFLOPS of FP16 tensor compute.
Memory
24 GB
GDDR6
Bandwidth
600
GB/s
TDP
150 W
Released
Apr 2021
Ampere
Interconnect
PCIe 4.0
64 GB/s
Est. on‑demand
~$1.00/hr
varies by provider
Compute Performance
| Precision | Peak throughput |
|---|---|
| FP64 (double precision) | 0.5 TFLOPS |
| FP32 (single precision) | 31.2 TFLOPS |
| FP24 | 62.4 TFLOPS |
| FP16 (tensor) | 250 TFLOPS |
| INT8 (tensor) | 500 TFLOPS |
| INT4 (tensor) | 1,000 TFLOPS |
Tensor figures use the vendor's peak numbers (with structured sparsity where supported).
System Requirements
Recommended CPU
AMD EPYC 7413 or Intel Xeon Gold 6338
Max VRAM per node (8 GPUs)
192 GB
System RAM (min / recommended)
128 / 256 GB
Minimum PSU
600 W
LLMs on the A10
Number of A10 GPUs needed to serve popular models at 8-bit quantization (including 20% overhead for activations and KV cache).
| Model | Params | VRAM (8-bit) | GPUs needed |
|---|---|---|---|
| GPT-5.6 Sol | 2400B | 2682 GB | 112x A10 |
| Qwen 3.8 Max (2.4T) | 2400B | 2682 GB | 112x A10 |
| GPT-5 Flagship | 2100B | 2347 GB | 98x A10 |
| GPT-5.6 Luna | 1400B | 1565 GB | 66x A10 |
| Kimi K3 (1.2T) | 1200B | 1341 GB | 56x A10 |
| Kimi K2.6 (1T) | 1000B | 1118 GB | 47x A10 |
| GPT-5.6 Terra | 800B | 894 GB | 38x A10 |
| GLM 5.3 (743B) | 743B | 830 GB | 35x A10 |
| DeepSeek V4 Pro (671B) | 671B | 750 GB | 32x A10 |
| Llama 4 Behemoth (500B) | 500B | 559 GB | 24x A10 |
| Claude 5 Fable (480B) | 480B | 536 GB | 23x A10 |
| Gemini 3.5 Pro | 400B | 447 GB | 19x A10 |
| GLM 5.2 (400B) | 400B | 447 GB | 19x A10 |
| Grok 4.5 | 350B | 391 GB | 17x A10 |
| Claude 4.8 Opus (300B) | 300B | 335 GB | 14x A10 |
| Claude 5 Opus (300B) | 300B | 335 GB | 14x A10 |
| Muse Spark 1.1 | 300B | 335 GB | 14x A10 |
| Grok 4 | 270B | 302 GB | 13x A10 |
| Gemini 3.1 Pro | 250B | 279 GB | 12x A10 |
| Qwen 3.7 Max (235B) | 235B | 263 GB | 11x A10 |
| Mistral Large 3 (200B) | 200B | 224 GB | 10x A10 |
| Grok 3 Mini | 190B | 212 GB | 9x A10 |
| Claude 5 Sonnet (175B) | 175B | 196 GB | 9x A10 |
| Gemini 3.7 Flash | 160B | 179 GB | 8x A10 |
| Gemini 3.5 Flash | 150B | 168 GB | 7x A10 |
| Gemini 2.5 Flash | 140B | 156 GB | 7x A10 |
| Llama 4 Maverick (128B) | 128B | 143 GB | 6x A10 |
| DeepSeek V4 Flash (120B) | 120B | 134 GB | 6x A10 |
| Qwen 3.6 Plus (110B) | 110B | 123 GB | 6x A10 |
| Nova Premier (80B) | 80B | 89 GB | 4x A10 |
| Qwen 3 Coder-Next (80B) | 80B | 89 GB | 4x A10 |
| Claude 4.5 Haiku (70B) | 70B | 78 GB | 4x A10 |
| Llama 3.3 Instruct (70B) | 70B | 78 GB | 4x A10 |
| Mistral Medium 3.5 (70B) | 70B | 78 GB | 4x A10 |
| Yi 1.5 (40B) | 40B | 45 GB | 2x A10 |
| Nova Core (34B) | 34B | 38 GB | 2x A10 |
| DeepSeek V3.1 (32B) | 32B | 36 GB | 2x A10 |
| Fara 1.5 27B | 27.4B | 31 GB | 2x A10 |
| Gemma 3 (27B) | 27B | 30 GB | 2x A10 |
| Qwen 3.8 (27B) | 27B | 30 GB | 2x A10 |
| Mistral Small 4 (24B) | 24B | 27 GB | 2x A10 |
| Yi 1.5 (15B) | 15B | 17 GB | 1x A10 |
| Phi 4 (14B) | 14B | 16 GB | 1x A10 |
| Gemma 4 (12B) | 12B | 13 GB | 1x A10 |
| Nova Lite (12B) | 12B | 13 GB | 1x A10 |
| Llama 3.2 Instruct (11B) | 11B | 12 GB | 1x A10 |
| Gemma 3 (9B) | 9B | 10 GB | 1x A10 |
| Yi 1.5 Lite (9B) | 9B | 10 GB | 1x A10 |
| Phi 4 Mini (7B) | 7B | 8 GB | 1x A10 |
| Mage VL | 4.7B | 5 GB | 1x A10 |
| Fara 1.5 4B | 4.5B | 5 GB | 1x A10 |
| Phi 3.5 (3.8B) | 3.8B | 4 GB | 1x A10 |
| LFM 2.5 3B | 3.1B | 3 GB | 1x A10 |
| LFM 2.5 2.6B | 2.7B | 3 GB | 1x A10 |
Frequently Asked Questions
How much VRAM does the NVIDIA A10 have?
The NVIDIA A10 has 24 GB of GDDR6 memory with 600 GB/s of memory bandwidth.
Which LLMs can run on a single A10?
At 8-bit quantization, a single A10 (24 GB) can serve models up to roughly 15B parameters, such as Yi 1.5 (15B). Larger models require multiple GPUs or more aggressive quantization.
How much does it cost to rent a NVIDIA A10?
On‑demand cloud pricing for the A10 is around $1.00/hour, i.e. about $730/month running 24/7. Actual prices vary by provider, region, and commitment.
What are the power and system requirements of the NVIDIA A10?
The A10 has a TDP of 150W. A power supply of at least 600W per GPU is recommended. Recommended host CPUs: AMD EPYC 7413 or Intel Xeon Gold 6338.
Deploy on a GPU cloud
Rent the NVIDIA A10 by the hour instead of buying hardware.