GPU Specifications / NVIDIA
NVIDIA P100 SXM2
Pascal datacenter GPU with 16 GB of HBM2 memory, 732 GB/s of bandwidth and up to 21.2 TFLOPS of FP16 tensor compute.
Key Specifications
Memory
16 GB HBM2
Memory Bandwidth
732 GB/s
TDP
300 W
Architecture
Pascal
Interconnect
NVLink 1.0 · 160 GB/s
Est. On-demand Price
~$0.60/h
Hourly rates are indicative on-demand estimates; actual pricing varies by provider and commitment.
Compute Performance
| Precision | Peak throughput |
|---|---|
| FP64 (double precision) | 5.3 TFLOPS |
| FP32 (single precision) | 10.6 TFLOPS |
| FP24 | 21.2 TFLOPS |
| FP16 (tensor) | 21.2 TFLOPS |
| INT8 (tensor) | 21.2 TFLOPS |
| INT4 (tensor) | 21.2 TFLOPS |
Tensor figures use the vendor's peak numbers (with structured sparsity where supported).
System Requirements
Recommended CPU
Intel Xeon E5-2698 v4
Max VRAM per node (8 GPUs)
128 GB
System RAM (min / recommended)
128 / 256 GB
Minimum PSU
800 W
LLMs on the P100 SXM2
Number of P100 SXM2 GPUs needed to serve popular models at 8-bit quantization (including 20% overhead for activations and KV cache).
| Model | Params | VRAM (8-bit) | GPUs needed |
|---|---|---|---|
| GPT-5.6 Sol | 2400B | 2682 GB | 168x P100 SXM2 |
| GPT-5 Flagship | 2100B | 2347 GB | 147x P100 SXM2 |
| GPT-5.6 Luna | 1400B | 1565 GB | 98x P100 SXM2 |
| Kimi K3 (1.2T) | 1200B | 1341 GB | 84x P100 SXM2 |
| Kimi K2.6 (1T) | 1000B | 1118 GB | 70x P100 SXM2 |
| GPT-5.6 Terra | 800B | 894 GB | 56x P100 SXM2 |
| DeepSeek V4 Pro (671B) | 671B | 750 GB | 47x P100 SXM2 |
| Llama 4 Behemoth (500B) | 500B | 559 GB | 35x P100 SXM2 |
| Claude 5 Fable (480B) | 480B | 536 GB | 34x P100 SXM2 |
| GLM 5.2 (400B) | 400B | 447 GB | 28x P100 SXM2 |
| Grok 4.5 | 350B | 391 GB | 25x P100 SXM2 |
| Claude 4.8 Opus (300B) | 300B | 335 GB | 21x P100 SXM2 |
| Muse Spark 1.1 | 300B | 335 GB | 21x P100 SXM2 |
| Grok 4 | 270B | 302 GB | 19x P100 SXM2 |
| Gemini 3.1 Pro | 250B | 279 GB | 18x P100 SXM2 |
| Qwen 3.7 Max (235B) | 235B | 263 GB | 17x P100 SXM2 |
| Mistral Large 3 (200B) | 200B | 224 GB | 14x P100 SXM2 |
| Grok 3 Mini | 190B | 212 GB | 14x P100 SXM2 |
| Claude 5 Sonnet (175B) | 175B | 196 GB | 13x P100 SXM2 |
| Gemini 3.5 Flash | 150B | 168 GB | 11x P100 SXM2 |
| Gemini 2.5 Flash | 140B | 156 GB | 10x P100 SXM2 |
| Llama 4 Maverick (128B) | 128B | 143 GB | 9x P100 SXM2 |
| DeepSeek V4 Flash (120B) | 120B | 134 GB | 9x P100 SXM2 |
| Qwen 3.6 Plus (110B) | 110B | 123 GB | 8x P100 SXM2 |
| Nova Premier (80B) | 80B | 89 GB | 6x P100 SXM2 |
| Qwen 3 Coder-Next (80B) | 80B | 89 GB | 6x P100 SXM2 |
| Claude 4.5 Haiku (70B) | 70B | 78 GB | 5x P100 SXM2 |
| Llama 3.3 Instruct (70B) | 70B | 78 GB | 5x P100 SXM2 |
| Mistral Medium 3.5 (70B) | 70B | 78 GB | 5x P100 SXM2 |
| Yi 1.5 (40B) | 40B | 45 GB | 3x P100 SXM2 |
| Nova Core (34B) | 34B | 38 GB | 3x P100 SXM2 |
| DeepSeek V3.1 (32B) | 32B | 36 GB | 3x P100 SXM2 |
| Gemma 3 (27B) | 27B | 30 GB | 2x P100 SXM2 |
| Mistral Small 4 (24B) | 24B | 27 GB | 2x P100 SXM2 |
| Yi 1.5 (15B) | 15B | 17 GB | 2x P100 SXM2 |
| Phi 4 (14B) | 14B | 16 GB | 1x P100 SXM2 |
| Nova Lite (12B) | 12B | 13 GB | 1x P100 SXM2 |
| Llama 3.2 Instruct (11B) | 11B | 12 GB | 1x P100 SXM2 |
| Gemma 3 (9B) | 9B | 10 GB | 1x P100 SXM2 |
| Yi 1.5 Lite (9B) | 9B | 10 GB | 1x P100 SXM2 |
| Phi 4 Mini (7B) | 7B | 8 GB | 1x P100 SXM2 |
| Phi 3.5 (3.8B) | 3.8B | 4 GB | 1x P100 SXM2 |
Frequently Asked Questions
How much VRAM does the NVIDIA P100 SXM2 have?
The NVIDIA P100 SXM2 has 16 GB of HBM2 memory with 732 GB/s of memory bandwidth.
Which LLMs can run on a single P100 SXM2?
At 8-bit quantization, a single P100 SXM2 (16 GB) can serve models up to roughly 14B parameters, such as Phi 4 (14B). Larger models require multiple GPUs or more aggressive quantization.
How much does it cost to rent a NVIDIA P100 SXM2?
On-demand cloud pricing for the P100 SXM2 is around $0.60/hour, i.e. about $438/month running 24/7. Actual prices vary by provider, region, and commitment.
What are the power and system requirements of the NVIDIA P100 SXM2?
The P100 SXM2 has a TDP of 300W. A power supply of at least 800W per GPU is recommended. Recommended host CPUs: Intel Xeon E5-2698 v4.
Deploy on a GPU cloud
Rent the NVIDIA P100 SXM2 by the hour instead of buying hardware.