NVIDIA B200 SXM
Blackwell datacenter GPU with 192 GB of HBM3e memory, 8000 GB/s of bandwidth and up to 4,500 TFLOPS of FP16 tensor compute.
Memory
192 GB
HBM3e
Bandwidth
8,000
GB/s
TDP
1000 W
Released
Nov 2024
Blackwell
Interconnect
NVLink 5.0
1800 GB/s
Est. on‑demand
~$14.00/hr
varies by provider
Compute Performance
| Precision | Peak throughput |
|---|---|
| FP64 (double precision) | 40 TFLOPS |
| FP32 (single precision) | 80 TFLOPS |
| FP24 | 160 TFLOPS |
| FP16 (tensor) | 4,500 TFLOPS |
| INT8 (tensor) | 9,000 TFLOPS |
| INT4 (tensor) | 18,000 TFLOPS |
Tensor figures use the vendor's peak numbers (with structured sparsity where supported).
System Requirements
Recommended CPU
Intel Xeon Platinum 8570 or NVIDIA Grace (GB200 NVL72)
Max VRAM per node (8 GPUs)
1,536 GB
System RAM (min / recommended)
1024 / 2048 GB
Minimum PSU
2000 W
LLMs on the B200 SXM
Number of B200 SXM GPUs needed to serve popular models at 8-bit quantization (including 20% overhead for activations and KV cache).
| Model | Params | VRAM (8-bit) | GPUs needed |
|---|---|---|---|
| GPT-5.6 Sol | 2400B | 2682 GB | 14x B200 SXM |
| Qwen 3.8 Max (2.4T) | 2400B | 2682 GB | 14x B200 SXM |
| GPT-5 Flagship | 2100B | 2347 GB | 13x B200 SXM |
| GPT-5.6 Luna | 1400B | 1565 GB | 9x B200 SXM |
| Kimi K3 (1.2T) | 1200B | 1341 GB | 7x B200 SXM |
| Kimi K2.6 (1T) | 1000B | 1118 GB | 6x B200 SXM |
| GPT-5.6 Terra | 800B | 894 GB | 5x B200 SXM |
| GLM 5.3 (743B) | 743B | 830 GB | 5x B200 SXM |
| DeepSeek V4 Pro (671B) | 671B | 750 GB | 4x B200 SXM |
| Llama 4 Behemoth (500B) | 500B | 559 GB | 3x B200 SXM |
| Claude 5 Fable (480B) | 480B | 536 GB | 3x B200 SXM |
| Gemini 3.5 Pro | 400B | 447 GB | 3x B200 SXM |
| GLM 5.2 (400B) | 400B | 447 GB | 3x B200 SXM |
| Grok 4.5 | 350B | 391 GB | 3x B200 SXM |
| Claude 4.8 Opus (300B) | 300B | 335 GB | 2x B200 SXM |
| Claude 5 Opus (300B) | 300B | 335 GB | 2x B200 SXM |
| Muse Spark 1.1 | 300B | 335 GB | 2x B200 SXM |
| Grok 4 | 270B | 302 GB | 2x B200 SXM |
| Gemini 3.1 Pro | 250B | 279 GB | 2x B200 SXM |
| Qwen 3.7 Max (235B) | 235B | 263 GB | 2x B200 SXM |
| Mistral Large 3 (200B) | 200B | 224 GB | 2x B200 SXM |
| Grok 3 Mini | 190B | 212 GB | 2x B200 SXM |
| Claude 5 Sonnet (175B) | 175B | 196 GB | 2x B200 SXM |
| Gemini 3.7 Flash | 160B | 179 GB | 1x B200 SXM |
| Gemini 3.5 Flash | 150B | 168 GB | 1x B200 SXM |
| Gemini 2.5 Flash | 140B | 156 GB | 1x B200 SXM |
| Llama 4 Maverick (128B) | 128B | 143 GB | 1x B200 SXM |
| DeepSeek V4 Flash (120B) | 120B | 134 GB | 1x B200 SXM |
| Qwen 3.6 Plus (110B) | 110B | 123 GB | 1x B200 SXM |
| Nova Premier (80B) | 80B | 89 GB | 1x B200 SXM |
| Qwen 3 Coder-Next (80B) | 80B | 89 GB | 1x B200 SXM |
| Claude 4.5 Haiku (70B) | 70B | 78 GB | 1x B200 SXM |
| Llama 3.3 Instruct (70B) | 70B | 78 GB | 1x B200 SXM |
| Mistral Medium 3.5 (70B) | 70B | 78 GB | 1x B200 SXM |
| Yi 1.5 (40B) | 40B | 45 GB | 1x B200 SXM |
| Nova Core (34B) | 34B | 38 GB | 1x B200 SXM |
| DeepSeek V3.1 (32B) | 32B | 36 GB | 1x B200 SXM |
| Fara 1.5 27B | 27.4B | 31 GB | 1x B200 SXM |
| Gemma 3 (27B) | 27B | 30 GB | 1x B200 SXM |
| Qwen 3.8 (27B) | 27B | 30 GB | 1x B200 SXM |
| Mistral Small 4 (24B) | 24B | 27 GB | 1x B200 SXM |
| Yi 1.5 (15B) | 15B | 17 GB | 1x B200 SXM |
| Phi 4 (14B) | 14B | 16 GB | 1x B200 SXM |
| Gemma 4 (12B) | 12B | 13 GB | 1x B200 SXM |
| Nova Lite (12B) | 12B | 13 GB | 1x B200 SXM |
| Llama 3.2 Instruct (11B) | 11B | 12 GB | 1x B200 SXM |
| Gemma 3 (9B) | 9B | 10 GB | 1x B200 SXM |
| Yi 1.5 Lite (9B) | 9B | 10 GB | 1x B200 SXM |
| Phi 4 Mini (7B) | 7B | 8 GB | 1x B200 SXM |
| Mage VL | 4.7B | 5 GB | 1x B200 SXM |
| Fara 1.5 4B | 4.5B | 5 GB | 1x B200 SXM |
| Phi 3.5 (3.8B) | 3.8B | 4 GB | 1x B200 SXM |
| LFM 2.5 3B | 3.1B | 3 GB | 1x B200 SXM |
| LFM 2.5 2.6B | 2.7B | 3 GB | 1x B200 SXM |
Frequently Asked Questions
How much VRAM does the NVIDIA B200 SXM have?
The NVIDIA B200 SXM has 192 GB of HBM3e memory with 8000 GB/s of memory bandwidth.
Which LLMs can run on a single B200 SXM?
At 8-bit quantization, a single B200 SXM (192 GB) can serve models up to roughly 160B parameters, such as Gemini 3.7 Flash. Larger models require multiple GPUs or more aggressive quantization.
How much does it cost to rent a NVIDIA B200 SXM?
On‑demand cloud pricing for the B200 SXM is around $14.00/hour, i.e. about $10,220/month running 24/7. Actual prices vary by provider, region, and commitment.
What are the power and system requirements of the NVIDIA B200 SXM?
The B200 SXM has a TDP of 1000W. A power supply of at least 2000W per GPU is recommended. Recommended host CPUs: Intel Xeon Platinum 8570 or NVIDIA Grace (GB200 NVL72).
Rent the B200 SXM in the cloud
Compare live B200 SXM rental prices across GPU cloud providers.
Deploy on a GPU cloud
Rent the NVIDIA B200 SXM by the hour instead of buying hardware.