xAI

Grok 4

Frontier reasoning model from xAI

Model Summary

Family

Grok

Version

4.0

Parameters

270B (est.)

Parameter counts for closed models are estimates; vendors rarely publish exact sizes.

VRAM Requirements by Quantization

Memory needed to serve Grok 4 for inference, including a 20% overhead for activations and KV cache.

PrecisionVRAM neededSmallest single GPU that fits
INT4 (4-bit)150.87 GBNVIDIA B200 SXM (192 GB)
INT8 (8-bit)301.75 GBMulti-GPU required
FP16 (16-bit)603.50 GBMulti-GPU required
FP32 (32-bit)1206.99 GBMulti-GPU required

Recommended GPU Configurations

Cheapest on-demand configurations to serve Grok 4 at 8-bit (302 GB VRAM).

3x AMD Instinct MI250X

384 GB total VRAM · CDNA 2

~$7.50/h

19x NVIDIA T4

304 GB total VRAM · Turing · multi-node

~$9.50/h

19x NVIDIA P100 SXM2

304 GB total VRAM · Pascal · multi-node

~$11.40/h

Quick GPU Planning

Use the calculator pre-filled with this exact version to estimate memory, speed, and compute requirements in a few clicks.

Access Pre-filled Calculator

Frequently Asked Questions

How much VRAM do you need to run Grok 4?

With an estimated 270B parameters, Grok 4 needs roughly 302 GB of VRAM in 8-bit (INT8), 151 GB in 4-bit, and 603 GB in FP16, including a 20% overhead for activations and KV cache.

Which GPUs can run Grok 4?

At 8-bit quantization, the most cost-effective option is 3x AMD Instinct MI250X (384 GB combined VRAM, around $7.50/hour on-demand). Higher-end cards like the NVIDIA B200 or AMD MI355X reduce the GPU count needed.

Can Grok 4 run on a single GPU?

Only with aggressive quantization: in 4-bit, a single NVIDIA B200 SXM (192 GB) can fit it. At 8-bit or higher, you need a multi-GPU setup.

How much does it cost to serve Grok 4 in the cloud?

Renting 3x Instinct MI250X costs on the order of $7.50/hour, i.e. about $5,475/month running 24/7. Actual prices vary by provider and commitment; spot and reserved capacity can be significantly cheaper.

Other Grok Versions