Alibaba
Qwen 3.8 Max (2.4T)
Alibaba's 2.4T MoE flagship (95B active), multimodal, 1M-token context
Model Summary
Family
Qwen
Version
3.8
Parameters
2400B (est.)
Parameter counts for closed models are estimates; vendors rarely publish exact sizes.
VRAM Requirements by Quantization
Memory needed to serve Qwen 3.8 Max (2.4T) for inference, including a 20% overhead for activations and KV cache.
| Precision | VRAM needed | Smallest single GPU that fits |
|---|---|---|
| INT4 (4-bit) | 1341.10 GB | Multi-GPU required |
| INT8 (8-bit) | 2682.21 GB | Multi-GPU required |
| FP16 (16-bit) | 5364.42 GB | Multi-GPU required |
| FP32 (32-bit) | 10728.84 GB | Multi-GPU required |
Recommended GPU Configurations
Cheapest on-demand configurations to serve Qwen 3.8 Max (2.4T) at 8-bit (2682 GB VRAM).
21x AMD Instinct MI250X
2688 GB total VRAM · CDNA 2 · multi-node
~$52.50/h
168x NVIDIA T4
2688 GB total VRAM · Turing · multi-node
~$84.00/h
14x AMD Instinct MI300X
2688 GB total VRAM · CDNA 3 · multi-node
~$84.00/h
Quick GPU Planning
Use the calculator pre-filled with this exact version to estimate memory, speed, and compute requirements in a few clicks.
Access Pre-filled CalculatorFrequently Asked Questions
How much VRAM do you need to run Qwen 3.8 Max (2.4T)?
With an estimated 2400B parameters, Qwen 3.8 Max (2.4T) needs roughly 2682 GB of VRAM in 8-bit (INT8), 1341 GB in 4-bit, and 5364 GB in FP16, including a 20% overhead for activations and KV cache.
Which GPUs can run Qwen 3.8 Max (2.4T)?
At 8-bit quantization, the most cost-effective option is 21x AMD Instinct MI250X (2688 GB combined VRAM, around $52.50/hour on-demand). Higher-end cards like the NVIDIA B200 or AMD MI355X reduce the GPU count needed.
Can Qwen 3.8 Max (2.4T) run on a single GPU?
No. Even in 4-bit, Qwen 3.8 Max (2.4T) exceeds the memory of any single current GPU, so a multi-GPU cluster is required.
How much does it cost to serve Qwen 3.8 Max (2.4T) in the cloud?
Renting 21x Instinct MI250X costs on the order of $52.50/hour, i.e. about $38,325/month running 24/7. Actual prices vary by provider and commitment; spot and reserved capacity can be significantly cheaper.
Deploy on a GPU cloud
Rent 21x Instinct MI250X by the hour instead of buying hardware.