DeepSeek AI
DeepSeek V4 Flash Vision Exp
DeepSeek AI model discovered on huggingface
Model Summary
Family
DeepSeek
Version
4
Parameters
304.6B (est.)
Parameter counts for closed models are estimates; vendors rarely publish exact sizes.
VRAM Requirements by Quantization
Memory needed to serve DeepSeek V4 Flash Vision Exp for inference, including a 20% overhead for activations and KV cache.
| Precision | VRAM needed | Smallest single GPU that fits |
|---|---|---|
| INT4 (4-bit) | 170.21 GB | NVIDIA B200 SXM (192 GB) |
| INT8 (8-bit) | 340.42 GB | Multi-GPU required |
| FP16 (16-bit) | 680.83 GB | Multi-GPU required |
| FP32 (32-bit) | 1361.67 GB | Multi-GPU required |
Recommended GPU Configurations
Cheapest on-demand configurations to serve DeepSeek V4 Flash Vision Exp at 8-bit (340 GB VRAM).
3x AMD Instinct MI250X
384 GB total VRAM · CDNA 2
~$7.50/h
22x NVIDIA T4
352 GB total VRAM · Turing · multi-node
~$11.00/h
2x AMD Instinct MI300X
384 GB total VRAM · CDNA 3
~$12.00/h
Quick GPU Planning
Use the calculator pre-filled with this exact version to estimate memory, speed, and compute requirements in a few clicks.
Access Pre-filled CalculatorFrequently Asked Questions
How much VRAM do you need to run DeepSeek V4 Flash Vision Exp?
With an estimated 304.6B parameters, DeepSeek V4 Flash Vision Exp needs roughly 340 GB of VRAM in 8-bit (INT8), 170 GB in 4-bit, and 681 GB in FP16, including a 20% overhead for activations and KV cache.
Which GPUs can run DeepSeek V4 Flash Vision Exp?
At 8-bit quantization, the most cost-effective option is 3x AMD Instinct MI250X (384 GB combined VRAM, around $7.50/hour on-demand). Higher-end cards like the NVIDIA B200 or AMD MI355X reduce the GPU count needed.
Can DeepSeek V4 Flash Vision Exp run on a single GPU?
Only with aggressive quantization: in 4-bit, a single NVIDIA B200 SXM (192 GB) can fit it. At 8-bit or higher, you need a multi-GPU setup.
How much does it cost to serve DeepSeek V4 Flash Vision Exp in the cloud?
Renting 3x Instinct MI250X costs on the order of $7.50/hour, i.e. about $5,475/month running 24/7. Actual prices vary by provider and commitment; spot and reserved capacity can be significantly cheaper.
Deploy on a GPU cloud
Rent 3x Instinct MI250X by the hour instead of buying hardware.