Alibaba
Qwen3.8 Flash Next
Alibaba model discovered on huggingface
Model Summary
Family
Qwen3.8
Version
3.8
Parameters
180B (est.)
Parameter counts for closed models are estimates; vendors rarely publish exact sizes.
VRAM Requirements by Quantization
Memory needed to serve Qwen3.8 Flash Next for inference, including a 20% overhead for activations and KV cache.
| Precision | VRAM needed | Smallest single GPU that fits |
|---|---|---|
| INT4 (4-bit) | 100.58 GB | AMD Instinct MI250X (128 GB) |
| INT8 (8-bit) | 201.17 GB | AMD Instinct MI325X (256 GB) |
| FP16 (16-bit) | 402.33 GB | Multi-GPU required |
| FP32 (32-bit) | 804.66 GB | Multi-GPU required |
Recommended GPU Configurations
Cheapest on-demand configurations to serve Qwen3.8 Flash Next at 8-bit (201 GB VRAM).
2x AMD Instinct MI250X
256 GB total VRAM · CDNA 2
~$5.00/h
13x NVIDIA T4
208 GB total VRAM · Turing · multi-node
~$6.50/h
13x NVIDIA P100 SXM2
208 GB total VRAM · Pascal · multi-node
~$7.80/h
Quick GPU Planning
Use the calculator pre-filled with this exact version to estimate memory, speed, and compute requirements in a few clicks.
Access Pre-filled CalculatorFrequently Asked Questions
How much VRAM do you need to run Qwen3.8 Flash Next?
With an estimated 180B parameters, Qwen3.8 Flash Next needs roughly 201 GB of VRAM in 8-bit (INT8), 101 GB in 4-bit, and 402 GB in FP16, including a 20% overhead for activations and KV cache.
Which GPUs can run Qwen3.8 Flash Next?
At 8-bit quantization, the most cost-effective option is 2x AMD Instinct MI250X (256 GB combined VRAM, around $5.00/hour on-demand). Higher-end cards like the NVIDIA B200 or AMD MI355X reduce the GPU count needed.
Can Qwen3.8 Flash Next run on a single GPU?
Yes. In 8-bit, a single AMD Instinct MI325X (256 GB) fits the model.
How much does it cost to serve Qwen3.8 Flash Next in the cloud?
Renting 2x Instinct MI250X costs on the order of $5.00/hour, i.e. about $3,650/month running 24/7. Actual prices vary by provider and commitment; spot and reserved capacity can be significantly cheaper.
Deploy on a GPU cloud
Rent 2x Instinct MI250X by the hour instead of buying hardware.