Alibaba

Qwen 3.6 Plus (110B)

Strong all-round model with high reasoning performance

Model Summary

Family

Qwen

Version

3.6

Parameters

110B (est.)

Parameter counts for closed models are estimates; vendors rarely publish exact sizes.

VRAM Requirements by Quantization

Memory needed to serve Qwen 3.6 Plus (110B) for inference, including a 20% overhead for activations and KV cache.

PrecisionVRAM neededSmallest single GPU that fits
INT4 (4-bit)61.47 GBNVIDIA H100 SXM5 (80 GB)
INT8 (8-bit)122.93 GBAMD Instinct MI250X (128 GB)
FP16 (16-bit)245.87 GBAMD Instinct MI325X (256 GB)
FP32 (32-bit)491.74 GBMulti-GPU required

Recommended GPU Configurations

Cheapest on-demand configurations to serve Qwen 3.6 Plus (110B) at 8-bit (123 GB VRAM).

1x AMD Instinct MI250X

128 GB total VRAM · CDNA 2

~$2.50/h

8x NVIDIA T4

128 GB total VRAM · Turing

~$4.00/h

8x NVIDIA P100 SXM2

128 GB total VRAM · Pascal

~$4.80/h

Quick GPU Planning

Use the calculator pre-filled with this exact version to estimate memory, speed, and compute requirements in a few clicks.

Access Pre-filled Calculator

Frequently Asked Questions

How much VRAM do you need to run Qwen 3.6 Plus (110B)?

With an estimated 110B parameters, Qwen 3.6 Plus (110B) needs roughly 123 GB of VRAM in 8-bit (INT8), 61 GB in 4-bit, and 246 GB in FP16, including a 20% overhead for activations and KV cache.

Which GPUs can run Qwen 3.6 Plus (110B)?

At 8-bit quantization, the most cost-effective option is 1x AMD Instinct MI250X (128 GB combined VRAM, around $2.50/hour on-demand). Higher-end cards like the NVIDIA B200 or AMD MI355X reduce the GPU count needed.

Can Qwen 3.6 Plus (110B) run on a single GPU?

Yes. In 8-bit, a single AMD Instinct MI250X (128 GB) fits the model.

How much does it cost to serve Qwen 3.6 Plus (110B) in the cloud?

Renting 1x Instinct MI250X costs on the order of $2.50/hour, i.e. about $1,825/month running 24/7. Actual prices vary by provider and commitment; spot and reserved capacity can be significantly cheaper.