How much VRAM to run DeepSeek R1 Distill Qwen3 8B?

About 8 GB atQ4_K_M with an 8K context — fits a RTX 3060 12GB. Full breakdown below, or check your exact hardware.

DeepSeek R1 Distill Qwen3 8B VRAM by quantisation

QuantisationWeightsTotal (8K ctx)Fits on
Q4_K_M5.0 GB7.8 GBRTX 3060 12GB, RTX 5060 Ti 16GB
Q5_K_M5.8 GB8.8 GBRTX 3060 12GB, RTX 5060 Ti 16GB
Q6_K6.7 GB9.7 GBRTX 3060 12GB, RTX 5060 Ti 16GB
Q8_08.7 GB11.9 GBRTX 3060 12GB, RTX 5060 Ti 16GB
FP16 / BF1616.4 GB20.4 GBRX 7900 XTX, RTX 4090

Check your hardware

About DeepSeek R1 Distill Qwen3 8B

DeepSeek R1 Distill Qwen3 8B is DeepSeek's 8.2B-parameter model released in May 2025, with a 128K-token context window. It uses a classic dense-attention design whose KV cache grows linearly with context: its KV cache is about 1.2 GB at an 8K context, 19.3 GB at 128K, and 19.3 GB at the full 128K window (FP16 cache).

For most people Q4_K_M is the sweet spot — the most popular quality/size trade-off — while Q8 is near-lossless if you have the memory. Totals above include the KV cache and a realistic framework overhead, so they are what you should expect to see in practice rather than just the download size. Weight sizes are calibrated against real GGUF files — see themethodology.

Frequently asked questions

How much VRAM does DeepSeek R1 Distill Qwen3 8B need?

At Q4_K_M with an 8K context, DeepSeek R1 Distill Qwen3 8B needs about 8 GB (weights 5 GB + KV cache + overhead). The smallest common hardware that fits is a RTX 3060 12GB.

Can an RTX 4090 (24GB) run DeepSeek R1 Distill Qwen3 8B?

Yes. An RTX 4090's 24 GB runs DeepSeek R1 Distill Qwen3 8B at FP16 / BF16 (about 20 GB at 8K context) — at full FP16 precision.

Can a Mac run DeepSeek R1 Distill Qwen3 8B?

Yes — Apple Silicon with 16 GB of unified memory or more (macOS lets the GPU use ~75% of it, ~12 GB) runs DeepSeek R1 Distill Qwen3 8B at Q4_K_M.

Related

VRAM calculator for any model ·Token counter

Last updated 2026-08-03. Architecture figures from the model's published config.json; see themethodology.