Skip to content
VRAMwise

Llama 3.3 70B Instruct: VRAM by quant

70.6B params · KV field 327680 · max context field 131072.

QuantBitsWeight GB sketch
Q4_K_M4.85 bpw~43.8 GB
Q5_K_M5.7 bpw~50.8 GB
Q6_K6.6 bpw~58.6 GB
Q8_08.5 bpw~75.5 GB
MXFP44.25 bpw~38.8 GB
F16/BF1616.0 bpw~141.2 GB

Add KV for your context length in context VRAM. Overall fit: VRAM fit.

Educational. MoE active params may differ from total params.

Frequently asked questions

Size?

70.6B.

Llama 3.3 70B Instruct VRAM requirements — by VRAMwise ↗