Skip to content
VRAMwise

GeForce RTX 4090: VRAM & local model fit

24 GB · ~1008 GB/s · GDDR6X · NVIDIA.

Capacity is the first filter; bandwidth is a speed hint. The table below is a weight-only educational sketch (rough Q4-ish), not a lab certificate. Add context KV in the context tool and overall fit in VRAM fit.

ModelRough needvs 24 GB
Llama 3.1 8B Instruct~5.9 GBlikely*
Llama 3.3 70B Instruct~40.3 GBtight/no*
Qwen3 8B~6.0 GBlikely*
Qwen3 14B~9.6 GBlikely*
Qwen3 32B~19.5 GBlikely*
Qwen3 30B-A3B (MoE)~18.3 GBlikely*
Gemma 3 12B IT~8.2 GBlikely*
Gemma 3 27B IT~16.6 GBlikely*
Mistral Small 3.2 24B~14.7 GBlikely*
Phi-4 (14.7B)~9.6 GBlikely*
gpt-oss-20b (MoE)~13.1 GBlikely*
gpt-oss-120b (MoE)~65.9 GBtight/no*

*Heuristic (params × factor + overhead). Apple unified memory uses different rules — see Apple silicon.

Educational estimates. Runtime overhead varies.

Frequently asked questions

VRAM?

24 GB.

GeForce RTX 4090 VRAM fit — by VRAMwise ↗