Skip to content
VRAMwise

GeForce RTX 5060: VRAM & local model fit

8 GB · ~448 GB/s · GDDR7 · NVIDIA.

Capacity is the first filter; bandwidth is a speed hint. The table below is a weight-only educational sketch (rough Q4-ish), not a lab certificate. Add context KV in the context tool and overall fit in VRAM fit.

ModelRough needvs 8 GB
Llama 3.1 8B Instruct~5.9 GBlikely*
Llama 3.3 70B Instruct~40.3 GBtight/no*
Qwen3 8B~6.0 GBlikely*
Qwen3 14B~9.6 GBtight/no*
Qwen3 32B~19.5 GBtight/no*
Qwen3 30B-A3B (MoE)~18.3 GBtight/no*
Gemma 3 12B IT~8.2 GBtight/no*
Gemma 3 27B IT~16.6 GBtight/no*
Mistral Small 3.2 24B~14.7 GBtight/no*
Phi-4 (14.7B)~9.6 GBtight/no*
gpt-oss-20b (MoE)~13.1 GBtight/no*
gpt-oss-120b (MoE)~65.9 GBtight/no*

*Heuristic (params × factor + overhead). Apple unified memory uses different rules — see Apple silicon.

Educational estimates. Runtime overhead varies.

Frequently asked questions

VRAM?

8 GB.

GeForce RTX 5060 VRAM fit — by VRAMwise ↗