GeForce RTX 5090: VRAM & local model fit
32 GB · ~1792 GB/s · GDDR7 · NVIDIA.
Capacity is the first filter; bandwidth is a speed hint. The table below is a weight-only educational sketch (rough Q4-ish), not a lab certificate. Add context KV in the context tool and overall fit in VRAM fit.
| Model | Rough need | vs 32 GB |
|---|---|---|
| Llama 3.1 8B Instruct | ~5.9 GB | likely* |
| Llama 3.3 70B Instruct | ~40.3 GB | tight/no* |
| Qwen3 8B | ~6.0 GB | likely* |
| Qwen3 14B | ~9.6 GB | likely* |
| Qwen3 32B | ~19.5 GB | likely* |
| Qwen3 30B-A3B (MoE) | ~18.3 GB | likely* |
| Gemma 3 12B IT | ~8.2 GB | likely* |
| Gemma 3 27B IT | ~16.6 GB | likely* |
| Mistral Small 3.2 24B | ~14.7 GB | likely* |
| Phi-4 (14.7B) | ~9.6 GB | likely* |
| gpt-oss-20b (MoE) | ~13.1 GB | likely* |
| gpt-oss-120b (MoE) | ~65.9 GB | tight/no* |
*Heuristic (params × factor + overhead). Apple unified memory uses different rules — see Apple silicon.
Educational estimates. Runtime overhead varies.