Skip to content
VRAMwise

Local LLM fit · before the 40 GB download

Parameter count × bits is not enough. Context eats VRAM.

VRAMwise estimates whether a GGUF quant fits on a discrete GPU: weights from published file sizes, KV cache from architecture fields, plus a visible overhead %. Verdicts are FITS / TIGHT / WON'T FIT — not marketing scores.

Run the fit checker →

How this differs from “params × 2 bytes” widgets

Workflow

  1. Fit checker — GPU + model + quant + context
  2. Context explorer — raise tokens until it breaks
  3. Quant picker — Q4/Q5/Q8 vs budget
  4. Spec tables — raw grids
VRAMwise — Will This Local Model Fit My GPU? — by VRAMwise ↗