How we calculate VRAM requirements
Every verdict on this site should be reproducible from plain text.
Weights
Prefer published GGUF sizes per quant from our model table. If missing, tools that need a size ask you to enter one rather than inventing silently.
KV cache
From architecture fields when present: bytes per token times context length.
Overhead
Default 10% on (weights+kv). Increase for vision stacks or concurrent batches.
What we do not claim
Exact tokens/sec, multi-GPU sharding, or vendor marketing “AI TOPS” conversions.
General information, not personalized advice.