Context budgeting for local LLMs
Context is a budget line, not a trophy setting.
Start from the task
Customer-support style chat may live at 4–8k. Agent workflows that stuff docs need more — and more KV.
Measure twice
Set quant and GPU first. Raise context only while the verdict stays FITS with comfortable headroom.
Write it down
Record model, quant, context, GPU, and measured tok/s. That note is more valuable than another undated Reddit screenshot.
General information, not personalized advice.
Models in band
Runtime tools
Methodology paths