How Much VRAM Do You Need to Run an LLM Locally? A Practical Sizing Guide (2026)
Sizing numbers in this guide are widely cited community rules of thumb for planning, not guarantees. Actual memory use varies with the runtime (llama.cpp, vLLM, Ollama), the quantization format, your context length, and how many users you serve at once. Always test with your own model and workload. The method below matters more than any single number.
On this page
The one equation that answers everything
Every "can I run this model?" question comes down to one simple formula:
VRAM needed ≈ (parameters in billions × bits per weight) ÷ 8
That gives you the model weights in gigabytes. A 7-billion-parameter model stored at 16-bit precision needs roughly 14 GB just for the weights. At 8-bit, roughly 7 GB. At 4-bit, roughly 3.5–4 GB. That single relationship — smaller precision, smaller footprint — is the entire reason local AI is possible on consumer hardware, and it leads directly to the next section.
The quantization ladder
Quantization means storing the model's weights in fewer bits than the original, at a modest cost in quality. The 4-bit formats (like GGUF Q4) have become the community standard for local inference: each step down the ladder roughly halves the memory while adding a little quantization error. Here is what the ladder looks like for a typical 7B model (approximate, weights only):
| Precision | Bits per weight | 7B model weights | Quality |
|---|---|---|---|
| FP16 | 16 | ~14 GB | Reference |
| Q8 | 8 | ~7 GB | Near-lossless |
| Q6 | 6 | ~5.5 GB | Very good |
| Q4 | 4 | ~4 GB | Good — the sweet spot |
| Q3 | 3 | ~3 GB | Noticeable loss |
That is why Q4 is the default answer: a 7B model goes from unfittable on most consumer cards at FP16 to comfortable on an 8 GB card at Q4, with quality close enough that most users can't tell the difference on everyday tasks. Going below Q4 is only worth it when the alternative is not running the model at all.
What model size fits on your GPU
Apply the equation, add roughly 15–20% on top for the runtime and a short context, and you get the practical tiers everyone plans against. The table below assumes 4-bit quantized dense models (the common case for local use):
| Your VRAM | Comfortable territory (4-bit) | Typical use |
|---|---|---|
| 8 GB | 7B–8B models | Entry point: chat assistants, short completions, autocomplete-length answers |
| 12 GB | 7B–14B models | Practical minimum for serious 14B-class workflows with realistic context |
| 16 GB | 14B–20B models | Strong range for professional workloads and larger mid-size models |
| 24 GB | 30B–35B models | Major jump: large local models with room for generous context |
| 40–48 GB | 70B-class models | High-quality 70B inference, longer contexts, multi-user serving |
| 80 GB+ | 100B+ models | Large open-weight models on datacenter cards or multi-GPU rigs |
Two caveats. First, fitting is not the same as comfortable. A model technically squeezed into VRAM leaves no headroom for the runtime, the KV cache, or anything else your machine is doing — and the moment a layer spills into system RAM, speed collapses. Second, mixture-of-experts models play by different rules: for MoE models, memory depends on total parameters, but compute depends on the parameters active per token — which is why some very large MoE models run surprisingly well on modest cards.
Don't forget the KV cache
The weights are only the first line of the bill. During generation, the model keeps a KV cache — a growing record of what it has processed — and the longer your context, the bigger it gets. A short Q&A session adds almost nothing. A long coding-agent session with tens of thousands of tokens of context can add gigabytes on top of the weights, easily pushing a "fits on paper" setup over the edge.
When you size a deployment, budget for the model plus the context you actually plan to run — not the context you wish you could run. If you serve multiple users or run parallel requests, each stream adds its own cache. This is the most common reason a setup that "should fit" runs out of memory in practice.
VRAM capacity isn't the whole story
Two cards with the same 24 GB can deliver very different token speeds. The hidden spec is memory bandwidth — how fast data moves between VRAM and the GPU cores, measured in GB/s. During generation, an LLM mostly waits on memory reads, so bandwidth largely sets your tokens per second. This is why older datacenter cards with huge bandwidth but modest capacity can feel faster for inference than newer consumer cards with the same VRAM.
The practical rule: when two cards both fit your model, pick the one with more bandwidth. And if your local model feels inexplicably slow, the usual culprit is that part of it spilled into system RAM — shrinking the context or moving down a quant usually fixes it.
When the model doesn't fit
Sometimes the model you want is bigger than any single card you own. You have four options, in roughly the order most people should consider them:
- Quantize more aggressively. Dropping from Q4 to Q3 (or running fewer layers on the GPU) buys you memory at some quality cost. A smaller quality hit beats not running the model.
- Split across GPUs. Tensor parallelism lets a model span two or more cards. See our guide on multi-GPU setups for when that makes sense.
- Use unified memory. Systems with unified memory (like Apple Silicon Macs with large RAM) let the model use system memory as one big pool — slower than VRAM, but it unlocks 70B-class and larger models without a datacenter card.
- Rent. If the model is much bigger than your hardware — and for the frontier open-weight models, it will be — renting is not a failure, it is the sensible default. Our rent-vs-buy guide and provider directory walk through the math.
The sizing checklist
- Start with the weights: parameters in billions × bits per weight ÷ 8. For MoE, use total parameters, not active.
- Add roughly 15–20% for the runtime and a short context.
- Budget the KV cache for the context length you actually plan to use, plus parallel users if you serve more than yourself.
- Pick the precision that fits: prefer a 4-bit larger model over a full-precision smaller one at the same memory budget.
- Match to a tier, with headroom — "fits exactly" is not a plan.
Run the numbers for your setup
Now that you know how to size the memory, size the money:
- Estimate your monthly usage with our GPU-hours estimation guide.
- Look up current rental rates for your GPU tier in the provider directory and understand the pricing models.
- Compare local ownership against renting in the GPU cost calculator — remember, a card you buy for inference is only worth it if it will actually be busy.
- Stretch the budget with the cost-saving strategies guide: right-sizing and quantization move every crossover point in your favor.
If the model you want dwarfs your hardware, that's not a failure of your setup — it's a signal. Rent until your volume justifies fixed hardware.
Related guides
- How to Estimate GPU Hours for Training and Fine-Tuning
- Rent vs Buy a GPU for AI: A Practical Decision Framework
- A100 vs H100 vs L40S: Which GPU for Your AI Workload
- Multi-GPU Setups: NVLink, Scaling, and When You Need More Than One GPU
- How GPU Cloud Pricing Works: On-Demand vs Spot vs Reserved
- Cost-Saving Strategies for Indie AI Developers