30-SECOND SUMMARY
What to take away
- Estimate weight size with parameters × bits per weight ÷ 8.
- Leave headroom for KV cache, the runtime, and the operating system.
- A smaller model that stays fully accelerated can feel better than a larger model that constantly spills to CPU memory.
What consumes memory
Model file size is only the first layer.
- 01WEIGHTS
Quantized parameters
- 02CONTEXT
KV cache
- 03RUNTIME
Buffers and backend
- 04SYSTEM
OS and applications
RAM, VRAM, and unified memory
System RAM is used by the operating system, CPU inference, and application state. Dedicated GPU VRAM stores data used directly by the GPU. Apple Silicon uses a unified memory pool shared by CPU and GPU.
These architectures are not directly interchangeable. Always test with the runtime and model you intend to use.
For verification, save the model and runtime versions, source input, relevant settings, and observed output together. Repeat the step while changing only one factor, and record unexpected results and untested limits as carefully as successes before applying the guidance to private or production data.
Estimate model-weight size
A rough lower bound is parameter count multiplied by average bits per weight, divided by eight. An 8B model at roughly 4.5 bits per weight is about 4.5GB for weights alone.
Metadata, alignment, runtime buffers, and architecture details can change the actual file and memory footprint.
Record the current version and settings before the example, then verify the expected response, file, or process afterward. Preserve the error and return to the smallest working command before adding options; this separates installation failures from input and integration failures.
weight GB ≈ parameters (billions) × bits per weight ÷ 8Context consumes additional memory
The KV cache grows with context length, model architecture, precision, and concurrent requests. Doubling context can materially increase memory even though the model file stays unchanged.
Use the smallest context that fits the real task and begin with one request at a time.
Define completion with an observable result instead of a general impression. Repeat the same input, and if the output changes, isolate whether the model, runtime settings, or source data changed before moving to the next stage.
Practical starting ranges
8GB systems are best treated as experiment machines for very small models. 16GB is a practical entry point for many 3B–8B Q4 models. 32GB opens more room for 8B–14B models and document workflows.
These are planning ranges, not universal minimums. CPU features and runtime support may prevent a model from running even when memory appears sufficient.
Treat the table as a recording framework, not as a universal answer. Fill it with measurements from the intended device and workload, and mark unavailable values as unknown instead of replacing them with zero or an estimate that could distort the comparison.
| Memory | Practical starting point | Typical constraint |
|---|---|---|
| 8GB | 1B–3B Q4 | Short context, close other apps |
| 16GB | 3B–8B Q4 | Limited large-model headroom |
| 32GB | 8B–14B Q4 candidates | Context and speed still matter |
| 64GB+ | 14B–32B Q4 candidates | Validate throughput and power |
Measure instead of guessing
Run the same prompt with one model at a time. Record peak memory, time to first output, sustained generation, and whether the model remains fully accelerated.
Stop when quality meets the task. Buying hardware for a larger model that does not improve your workflow is wasted capacity.
Preserve the before-and-after state and the time of the check so that another run can reproduce the result. Include at least one failure condition—such as empty input, constrained resources, or a restart—to reveal the boundary of the step rather than documenting only the happy path.
Frequently asked questions
Is file size equal to required RAM?
No. The runtime, context cache, operating system, and other applications require additional memory. For a practical check, follow the “RAM, VRAM, and unified memory” section, change one condition at a time, and record the result.
Is more VRAM always better?
Capacity matters, but bandwidth, software support, model architecture, and workload also affect performance. Leave headroom for KV cache, the runtime, and the operating system. For a practical check, follow the “Estimate model-weight size” section, change one condition at a time, and record the result.
Can a model larger than VRAM still run?
Some runtimes can split work across GPU and system memory, but performance may fall substantially. A smaller model that stays fully accelerated can feel better than a larger model that constantly spills to CPU memory. For a practical check, follow the “Context consumes additional memory” section, change one condition at a time, and record the result.
Primary sources
Check the original documentation for version-specific details.
Hugging Face GGUF Ollama Context Length llama.cpp Repository