Spektiq

Is Your GPU's VRAM Enough for a Local LLM?

· 13 min read

The short answer, and why it depends

If you searched whether a card with only a little VRAM is enough to run an LLM locally, you probably own a laptop with a GeForce RTX 4050 or an older desktop card, and want to know whether to try a model or save up first. NVIDIA's own laptop spec table (read on 2026-10-05) places the RTX 4050 Laptop GPU at the entry memory tier of that generation, with the RTX 4060 and RTX 4070 Laptop GPUs above it.

The short answer is yes, with conditions. In the table below (a snapshot as of 2026-10-05), only the smallest model in its most compact format leaves clear headroom on a card at this entry tier: the next format up of the same model already comes level with the card's memory, and every larger model exceeds it with the weights alone. That one row may sit entirely on the GPU, but the table cannot confirm it, because it leaves out the context and the runtime's buffers. Whether a given model does depends on three things that the usual rule-of-thumb lists skip: how many bits per weight the file uses, how long a context you ask the runtime to reserve, and what the runtime does when the model spills over.

There are really two worries. The first is "will it run at all?", and the answer is usually yes, as long as the weights fit in system RAM and VRAM combined, because llama.cpp and Ollama split a model between CPU and GPU rather than refuse it. The second is "will it be usably fast?", and that is where the real threshold sits: not whether the model loads, but whether it fits entirely on the GPU.

Below: a sourced table of weight sizes (a lower bound), what it leaves out, and how to check where a model landed. Everything was read and checked on 2026-10-05.

What the weights alone take

A model file is mostly its weights. llama.cpp states that models are fully loaded into memory, so the weight file has to fit in system RAM, in VRAM, or across both.

The formats below are GGUF quantization levels from llama.cpp, which offers integer quantization "for faster inference and reduced memory use" on NVIDIA, AMD, Vulkan, SYCL and Apple Silicon backends, among others. F16 is the half-precision baseline.

Model Parameters (model card) Q4_K_M (4.8944 bits/weight) Q6_K (6.5633 bits/weight) Q8_0 (8.5008 bits/weight) F16 (16.0005 bits/weight) Sources
Qwen2.5-7B-Instruct 7.61 billion 4.7 GB 6.2 GB 8.1 GB 15.2 GB [1] [2]
Qwen2.5-14B-Instruct 14.7 billion 9.0 GB 12.1 GB 15.6 GB 29.4 GB [1] [3]
Qwen2.5-32B-Instruct 32.5 billion 19.9 GB 26.7 GB 34.5 GB 65.0 GB [1] [4]
Qwen2.5-72B-Instruct 72.7 billion 44.5 GB 59.6 GB 77.3 GB 145.4 GB [1] [5]

Each value is computed, not measured: parameters × bits per weight ÷ 8, in decimal gigabytes (1 GB = 10^9 bytes, slightly less than one GiB).

The bits-per-weight figures were measured by the llama.cpp project on Llama 3.1 8B files. Applied to other models they estimate the size of the weight file; they are not a measured size for that model.

llama.cpp states that memory and disk requirements are the same because models are fully loaded into memory: the weight file has to fit in RAM, in VRAM, or across both.

This table counts the weights only. It does not count the memory a runtime uses for the conversation context or its own buffers, so it is a lower bound, not a full hardware requirement.

Lower-bit formats trade quality for size. The table says what fits, not how well a given format answers.

Sources

  1. llama-quantize README, Memory/Disk Requirements and Quantization sections (measured on Llama 3.1 8B), ggml-org/llama.cpp (GitHub), read on 2026-10-05
  2. Qwen2.5-7B-Instruct model card, Qwen (Hugging Face), read on 2026-10-05
  3. Qwen2.5-14B-Instruct model card, Qwen (Hugging Face), read on 2026-10-05
  4. Qwen2.5-32B-Instruct model card, Qwen (Hugging Face), read on 2026-10-05
  5. Qwen2.5-72B-Instruct model card, Qwen (Hugging Face), read on 2026-10-05

As the notes under the table say, each value is computed: for other models, it is an estimate of the weight file, not a measured size.

Most importantly, the table is a lower bound: read a row as "its weights alone come to this value", never "this model needs this much". The runtime still needs memory for the context and its own buffers, which the next section covers.

To use a row: compare it with the free memory your own system reports (see the checking section), after converting both values to the same unit. The table is in decimal gigabytes, while nvidia-smi and llama.cpp's logs report binary units (MiB or GiB): convert explicitly rather than comparing raw numbers. If the weights alone already exceed the free memory, the model cannot sit entirely on the GPU. If they leave headroom, the context and buffers decide, and only a test on your machine settles it.

Finally, fitting is not answering well. The llama.cpp quantize documentation (see the table's sources) says reducing weight precision can introduce accuracy loss (measured with perplexity and KL divergence), which an importance matrix (--imatrix) can limit. Hugging Face adds that a lower-bit format can give different results, so test on your own task.

What the table leaves out: context is memory too

While generating, a model keeps a key and a value for every token already seen, so it need not recompute them: the key-value cache, or KV cache. Hugging Face's optimization guide (read 2026-10-05) explains that for short inputs, memory is dominated by the weights, but the KV cache grows linearly with the sequence length and becomes a serious memory cost for long inputs and multi-turn chat.

So the context length you set changes the answer for the same model file: weights that fit with room to spare can stop fitting once a long context is reserved.

Architecture matters too: Grouped-Query and Multi-Query Attention shrink the cache by sharing key and value heads across query heads. Recent models such as Qwen3 use Grouped-Query Attention, so similar weight sizes can still mean different cache sizes.

Defaults often decide for you. In llama-server, an unset -c / --ctx-size is loaded from the model's metadata, so it can reserve the model's full native context. Ollama picks a default context length based on the VRAM it detects, with the smallest default for cards with little VRAM, and its context length documentation is explicit that a larger context increases the memory needed to run the model.

Runtimes also keep working buffers, and other processes (the desktop, a browser) can hold VRAM: nvidia-smi or the runtime's startup log shows what is in use and what is free.

When it doesn't fit: the split and the cliff

The llama.cpp README describes CPU+GPU hybrid inference, used to partially accelerate models larger than the total VRAM capacity. A model that does not fit is not refused: some layers go on the GPU and the rest run on the CPU. Ollama also splits a model between CPU and GPU when it does not fit, and tells you when it happens, as the checking section shows.

So "it loads" is not a useful test. The useful test is whether it sits entirely on the GPU.

The best independent evidence I found is an InventiveHQ sweep published on 2026-06-26. They ran llama.cpp on an RTX 5060 Ti with a mid-sized Qwen2.5-Coder model in Q4_K_M and moved -ngl step by step from all layers on the CPU to all layers on the GPU. Their conclusion: "partial offload is a cliff, not a slope". Throughput stayed far below the full-GPU figure until the last layers moved over, with no sweet spot.

That is one GPU, one model and one date: a strong signal, not a law.

The takeaway for a card at this tier: choose model, format and context so everything lands on the GPU, rather than a bigger model that half-fits.

Levers that trade memory for something else

Each of these frees memory, and most of them cost something in return (Flash Attention is the exception on output quality, but it depends on hardware support). Flags are from the llama-server README and Ollama docs as read on 2026-10-05; check --help on your version.

A shorter context

Set the context instead of inheriting a default. The cost: the model sees less of the conversation at once.

llama-server -m <model.gguf> -c <context>
OLLAMA_CONTEXT_LENGTH=<context> ollama serve

These lines use Unix shell syntax. On Windows, quit the Ollama app from the taskbar, set the variable in your account's environment variables and restart the app, as Ollama's FAQ describes (or set $env:OLLAMA_CONTEXT_LENGTH="<context>" in PowerShell before ollama serve).

In the Ollama app, it is a slider.

A quantized KV cache

llama.cpp lets you pick the cache data type separately for keys and values with --cache-type-k and --cache-type-v; the default is f16, and quantized types such as q8_0 are available. In Ollama, the FAQ documents OLLAMA_KV_CACHE_TYPE, with f16 by default and q8_0 or q4_0 as options, applied only when Flash Attention is enabled. Ollama itself describes q8_0 as a "very small loss in precision" and q4_0 as a "small-medium loss": that is the vendor's description, not an independent measurement.

llama-server -m <model.gguf> --cache-type-k q8_0 --cache-type-v q8_0
OLLAMA_KV_CACHE_TYPE=q8_0 ollama serve

Flash Attention

Flash Attention computes attention that is mathematically equivalent to standard attention, with possible small differences from numerical rounding. According to Hugging Face, its memory grows linearly, not quadratically, with input length. In llama-server it is set to auto by default, and Ollama turns it on automatically when the hardware supports it; OLLAMA_FLASH_ATTENTION=1 forces it.

Fewer layers on the GPU

-ngl / --n-gpu-layers sets the maximum number of layers stored in VRAM (an exact number, auto or all), and --fit adjusts unset arguments so the model fits in device memory. Remember the cliff: this lever makes a model run, not run fast.

llama-server -m <model.gguf> -ngl <layers>

For Mixture-of-Experts models, --cpu-moe (or --n-cpu-moe <layers> for the first layers) keeps the expert weights in system RAM while the rest stays on the GPU. The log's layer counter then misses part of the picture (see the checking section). Test it yourself; I have no verified speed figure.

A smaller weight format

A lower-bit column of the table is the most direct saving, with the earlier caveat: it fits more easily, but does not promise the same answers.

Check where your model actually landed

Stop trusting tables here, including mine. With Ollama, load the model, then ask where it sits:

ollama run <model>
ollama ps

The Processor column of ollama ps tells you whether the model is entirely on the GPU, entirely on the CPU, or split between the two. A split is the warning sign from the previous section: it runs, but going by the InventiveHQ sweep it can be much slower.

On NVIDIA hardware, nvidia-smi shows memory in use and free, in its own units, and the processes holding VRAM.

nvidia-smi

With llama.cpp directly, read the startup log: its offload line reports how many layers went to the GPU out of a total the runtime counts itself (wording varies between versions). Compare the offloaded count with the total on that line, not with the model card's block count: the runtime's total also includes the output layer. An offloaded count below the total proves the model is only partly on the GPU, but not yet why. Before blaming free memory, check whether you set an explicit layer limit with -ngl, which --fit leaves as you set it, and read what the log says was adjusted. If you forced all layers, do not assume that a model which started fits: check free memory in nvidia-smi and watch the speed.

That counter follows the layer setting, not where every tensor ended up. On a Mixture-of-Experts model, expert weights can stay in system RAM while the counter reads complete, with --cpu-moe or --n-cpu-moe but also under auto or --fit, which can move expert tensors to the CPU by itself. For any such model, whatever the options, also read the log lines on tensor placement and model buffers. A small CPU-side buffer alone is normal: llama.cpp usually keeps the input embedding there. Expert tensors placed on the CPU, or a CPU model buffer far larger than that input part, mean part of the model runs there.

A protocol:

  1. Pick the smallest model in the table (a snapshot as of 2026-10-05), in its most compact format, and check its weights against your card's free memory, both converted to the same unit, as described earlier.
  2. Set the context explicitly, at the length you actually plan to use.
  3. Check the Processor column in ollama ps; with llama.cpp, compare the offloaded count with the log's total and, for any Mixture-of-Experts model (auto and --fit included), also read the tensor placement and model buffer lines.
  4. If it shows a split, apply one lever from the previous section and check again, one change at a time.

The test is "entirely on the GPU, at the context I actually use", not "it started".

If it still doesn't fit

A card at this tier is one honest constraint of running a model locally, not a dead end. It can hold small models in compact formats entirely on the GPU, if you set the context deliberately and check where the model landed. Beyond that, the split still works, but going by the InventiveHQ sweep, expect it to be much slower.

When a model does not fit, work through the options in order of cost: a smaller model or format (which may answer less well), then a shorter context and a quantized KV cache. A partial offload can suit occasional, non-interactive jobs where waiting is fine. Only if your use genuinely needs more is it time to look at more VRAM, or at a hosted model, and it is worth knowing what a hosted model really costs, from an independent look before you decide.

Whatever you choose, test it with the commands above: the answer depends on your settings and runtime version, and this article is a snapshot dated 2026-10-05.