All how to guidesSizing

How much memory do you need to run an LLM locally?

A simple rule of thumb to work out whether a model fits on your hardware, what quantisation does, and why context length matters.

3 min read ai StoreLabs

Asus Ascent GX10 with NVIDIA GB10

The first question when buying AI hardware is almost always: will my model fit? You can estimate it in a minute.

The rule of thumb

Memory for the weights ≈ number of parameters × bytes per parameter

Precision Bytes per parameter 8B model 70B model 200B model
FP16 / BF16 2 ~16 GB ~140 GB ~400 GB
FP8 / INT8 1 ~8 GB ~70 GB ~200 GB
4-bit (FP4, INT4, Q4) ~0.5 ~4–5 GB ~35–40 GB ~100–110 GB

Then add headroom of roughly 20–30% for the KV cache (the model's working memory for your conversation), activations and the software itself.

Context length eats memory too

The KV cache grows with context length and with the number of people using the model at once. A model that fits comfortably with a short prompt can run out of memory at 128K tokens of context or with ten users. If you work with long documents or serve a team, size generously.

Quantisation: the trade-off

Quantisation stores weights with fewer bits. 8-bit is close to lossless for most uses. Good 4-bit formats (including NVIDIA's FP4 on Blackwell) keep quality surprisingly close to the original while cutting memory by about 4× versus FP16. That is how a 128 GB GB10 system can run models of up to roughly 200 billion parameters.

What fits on what we sell

Hardware Memory Comfortable at 4-bit
RTX PRO 6000 (96 GB) 96 GB GDDR7 Up to ~120B, very fast
One GB10 system 128 GB unified Up to ~200B
Two linked GB10 systems 2 × 128 GB Up to ~405B

These are rough guides; actual limits depend on the model architecture, the runtime and your context length.

Memory size vs memory speed

Capacity decides whether a model runs; memory bandwidth largely decides how fast it writes tokens. A discrete GPU like the RTX PRO 6000 (about 1.8 TB/s) generates text much faster than a GB10 system (about 273 GB/s) for any model that fits in its 96 GB. GB10 systems trade speed for capacity in a small, quiet box. Pick on what matters more for you: the biggest models, or the fastest responses.

Worked example

You want to run a 70-billion-parameter model for a team of five with long documents:

  1. Weights at 4-bit: ~40 GB
  2. Headroom for long context and five users: plan for another ~20–40 GB
  3. Total: ~60–80 GB, which fits on one GB10 system or one RTX PRO 6000

Want a second opinion on your setup? Send us your model and use case.

Keep reading

More guides.