The first question when buying AI hardware is almost always: will my model fit? You can estimate it in a minute.
The rule of thumb
Memory for the weights ≈ number of parameters × bytes per parameter
| Precision | Bytes per parameter | 8B model | 70B model | 200B model |
|---|---|---|---|---|
| FP16 / BF16 | 2 | ~16 GB | ~140 GB | ~400 GB |
| FP8 / INT8 | 1 | ~8 GB | ~70 GB | ~200 GB |
| 4-bit (FP4, INT4, Q4) | ~0.5 | ~4–5 GB | ~35–40 GB | ~100–110 GB |
Then add headroom of roughly 20–30% for the KV cache (the model's working memory for your conversation), activations and the software itself.
Context length eats memory too
The KV cache grows with context length and with the number of people using the model at once. A model that fits comfortably with a short prompt can run out of memory at 128K tokens of context or with ten users. If you work with long documents or serve a team, size generously.
Quantisation: the trade-off
Quantisation stores weights with fewer bits. 8-bit is close to lossless for most uses. Good 4-bit formats (including NVIDIA's FP4 on Blackwell) keep quality surprisingly close to the original while cutting memory by about 4× versus FP16. That is how a 128 GB GB10 system can run models of up to roughly 200 billion parameters.
What fits on what we sell
| Hardware | Memory | Comfortable at 4-bit |
|---|---|---|
| RTX PRO 6000 (96 GB) | 96 GB GDDR7 | Up to ~120B, very fast |
| One GB10 system | 128 GB unified | Up to ~200B |
| Two linked GB10 systems | 2 × 128 GB | Up to ~405B |
These are rough guides; actual limits depend on the model architecture, the runtime and your context length.
Memory size vs memory speed
Capacity decides whether a model runs; memory bandwidth largely decides how fast it writes tokens. A discrete GPU like the RTX PRO 6000 (about 1.8 TB/s) generates text much faster than a GB10 system (about 273 GB/s) for any model that fits in its 96 GB. GB10 systems trade speed for capacity in a small, quiet box. Pick on what matters more for you: the biggest models, or the fastest responses.
Worked example
You want to run a 70-billion-parameter model for a team of five with long documents:
- Weights at 4-bit: ~40 GB
- Headroom for long context and five users: plan for another ~20–40 GB
- Total: ~60–80 GB, which fits on one GB10 system or one RTX PRO 6000
Want a second opinion on your setup? Send us your model and use case.