To run an AI model for inference, the GPU must have enough memory (VRAM) to hold the model's weights plus a working cache. The weights need roughly parameters × bytes per parameter: about 2 bytes in FP16/BF16, 1 byte in 8-bit and 0.5 byte in 4-bit. A 7-billion-parameter model therefore needs roughly 14 GB in FP16, and a 70-billion-parameter model roughly 140 GB. This guide shows how to size a GPU server for inference, and where refurbished GPU servers fit.
Training stores gradients and optimizer states as well as the weights; a commonly cited figure for mixed-precision training with the Adam optimizer is around 16 bytes per parameter. Inference stores the weights and a cache, so the same model needs far less memory to serve than to train. That is why inference can often run on one GPU where training needs a multi-GPU system. For a comparison of running your own hardware and other options, see our GPU servers buy vs rent guide.
| Model size | FP16 / BF16 | 8-bit | 4-bit |
|---|---|---|---|
| 7B | ~14 GB | ~7 GB | ~3.5 GB |
| 13B | ~26 GB | ~13 GB | ~6.5 GB |
| 34B | ~68 GB | ~34 GB | ~17 GB |
| 70B | ~140 GB | ~70 GB | ~35 GB |
These are weights only. Real usage is higher because of the cache and runtime overhead, so treat the numbers as a minimum.
Language models keep a key-value (KV) cache so they do not recompute the whole conversation for every new token. It grows with the length of the context and with the number of requests handled at once. A model that fits on paper can therefore run out of memory when many users connect, or when prompts are long. Leave headroom and test with your own prompt lengths and expected concurrency.
Serverwale lists these GPUs on its GPU servers page. VRAM is per GPU:
| GPU | VRAM | Weights that fit on one card |
|---|---|---|
| NVIDIA RTX 4090 | 24 GB | 7B in FP16; 13B in 8-bit; about 34B in 4-bit |
| NVIDIA V100 (refurbished) | 32 GB | 13B in 8-bit; about 34B in 4-bit |
| NVIDIA RTX A6000 / L40S | 48 GB | 13B in FP16; 34B in 8-bit; 70B in 4-bit |
| NVIDIA A100 / H100 | 80 GB | 34B in FP16; 70B in 8-bit with little headroom |
A 70B model in FP16 needs at least two 80 GB GPUs. Quantization (8-bit or 4-bit) lets a larger model fit on a smaller card at some cost in output quality, which you should measure for your task.
Once a model fits, the next question is how many users it can serve. More concurrent requests need more cache memory and more compute; interactive chat needs low latency, while batch jobs can run on a smaller GPU. Serving software such as vLLM, NVIDIA Triton or Ollama batches requests to use the GPU efficiently.
Inference runs continuously, so owning hardware often works out cheaper than renting when usage is steady. Refurbished data-centre GPUs cost less than new ones; Serverwale lists refurbished V100 and A100 systems at 50–60% below new. Older architectures may lack newer features (for example, V100 predates BF16 support), so check that your software supports the card. See the GPU server price guide for budget ranges, and remember that GPU servers need proper power and cooling: cooling and airflow, UPS sizing and PSU redundancy.
For a fuller walk-through, read Serverwale's GPU server inference sizing guide and the A100 vs RTX A6000 vs RTX 4090 comparison. For full-stack builds, see AI infrastructure.
How much VRAM do I need for a 7B model? About 14 GB for the weights in FP16, or around 7 GB in 8-bit, plus memory for the KV cache. A 24 GB card runs a 7B model in FP16 comfortably.
Can I run a 70B model on one GPU? Only when quantized: about 35 GB of weights in 4-bit fits on a 48 GB card, and about 70 GB in 8-bit only just fits on an 80 GB card. In FP16 it needs at least two 80 GB GPUs.
What is the KV cache? The memory a language model uses to remember the conversation so far. It grows with context length and concurrent users.
Do I need multiple GPUs for inference? Not if the model and cache fit on one GPU. Multiple GPUs are needed for models that exceed one card's memory or to serve more users at once.
Is a refurbished GPU server good enough for inference? Yes, if the card supports your software and the system is tested and warranted. Confirm the GPU generation supports the precision and framework you plan to use.
Share the model, precision and number of users and we will recommend a configuration. Serverwale · Phone +91-87962-44410 · WhatsApp https://wa.me/918796244410 · Burari, Delhi 110084, India. See GPU servers and live stock.
Want a new custom build instead? ProStation Systems builds workstations and servers to order.