How Much GPU Memory Does AI Model Training Need?
Key Takeaway
Calculate separately for three scenarios: inference depends on weights plus KV cache; fine-tuning depends on weights, gradients, optimizer states and activations; only training from scratch is a cluster-scale requirement. Quick rule of thumb: full fine-tuning is roughly "parameters (B) x 16 to 20 GB", LoRA roughly "x 1.2", and 4-bit QLoRA roughly "x 0.7".
Key Comparison
| Model Size | Full Fine-Tuning (minimum) | LoRA (rank 16 to 64) | 4-bit QLoRA | Inference Only (FP16) |
|---|---|---|---|---|
| 7B | 112 to 160 | 16 to 24 | 6 to 10 | 15 to 18 |
| 13B to 14B | 210 to 280 | 28 to 40 | 12 to 18 | 27 to 32 |
| 32B to 34B | 510 to 680 | 60 to 80 | 24 to 34 | 65 to 75 |
| 70B to 72B | 1,120 to 1,440 | 130 to 170 | 48 to 60 | 140 to 155 |
| 235B MoE | Cluster-scale, multi-node required | Needs 4 to 8 cards of 80GB or more | Needs 2 to 4 cards of 80GB or more | Needs 4 to 8 cards of 80GB or more |
(Units in GB, on the assumption that weights are stored in FP16/BF16, the optimizer is AdamW and there is no CPU offload.)
Practical Takeaways
- During training, memory is consumed by five items: weights (2 bytes per parameter), gradients (2 bytes), optimizer states (8 bytes for Adam), activations, and 1 to 2 GB of framework overhead - about 16 to 20 bytes per parameter in total.
- Inference memory = weights + KV cache + framework overhead. For a 7B model in FP16, the KV cache for a single request with 2,048-token context is about 1.4 GB (only 10% of the weights), but with 32 concurrent requests and the context stretched to 32,768 tokens, it can balloon tens of times over and become the main bottleneck.
- Precision directly determines the weight footprint: FP32 is 4 bytes per parameter, FP16/BF16 2, FP8/INT8 1 and FP4/INT4 0.5 (a 7B model would be 28GB, 14GB, 7GB and 3.5GB respectively).
- Six solutions for insufficient memory: lower the precision; LoRA/QLoRA (training only 0.1% to 1% of parameters, cutting memory by 70% to 90%); gradient checkpointing (reducing activation memory by 50% to 70%, with training 20% to 30% slower); offloading the optimizer to CPU; distributed strategies (DDP saves no memory; ZeRO-3 saves the most but carries the highest communication volume; TP requires NVLink's high bandwidth; PP involves less cross-node communication); and shrinking batch size and sequence length.
- Look beyond capacity: H100 bandwidth is 3.35 TB/s, H200 4.8 TB/s and B200 8 TB/s - far above the 1.79 TB/s of consumer cards; for multi-GPU training, NVLink (900 GB/s on H100) far outperforms PCIe 5.0 (about 64 GB/s unidirectional).
Tip: Practical reminder: when planning an inference server, allow 20% to 30% memory headroom, otherwise the shortfall often only surfaces under concurrent load testing at go-live. As for buying versus renting cloud GPUs, if utilisation is above 60%, the three-year TCO usually favours buying; at lower utilisation, the cloud is the better fit.