How Much GPU Memory Does AI Model Training Need?

Key Takeaway

Calculate separately for three scenarios: inference depends on weights plus KV cache; fine-tuning depends on weights, gradients, optimizer states and activations; only training from scratch is a cluster-scale requirement. Quick rule of thumb: full fine-tuning is roughly "parameters (B) x 16 to 20 GB", LoRA roughly "x 1.2", and 4-bit QLoRA roughly "x 0.7".


Key Comparison


Model Size Full Fine-Tuning (minimum) LoRA (rank 16 to 64) 4-bit QLoRA Inference Only (FP16)
7B 112 to 160 16 to 24 6 to 10 15 to 18
13B to 14B 210 to 280 28 to 40 12 to 18 27 to 32
32B to 34B 510 to 680 60 to 80 24 to 34 65 to 75
70B to 72B 1,120 to 1,440 130 to 170 48 to 60 140 to 155
235B MoE Cluster-scale, multi-node required Needs 4 to 8 cards of 80GB or more Needs 2 to 4 cards of 80GB or more Needs 4 to 8 cards of 80GB or more

(Units in GB, on the assumption that weights are stored in FP16/BF16, the optimizer is AdamW and there is no CPU offload.)


Practical Takeaways

  1. During training, memory is consumed by five items: weights (2 bytes per parameter), gradients (2 bytes), optimizer states (8 bytes for Adam), activations, and 1 to 2 GB of framework overhead - about 16 to 20 bytes per parameter in total.

  2. Inference memory = weights + KV cache + framework overhead. For a 7B model in FP16, the KV cache for a single request with 2,048-token context is about 1.4 GB (only 10% of the weights), but with 32 concurrent requests and the context stretched to 32,768 tokens, it can balloon tens of times over and become the main bottleneck.

  3. Precision directly determines the weight footprint: FP32 is 4 bytes per parameter, FP16/BF16 2, FP8/INT8 1 and FP4/INT4 0.5 (a 7B model would be 28GB, 14GB, 7GB and 3.5GB respectively).

  4. Six solutions for insufficient memory: lower the precision; LoRA/QLoRA (training only 0.1% to 1% of parameters, cutting memory by 70% to 90%); gradient checkpointing (reducing activation memory by 50% to 70%, with training 20% to 30% slower); offloading the optimizer to CPU; distributed strategies (DDP saves no memory; ZeRO-3 saves the most but carries the highest communication volume; TP requires NVLink's high bandwidth; PP involves less cross-node communication); and shrinking batch size and sequence length.

  5. Look beyond capacity: H100 bandwidth is 3.35 TB/s, H200 4.8 TB/s and B200 8 TB/s - far above the 1.79 TB/s of consumer cards; for multi-GPU training, NVLink (900 GB/s on H100) far outperforms PCIe 5.0 (about 64 GB/s unidirectional).

Tip: Practical reminder: when planning an inference server, allow 20% to 30% memory headroom, otherwise the shortfall often only surfaces under concurrent load testing at go-live. As for buying versus renting cloud GPUs, if utilisation is above 60%, the three-year TCO usually favours buying; at lower utilisation, the cloud is the better fit.