AivexaNewsSearch
AI news for builders and product teamsChecked every hour

How to Train Really Large Models on Many GPUs?

Collected Oct 1, 2026

Lilian Weng published an article on training large and deep neural networks across multiple GPUs, addressing the challenge that large models demand more GPU memory and training time than a single GPU can provide. The post was updated on 2022-03-13 to add expert choice routing, and on 2022-06-10 Weng noted that she and Greg wrote a shorter, upgraded version published on the OpenAI Blog titled "Techniques for Training Large Neural Networks."

The article states that the main bottleneck for training very large models is intense demand for GPU memory, exceeding what one GPU machine can host. Beyond model weights, storing intermediate outputs such as gradients and optimizer states is usually even more expensive. Large models often pair with large training corpora, so a single process may take too long. Parallelism is therefore necessary and can occur along data, model architecture, and tensor operation dimensions.

The post covers data parallelism, including naive DP, offloading unused parameters to CPU in methods like GeePS, and synchronization approaches: bulk synchronous parallels, asynchronous parallel, and gradient accumulation as in Distributed Data Parallel since PyTorch v1.5. Model parallelism partitions computation and parameters across machines. Pipeline parallelism combines model and data parallelism to reduce idle time bubbles, with GPipe aggregating gradients synchronously and PipeDream using 1F1B scheduling with weight stashing and optional vertical sync. Variations PipeDream-flush and PipeDream-2BW reduce memory by maintaining one or two weight versions.

Tensor parallelism horizontally partitions tensor operations, illustrated by Megatron-LM for transformer MLP and self-attention. PTD-P combines pipeline, tensor, and data parallelism with interleaved scheduling. The article also covers Mixture-of-Experts, including sparsely gated MoE layers, noisy top-k gating, auxiliary importance loss, GShard, and Switch Transformer, plus memory-saving designs such as CPU offloading, activation recomputation, mixed precision, compression, and memory-efficient optimizers.

Read at Lilian Weng

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

[Updated on 2022-03-13: add expert choice routing .] [Updated on 2022-06-10]: Greg and I wrote a shorted and upgraded version of this post, published on OpenAI Blog: “Techniques for Training Large Neural Networks”