Scaling MoE reinforcement learning on Amazon EKS with EFA and DeepEP with 40% more throughput
AWS detailed an architecture for post-training Mixture-of-Experts (MoE) models with reinforcement learning at scale, combining Amazon Elastic Kubernetes Service (Amazon EKS), Elastic Fabric Adapter (EFA), and DeepEP. The post states that across 48 P5en instances—16 dedicated to training and 32 to inference—running a super-sparse MoE model, enabling DeepEP over EFA increased aggregate RL rollout throughput by 40 percent.
The post describes three interrelated challenges in large-scale RL training: balancing the competing resource demands of rollout generation and policy training, managing accelerator compute, memory, and network bandwidth simultaneously, and handling the shift from high-bandwidth intra-node communication to lower-bandwidth inter-node links as jobs scale beyond a single instance. Rollout generation is described as large-scale distributed inference focused on aggregate throughput, while policy training requires tightly coupled workers that progress in lockstep, where latency spikes or straggling workers can stall jobs or trigger NCCL timeouts.
In the described architecture, EKS manages lifecycle and placement of heterogeneous workers, EFA provides an inter-node data path for communication-intensive GPU workloads, and Amazon S3 stores datasets, model checkpoints, and completed training artifacts. The EKS cluster contains separate node groups: GPU instances for rollout generation, reward-model inference, and policy training; CPU instances for environments and preprocessing; and memory-optimized instances for experience buffers and checkpoint caches.
Amazon says it contributed features to migrate DeepEP's communication primitives to libfabric, making the transport layer portable across libfabric-supported fabrics and optimizing MoE training over EFA. With these changes, DeepEP v2 gains native EFA support, and NCCL 2.31 incorporates EFA optimizations for dense collective communication. DeepEP replaces standard NCCL all-to-all collectives with dispatch and combine GPU kernels, using NVLink for intra-node transfers and libfabric over EFA for inter-node transfers. On supported instance types such as P5 and P6, the post states EFA works with NVIDIA GPUDirect RDMA to transfer data directly between GPU memory buffers across instances, bypassing CPU and operating system.
The post also describes using Amazon EC2 Spot Instances for rollout generation, since interrupted rollout workers do not require the entire RL job to stop, while policy-training workers run on stable capacity. Listed prerequisites include Amazon EKS 1.31 or later, EFA installer 1.49 with the AWS OFI NCCL plugin, DeepEP 2.0.0, NCCL 2.31.2, SGLang 0.5.17, and PyTorch 2.12.1 (CUDA 13.0).
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
Learn how to scale Mixture-of-Experts (MoE) reinforcement learning on Amazon EKS using Elastic Fabric Adapter (EFA) and DeepEP. This post presents an architecture that combines Amazon EKS, EFA, and Amazon S3 and increased aggregate reinforcement learning rollout throughput by 40% for large-scale RLHF and GRPO training.