Async GRPO with LoRA across HF Jobs: a bucket, a proxy, and no NCCL
TRL's AsyncGRPOTrainer now supports training a LoRA adapter and syncing only that adapter to vLLM, shipping in TRL v1.14 via PR #7017. Hugging Face detailed a project built on this in which the trainer and vLLM replicas run as separate Hugging Face Jobs on separate machines, with no NCCL connection between them.
A rank-1 adapter for a 1.5B model is a few megabytes, versus roughly 3 GB for the full model, so the adapter can travel through a Storage Bucket mounted as a FUSE filesystem in every Job at the same path rather than over NCCL. The post cites Thinking Machines's "LoRA Without Regret" as showing LoRA can match full fine-tuning for policy-gradient RL, even at rank 1.
The setup uses three Jobs: a trainer running AsyncGRPOTrainer with LoRA and FSDP, and two vLLM Jobs each serving the base model plus the last published adapter. Every few optimizer steps the trainer saves the adapter under an output directory, publishes it with an atomic rename, and sends the path to vLLM's /v1/load_lora_adapter endpoint. A small proxy on the trainer Job at 127.0.0.1:8000 adds the Authorization header required by exposed Job ports, routes each rollout to the replica likely holding its KV prefix, and broadcasts state-changing requests such as adapter loads, pause, and resume to all replicas.
With max_staleness=4, the trainer keeps max_staleness + 1 adapter versions registered, giving --max-loras 6. Adapters use versioned names because vLLM keys its prefix cache by adapter name. The dataset is sail/Sanity-Test-R1D-1.5B from "Defeating the Training-Inference Mismatch via FP16" (Qi et al., 2025), with 1,460 MATH questions and hyperparameters from the paper's LoRA scripts. Five runs took the same recipe from 3 h 27 min to 53 min for 500 steps, according to the post.
Based on reporting from the original publisher. Visit the source for full context and later updates.