AivexaNewsSearch
AI news for builders and product teamsChecked every hour
AWS Machine Learning BlogFirst partyDeveloper tools

Accelerate multimodal RL training with SkyRL on Amazon SageMaker HyperPod

Collected Sep 30, 2026

AWS published a walkthrough on running SkyRL, an open-source reinforcement learning framework, on Amazon SageMaker HyperPod to train a Qwen3-VL-8B vision-language model to navigate visual mazes using Group Relative Policy Optimization (GRPO).

Starting from the VisGym SFT checkpoint, a supervised fine-tuning starting point, AWS reports that GRPO post-training on HyperPod improved the maze solve rate from 43.75% to more than 95% on a fixed 64-maze evaluation set. In the described run, the model reached a 75% solve rate around step 100 and peaked at 96.875% (62/64 mazes) at step 160, compared with a baseline of 43.75% (28/64 mazes) before post-training. AWS notes results will vary based on hyperparameters and maze configuration.

The post says the setup requires a SageMaker HyperPod cluster with Amazon EKS orchestration, at least 3 ml.g7e.12xlarge instances and one ml.r5d.16xlarge instance, the KubeRay operator, the HyperPod Observability EKS add-on, the HyperPod Ray Endpoint Operator, and the Amazon FSx for Lustre CSI driver, plus a SageMaker Studio domain and the toolkit-for-ray-on-sagemaker-ai Python package.

The described topology uses three GPU worker nodes and a CPU head node, with SkyRL colocating inference and training on the same GPUs: vLLM engines generate rollouts while a policy model sharded with Fully Sharded Data Parallel handles gradient updates. Updated LoRA adapter weights sync to inference engines through Amazon FSx for Lustre shared storage. The head node consolidates LoRA adapter shards at each checkpoint save.

HyperPod cluster resiliency continuously monitors node health and automatically replaces faulty nodes, and checkpointing lets a job resume from its last saved step. According to the post, both the Ray Dashboard for job-level visibility and Amazon Managed Grafana for infrastructure and training metrics are accessible from the Tasks tab in SageMaker Studio, with the HyperPod Observability EKS add-on provisioning four pre-built Ray dashboards. The post also covers hosting the resulting LoRA adapter for inference with Ray Serve and an OpenAI-compatible endpoint.

Read at AWS Machine Learning Blog

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

Learn how to run SkyRL, an open-source reinforcement learning framework, on Amazon SageMaker HyperPod to post-train a Qwen3-VL-8B vision-language model with GRPO. This walkthrough covers building the container image, launching a Ray cluster from SageMaker Studio, submitting and monitoring the job, and hosting the trained LoRA adapter for inference.