RISED: Rubrics for Agentic Multi-Environment Selection and Self-Distillation

Researchers at Apple, with co-authors affiliated with the National University of Singapore, introduced RISED, a method for training a single LLM agent jointly across diverse interactive environments. The work is described in a publication titled "RISED: Rubrics for Agentic Multi-Environment Selection and Self-Distillation."
The researchers state that existing curriculum and data-selection strategies often allocate training at the environment level or prioritize local reward-based signals, without explicitly considering relationships between current rollouts across environments for prompt-group selection. They also note that because environments are learned at different rates, all-failure and all-success rollout groups can coexist within a batch, leaving that data without group-relative reward signals. According to the work, these challenges highlight limitations of relying solely on scalar rewards in multi-environment reinforcement learning: limited information about cross-environment relationships and no within-group reward contrast when rewards are identical.
RISED repurposes rubrics beyond their use as reward, applying them to guide online data selection and policy supervision. An LLM judge tags each rollout using a predefined rubric vocabulary shared across environments. The resulting profiles guide selection of data that aligns with the overall behavioural composition of the mixed-environment batch while limiting overlap with already-selected data. Positive rubrics, describing desired behaviours, provide privileged context for an on-policy self-distillation teacher, supplying additional token-level supervision. Negative rubrics, describing undesired behaviours, guide subsequent rollout generation away from recurring failure modes.
The researchers report that across model backbones, RISED achieves the highest mean pass rate across environments and ranks first or second in every individual environment. They add that rubric-based analysis of RISED can characterize the behavioural changes accompanying these gains. The listed authors are Jingtan Wang, Sirajul Salekin, Young mok Jung, Javier Movellan, Bryan Kian Hsiang Low, and Manjot Bilkhu. The work was done while Jingtan Wang was at Apple.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
Training a single LLM agent jointly across diverse interactive environments has attracted increasing attention as a route to generalist agents. Existing curriculum and data-selection strategies often allocate training at the environment level or prioritize local reward-based signals, without explicitly considering relationships between current rollouts across environments for prompt-group selection. Meanwhile, as environments are learned at different rates, all-failure and all-success rollout groups can coexist wit