RISED: Rubrics for Agentic Multi-Environment Selection and Self-Distillation
Apple Machine Learning Research
RISED trains a single LLM agent across multiple interactive environments using rubric-based rollout tagging to select training data and to add token-level supervision via self-distillation. In evaluation across model backbones, RISED achieves the highest mean pass rate across environments and ranks first or second in every individual environment. As a result, learning is guided by richer within- and cross-environment behavioural feedback instead of only scalar rewards, improving pass rates across all tested environments.
Why it matters
Training a single LLM agent jointly across diverse interactive environments has attracted increasing attention as a route to generalist agents. Existing curriculum and data-selection strategies often allocate training at the environment level or prioritize local reward-based signals, without explicitly considering relationships between current rollouts across environments for prompt-group selection. Meanwhile, as environments are learned at different rates, all-failure and all-success rollout groups can coexist within a batch, leaving those data without group-relative reward signals. Both…