TLDRocket
Sign in

RISED: Rubrics for Agentic Multi-Environment Selection and Self-Distillation

Apple Machine Learning Research

RISED trains a single LLM agent across multiple interactive environments using rubric-based rollout tagging to select training data and to add token-level supervision via self-distillation. In evaluation across model backbones, RISED achieves the highest mean pass rate across environments and ranks first or second in every individual environment. As a result, learning is guided by richer within- and cross-environment behavioural feedback instead of only scalar rewards, improving pass rates across all tested environments.

Why it matters

Training a single LLM agent jointly across diverse interactive environments has attracted increasing attention as a route to generalist agents. Existing curriculum and data-selection strategies often allocate training at the environment level or prioritize local reward-based signals, without explicitly considering relationships between current rollouts across environments for prompt-group selection. Meanwhile, as environments are learned at different rates, all-failure and all-success rollout groups can coexist within a batch, leaving those data without group-relative reward signals. Both…

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.