SCLATE: A Substrate for Continual-Learning Agent Training and Evaluation
Apple Machine Learning Research
Apple researchers built SCLATE to train and test continual-learning agents on one shared schedule. It can shrink a month-long run into hours and shows extra memory isn’t always better.
Based on reporting by Apple Machine Learning Research — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Apple’s ML Research team has put a name to a mess that’s been hiding in plain sight: if you want to train or evaluate a continual-learning agent, you usually end up wiring together a custom event loop for every benchmark and every agent. That’s because the benchmarks handle their own timing, while the agent has its own life to live — sessions start and stop, cron-like events fire, memory gets consolidated, and all of it has to stay in sync over long horizons.
SCLATE is the company’s answer. It is an execution substrate where a benchmark and an unmodified agent both contribute events to a single scheduler through an adapter. A hybrid simulated clock keeps the whole thing moving on one timeline, running in real time when the agent is busy and skipping dead air when nothing is happening. That makes a month-long scenario finish in hours instead of dragging on forever.
The system is also built as a rollout engine, so it can run an agent’s harness and memory without changing them. To make the runs auditable, it records the tokens and log probabilities from every model call through an in-container proxy. Apple says it ported seven benchmarks to SCLATE and compared ten unmodified harness and memory setups across ten models.
The results push back on a comforting assumption. Adding a memory system did not reliably outperform the harness’s built-in memory, and the same harness-memory setup could be used very differently by different models. That matters, because a lot of agent work still treats “better memory” as a default win. SCLATE makes that claim harder to get away with.
Apple also used the setup to post-train Qwen3.5-4B through unmodified harnesses and memory systems. After that, the model read 6.8× fewer file lines, gained 16.7 points on SWE-bench Verified pass rate, wrote richer memory records, and improved held-out MetaClaw accuracy by up to 11.8 points. The broader point is simple: if you want to know whether an agent is actually learning, you need a testbed that can watch the whole machine, not just the final answer.
My take — AI-written commentary, not fact-checked reporting
This is the right kind of unglamorous infrastructure work: make the timing honest, then see which tricks were carrying their own weight. The industry loves slapping “memory” on a slide deck and calling it progress; SCLATE suggests the bar should be a lot meaner. Open evaluation beats wishful thinking, every time.
Read more about this at: Apple Machine Learning Research
Related stories
Contrastive-LM Releases CLM-8B: An Open System One Model That Scores Agent Actions Up to 9× Faster Than Jev
MarkTechPost · 2 weeks ago ·
31
The Sequence Knowledge - Issue 945: Learning RSI: Agents that Rewrite their Own Scaffolding
TheSequence · 1 day ago ·
18
UC Berkeley Researchers Release CUA-Lite, an Open Platform Unifying Sandboxes, Data, Evaluation and RL for Computer-Use Agents
MarkTechPost · 1 month ago ·
8