Perplexity Trains Its Computer Agent on Real Mistakes With Hint-Guided Self-Distillation
MarkTechPost Michal Sutter
Perplexity Research trained a model in its Computer agent using real user sessions that included failures, applying rejection-sampling fine-tuning plus hint-guided self-distillation. In a live A/B test, tool-call failures fell from 2.24% to 1.77% across two trained checkpoints, a reported 21.2% relative reduction. The improved checkpoint runs only as an internal option inside Perplexity Computer and the post-trained weights or training code were not released, so it is not directly deployable elsewhere.
Why it matters
Perplexity Research published a new post-training study. It trains a model inside Perplexity Computer on real user sessions, including failed ones. The method pairs rejection sampling fine-tuning with hint-guided self-distillation. In a live A/B test, tool-call failures fell from 2.24% to 1.77% between 2 trained checkpoints. Perplexity team reports this as a statistically significant 21.2% […] The post Perplexity Trains Its Computer Agent on Real Mistakes With Hint-Guided Self-Distillation appeared first on MarkTechPost.
Related stories
Perplexity Releases Hybrid Compute on Mac: Cloud Agents Orchestrate Down to a Local Model, Gated On Device
MarkTechPost · 3 weeks ago ·
32
Embarrassingly Simple Self-Distillation Improves Code Generation
Apple · 2 months ago ·
50