TLDRocket
Sign in

Perplexity Trains Its Computer Agent on Real Mistakes With Hint-Guided Self-Distillation

MarkTechPost Michal Sutter

Perplexity Research trained a model in its Computer agent using real user sessions that included failures, applying rejection-sampling fine-tuning plus hint-guided self-distillation. In a live A/B test, tool-call failures fell from 2.24% to 1.77% across two trained checkpoints, a reported 21.2% relative reduction. The improved checkpoint runs only as an internal option inside Perplexity Computer and the post-trained weights or training code were not released, so it is not directly deployable elsewhere.

Why it matters

Perplexity Research published a new post-training study. It trains a model inside Perplexity Computer on real user sessions, including failed ones. The method pairs rejection sampling fine-tuning with hint-guided self-distillation. In a live A/B test, tool-call failures fell from 2.24% to 1.77% between 2 trained checkpoints. Perplexity team reports this as a statistically significant 21.2% […] The post Perplexity Trains Its Computer Agent on Real Mistakes With Hint-Guided Self-Distillation appeared first on MarkTechPost.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.