TLDRocket
Sign in

Perplexity Introduces Photon: A Rust-Based Retrieval Engine That Cuts p99 Latency From 800 ms to 65 ms

MarkTechPost Asif Razzaq

Perplexity built Photon, a Rust search engine that slashed p99 latency from about 800 ms to about 65 ms. It also powers a cheaper Fast Search mode for agents, but the engine itself isn’t open source.

Based on reporting by MarkTechPost, Asif Razzaq — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Perplexity has swapped out the open-source retrieval engine it had forked and put a homegrown one in its place. The new system is called Photon, it’s written in Rust, and it now handles retrieval and ranking for all production traffic. It also sits behind a new Fast Search mode in the Perplexity Search API.

The company’s main complaint about the old setup was simple: it could not keep up as the index grew. Tail latency was the first pain point, with production p99 hovering near 800 ms. The dataset had outgrown RAM, so mlock was off the table. Cold reads caused major page faults. Merge periods were worse, with p99 climbing to about 1.2 s for 10 to 15 minutes while disk indexes fused. Recovery was slow too — bringing up and syncing an extra cluster could take more than a week, and that process also increased partial responses.

Photon attacks the problem at several layers. Requests go through a load balancer to a broker, which fans out to shard groups and watches for timeouts. Shards do retrieval, initial ranking, and second-stage ranking. The broker then merges candidates and pulls key document fields. Under the hood, Photon uses adaptive posting lists, a WAND-like budgeted traversal scheme, compact docblob records, and batched async reads through io_uring. Index building and serving are also separated, with versioned shard indexes built on dedicated nodes and rolled into service one group at a time.

The result is a sharp drop in retrieval and ranking latency: from about 800 ms to about 65 ms at p99. Perplexity says Photon runs on about 20% fewer serving machines than the old content nodes and stores about 2.5x as much data per document. It also says that pinning the same dataset with mlock would require an estimated 4.6x more resident memory than Photon uses today. Index version switches no longer trigger latency spikes, and a full web index now builds in a single-digit number of hours.

Fast Search is the commercial face of all this. It’s a hosted API feature, not something that can be self-hosted, and it costs $1 per 1,000 requests when search_type is set to fast on POST /search. Perplexity says it tested the mode across 3,554 tasks on six benchmarks, where Fast scored 64.3% at an estimated model-plus-search cost of $59.73. The default preset scored 64.0% at $187.60, so Fast came out much cheaper, with some loss in broader search quality.

Perplexity is doing the sensible thing here: build the boring infrastructure properly, then sell the speed as an API knob. That’s the pattern. The part the industry keeps forgetting is that search systems live or die on latency spikes, not just average numbers on a slide.

My take — AI-written commentary, not fact-checked reporting

This is the rare AI infra story that actually sounds like engineering, not confetti. A Rust rewrite, fewer machines, better p99s, and a paid fast path: that’s a business model with a pulse. The open-source fork era was always going to end in someone paying the maintenance bill.

Read more about this at: MarkTechPost

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.