Perplexity Introduces Photon: A Rust-Based Retrieval Engine That Cuts p99 Latency From 800 ms to 65 ms
MarkTechPost Asif Razzaq
Perplexity built Photon, a Rust search engine that slashed p99 latency from about 800 ms to about 65 ms. It also powers a cheaper Fast Search mode for agents, but the engine itself isn’t open source.
Based on reporting by MarkTechPost, Asif Razzaq — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Perplexity has swapped out the open-source retrieval engine it had forked and put a homegrown one in its place. The new system is called Photon, it’s written in Rust, and it now handles retrieval and ranking for all production traffic. It also sits behind a new Fast Search mode in the Perplexity Search API.
The company’s main complaint about the old setup was simple: it could not keep up as the index grew. Tail latency was the first pain point, with production p99 hovering near 800 ms. The dataset had outgrown RAM, so mlock was off the table. Cold reads caused major page faults. Merge periods were worse, with p99 climbing to about 1.2 s for 10 to 15 minutes while disk indexes fused. Recovery was slow too — bringing up and syncing an extra cluster could take more than a week, and that process also increased partial responses.
Photon attacks the problem at several layers. Requests go through a load balancer to a broker, which fans out to shard groups and watches for timeouts. Shards do retrieval, initial ranking, and second-stage ranking. The broker then merges candidates and pulls key document fields. Under the hood, Photon uses adaptive posting lists, a WAND-like budgeted traversal scheme, compact docblob records, and batched async reads through io_uring. Index building and serving are also separated, with versioned shard indexes built on dedicated nodes and rolled into service one group at a time.
The result is a sharp drop in retrieval and ranking latency: from about 800 ms to about 65 ms at p99. Perplexity says Photon runs on about 20% fewer serving machines than the old content nodes and stores about 2.5x as much data per document. It also says that pinning the same dataset with mlock would require an estimated 4.6x more resident memory than Photon uses today. Index version switches no longer trigger latency spikes, and a full web index now builds in a single-digit number of hours.
Fast Search is the commercial face of all this. It’s a hosted API feature, not something that can be self-hosted, and it costs $1 per 1,000 requests when search_type is set to fast on POST /search. Perplexity says it tested the mode across 3,554 tasks on six benchmarks, where Fast scored 64.3% at an estimated model-plus-search cost of $59.73. The default preset scored 64.0% at $187.60, so Fast came out much cheaper, with some loss in broader search quality.
Perplexity is doing the sensible thing here: build the boring infrastructure properly, then sell the speed as an API knob. That’s the pattern. The part the industry keeps forgetting is that search systems live or die on latency spikes, not just average numbers on a slide.
My take — AI-written commentary, not fact-checked reporting
This is the rare AI infra story that actually sounds like engineering, not confetti. A Rust rewrite, fewer machines, better p99s, and a paid fast path: that’s a business model with a pulse. The open-source fork era was always going to end in someone paying the maintenance bill.
Read more about this at: MarkTechPost
Related stories
Perplexity Releases pplx, a Single-Binary CLI That Puts Its Search API in the Terminal for Coding Agents
MarkTechPost · 2 months ago ·
24
Perplexity Open Sources Lily: A Rust + Metal Inference Engine for Qwen3.6-35B-A3B on Apple Silicon
MarkTechPost · 3 weeks ago ·
49