Perplexity's Photon: A Rust-Based Retrieval Engine Cutting p99

Perplexity Introduces Photon: A Rust-Based Retrieval Engine That Cuts p99 Latency From 800 ms to 65 ms


Perplexity has released Photon, an in-house retrieval and ranking engine written in Rust. It replaces an open-source engine that Perplexity had previously forked for its AI-native search stack. Photon now handles retrieval and ranking for all production traffic and also powers a new Fast Search mode in the Perplexity Search API. According to Perplexity, the engine delivers single-call latency of 160 ms at p50 and 230 ms at p95.


Is it deployable? Yes, as a hosted API. Set search_type: "fast" on POST /search and pay $1 per 1,000 requests. Photon itself is not open source, so the engine cannot be self-hosted.


Why Perplexity Replaced Its Old Engine


The previous engine hit three limits as the index grew:


  • Tail latency: Production p99 sat near 800 ms. The dataset exceeded available RAM, so mlock was not an option. Cold reads triggered major page faults that stalled queries.
  • Merge spikes: During disk index fusion, p99 climbed to roughly 1.2 seconds for 10 to 15 minutes at a time.
  • Slow recovery: Deploying and syncing an extra cluster could take more than a week. Recovery also increased the share of partial responses.

The Perplexity team concluded that building from scratch was simpler and cheaper than maintaining its fork.


How Photon Works


A load balancer routes each request to a Photon broker. The broker fans out to a shard group and monitors for timeouts. Each shard runs retrieval, initial ranking, and second-stage ranking. The broker then aggregates and returns the results.

via MarkTechPost

Related