Perplexity Introduces Photon: A Rust-Based Retrieval Engine That Cuts p99 Latency From 800 ms to 65 ms
Perplexity has released Photon, an in-house retrieval and ranking engine written in Rust. It replaces an open-source engine that Perplexity had previously forked for its AI-native search stack. Photon now handles retrieval and ranking for all production traffic and also powers a new Fast Search mode in the Perplexity Search API. According to Perplexity, the engine delivers single-call latency of 160 ms at p50 and 230 ms at p95.
Is it deployable? Yes, as a hosted API. Set search_type: "fast" on POST /search and pay $1 per 1,000 requests. Photon itself is not open source, so the engine cannot be self-hosted.
Why Perplexity Replaced Its Old Engine
The previous engine hit three limits as the index grew:
- Tail latency: Production p99 sat near 800 ms. The dataset exceeded available RAM, so
mlockwas not an option. Cold reads triggered major page faults that stalled queries. - Merge spikes: During disk index fusion, p99 climbed to roughly 1.2 seconds for 10 to 15 minutes at a time.
- Slow recovery: Deploying and syncing an extra cluster could take more than a week. Recovery also increased the share of partial responses.
The Perplexity team concluded that building from scratch was simpler and cheaper than maintaining its fork.
How Photon Works
A load balancer routes each request to a Photon broker. The broker fans out to a shard group and monitors for timeouts. Each shard runs retrieval, initial ranking, and second-stage ranking. The broker then aggregates and returns the results.
via MarkTechPost
