Perplexity GEO: Photon Replaces Its Third-Party Retrieval Engine

Key takeaways

Perplexity says its in-house Photon engine cut retrieval and ranking p99 latency from about 800 ms to 65 ms, and shipped a Fast Search tier at $1.00 per 1,000 requests.

Perplexity published an engineering account on September 24, 2026 of Photon, an in-house retrieval and ranking engine that now runs in its production search stack in place of a third-party open-source engine the company had adapted and maintained as its own fork. The same post introduced Fast Search, a lower-cost option on Perplexity's Search API, which the company's documentation prices at $1.00 per 1,000 requests against $5.00 for standard web search.

What was announced

Photon is a retrieval and ranking engine built inside Perplexity. The post states it "now powers our search pipeline" and describes it as the foundation of a new fast preset for the company's Search API.

It replaces an engine Perplexity did not name. The post describes the previous system as "a mature, general-purpose retrieval and ranking engine that met our initial needs," which the company then adapted to its architecture and maintained as a fork. Perplexity says it reached the point where "it would be simpler and cheaper to build from scratch rather than continuing to adapt and maintain our fork."

The company places Photon in a sequence of in-house infrastructure work, alongside CobbleDB, described as a distributed key-value store for serving page records, and its Search API. The conclusion states Perplexity is "steadily moving our search stack onto Rust-based, in-house infrastructure."

The post also states that "a small engineering team worked alongside a swarm of persistent coding agents to build Photon."

What changed

The post reports the following production measurements after migration, describing them as covering all stages inside Photon but excluding later stages in the higher-level search stack:

  • p99 response time in retrieval and ranking fell from roughly 800 ms to roughly 65 ms.
  • Photon uses approximately 20% fewer equivalent serving machines than the previous system's content nodes while storing approximately 2.5 times as much data per document.
  • Pinning the same dataset in RAM with mlock would require an estimated 4.6 times as much resident memory as Photon currently uses.
  • Perplexity says it "no longer observe[s] the latency spikes previously associated with switching index versions." In the old system, disk index fusion raised p99 latency to approximately 1.2 seconds for 10 to 15 minutes at a time.

Index building was moved off the serving cluster. Indexers now run as distributed jobs that produce versioned, ready-to-serve index structures, and serving nodes load them. Perplexity reports that the full web index can be built "in a single-digit number of hours," that an additional copy of the data can be deployed "in around half an hour," and that for verticals with high freshness requirements "document delivery takes a single-digit number of minutes." Under the previous architecture, deploying an additional cluster and synchronizing its data "could take more than a week."

On the API side, Fast Search is billed at $1.00 per 1,000 requests. Standard web search on the same API is billed at $5.00 per 1,000 requests.

Details and availability

Fast Search is enabled by setting search_type: "fast" on a POST /search request. Perplexity's documentation states that results use the same response format as standard web search, that max_results accepts values from 1 to 20, and that omitting search_type uses standard web search.

The documentation adds two SDK notes: Python library versions 0.43.4 and 0.43.5 require extra_body={"search_type": "fast"} because passing the value directly "fails the library's request validation," and the TypeScript SDK 0.38.5 requires the value cast with as any until its types include it. The documentation also notes that Fast Search is a search type on the Search API and not the Agent API fast preset.

Perplexity reports single-search-call latency for the fast preset of 160 ms at p50 and 230 ms at p95. It evaluated the preset on six agentic benchmarks: WideSearch, BrowseComp, DSQA, FRAMES, SEAL-0 and SEAL-Hard. On 3,554 selected tasks, the fast preset scored 64.3% at an estimated $59.73, against 64.0% at $187.60 for the default preset, which the post describes as approximately 68% lower estimated cost for the same number of tasks.

The post also states a trade-off. On Perplexity's internal benchmarks covering long-tail queries, broad coverage and result diversity, the fast preset "scores 0.24 points lower on relevance and approximately 3 percentage points lower on answer availability." An accompanying figure gives relevance falling from 2.45 to 2.21 and answer availability falling from 0.596 to 0.567. Perplexity recommends the fast preset for "day-to-day agentic tasks" and the default preset "for more challenging or ambiguous questions that are less time-sensitive."

Context

A retrieval and ranking engine sits between a query and the model that answers it. For each query it searches a large distributed index to select and order a small set of relevant pages within the stack's latency budget, and the model grounds its answer in those pages. Perplexity describes the engine as one of three constraints it had to solve for: cost, tail latency, and node recovery time.

That stage is where a page either enters an answer or does not. The changes Perplexity documents are attached to it: the engine itself was replaced, the amount of data stored per document grew by roughly 2.5 times, and a second API tier now accepts lower relevance and lower answer availability in exchange for lower latency and cost.

Perplexity frames the rebuild as a control decision. The conclusion states Photon "gives us full control over data representations, how reads are organized, and how new indexes are deployed."

The post carries a caveat on its own latency comparison. A figure caption states that "the reported percentiles and measurement setups are not a controlled like-for-like comparison," noting that Exa Instant reports p90 rather than p95 for its high-tail value and that Parallel's high-tail values are not published.

What has not been confirmed

Perplexity has not said what the additional per-document data contains. The post states that Photon stores approximately 2.5 times as much data per document and that "after the migration, we used this additional data to improve ranking quality beyond that of the old system," but it does not describe the fields or signals involved.

The post does not say whether Fast Search will become the default for the Search API. The documentation states that standard web search remains the default when search_type is omitted.

The post does not address publisher outcomes. It contains no statement about citation behavior, referral traffic, or how the ranking changes affect pages outside Perplexity's own benchmarks.

Perplexity did not name the third-party engine that Photon replaced.

The 68% figure is an estimated model-plus-search cost per task measured across six agentic benchmarks. It is not a per-request price. The per-request difference between the two tiers is documented separately, at $1.00 and $5.00 per 1,000 requests.

Sources

Perplexity Engineering. "Photon: Building a retrieval and ranking engine from scratch." Perplexity, September 24, 2026. https://www.perplexity.ai/hub/blog/photon

Perplexity. "Fast Search." Perplexity API documentation, accessed September 28, 2026. https://docs.perplexity.ai/docs/search/fast-search

Explore this topic

Keep following the same growth thread