Perplexity has replaced the third-party platform behind its search with an engine it built itself, it said in an engineering post on 24 September. As its index and workloads grew it hit hard limits on cost, tail latency and node recovery time that it could not resolve within the platform it had been running.
The engine, Photon, uses compact data formats so each query reads and decodes only what it needs, adds batch-aware caching and asynchronous reads to overlap waits, and separates index building from, updating one serving group at a time and warming its caches before routing live queries to it.
Why it matters: retrieval decides what the model is allowed to know, so owning it changes the cost structure of every answer. Perplexity’s headline numbers are about cost rather than quality. Across six public benchmarks it reports its new fast preset cutting estimated model-plus-search cost per task by 68%, scoring 64.3% on aggregate task quality against 64.0% for its default preset.
The post is unusually candid about its own comparisons. The 160ms median and 230ms 95th-percentile latency figures are labelled self-reported, and it notes that one rival publishes a 90th-percentile figure rather than a 95th while another does not publish tail-latency figures at all.
