tech
Perplexity splits AI inference between PCs and cloud to cut costs
Perplexity AI announced a platform at Computex that dynamically routes AI inference between PCs and cloud servers in real time, acting as an “air-traffic controller” for AI tasks. The chip-agnostic system targets the cost crisis of centralised inference as Perplexity’s revenue hits $500 million.

TL;DR
- Perplexity AI developed a platform that dynamically splits AI workloads between PCs and cloud servers.
- The system acts as an "air-traffic controller for AI tasks" to reduce inference costs.
- Simple tasks run locally on PCs, while complex tasks are routed to cloud servers.
- This hybrid approach offloads inference work to existing PCs, reducing strain on data centers.
- The platform is "chip agnostic" and works with processors from Intel and Nvidia.
- Perplexity's revenue grew fivefold to $500 million, with significant growth per employee added.
- By using user hardware, Perplexity can reduce marginal cost per query and improve response latency.