Perplexity Lily inference engine performance benchmark results displayed on a computer screen showing speed comparisons with Apple MLX

Perplexity Lily Launch Delivers 23% Faster Performance Than Apple MLX

Perplexity Lily launched Wednesday as an open source local inference engine for Apple silicon, delivering 23% higher average prefill throughput than Apple‘s (AAPL) MLX-LM library across ten prompt lengths on an M5 Max MacBook Pro.

Key Takeaways

  • Perplexity Lily launched Wednesday as an open source local inference engine built for Apple silicon
  • Lily delivered 23% higher average prefill throughput than Apple’s MLX-LM library across ten prompt lengths
  • The engine was built specifically to run the Qwen3.6-35B-A3B model, a mixture-of-experts model with 35 billion parameters
  • Perplexity said Lily was built for hybrid compute inside Perplexity Computer, treating Apple silicon as a distinct platform

Perplexity Lily Takes Aim At Apple’s Own Playbook

The company posted benchmark results on X, testing the engine across ten prompt lengths and decode contexts on Apple’s M5 Max chip. Perplexity wrote that the tool is now available to developers, built specifically to run the Qwen3.6-35B-A3B model on Apple hardware.

Rather than using a general runtime built for many processors, it maps Qwen’s operations directly onto the chip’s compute and memory architecture, the kind of hardware-specific tuning that drives real cost-per-inference gains at the silicon level.

Inside The Benchmark That Pits Lily Against MLX

Prefill throughput measures how fast a model processes a prompt’s existing text before it replies. Decode throughput measures how fast it produces each new word afterward.

Also Read: 21 Global Banks Plan First Joint Dollar Stablecoin for 2027

Qwen3.6-35B-A3B is a mixture-of-experts model holding 35 billion parameters but activating roughly 3 billion for any single query, allowing it to run on a laptop rather than a data center rack.

MLX-LM is Apple’s own open source library for running language models on M-series chips, built on the MLX framework the company released in 2023.

From Search Upstart To Local Hardware Contender

Perplexity built its reputation as an AI-powered answer engine challenging Google‘s (GOOGL) hold on web search, then expanded into a browser called Comet. The company said Lily was built for hybrid compute inside Perplexity Computer, treating Apple silicon as a distinct platform rather than a generic target, an approach that mirrors how Nvidia (NVDA) used CUDA to lock developers to its chips.

The Race To Own Every Chip Running Your AI

Perplexity did not disclose download numbers since release, nor whether Windows or Android versions are planned.

Apple has not responded publicly to the benchmark claims. As labs push smaller models onto consumer hardware, the contest over local inference throughput and power efficiency will tighten, and the engineering spend behind Lily signals Perplexity is treating on-device compute as a core capacity bet, not a side project.

Read Next: Astra Builds Breakthrough Game Prototypes With Half the Fix Time

Similar Posts