Optimizing On-Device Inference for Apple Silicon
Perplexity describes Lily, a lightweight local inference engine built for Apple silicon and Qwen3.6-35B-A3B, with separate optimizations for prefill…
Perplexity has built Lily, a lightweight local inference engine for running Qwen3.6-35B-A3B on Apple silicon, with separate optimizations for the prefill and decode stages.
The announcement walks through how Qwen's architecture opens up model-specific optimizations on Apple silicon, where further tuning stops paying off, and gives an end-to-end comparison against MLX-LM. Lily is to be open-sourced.
Perplexity
AI-powered answer engine — ask anything and get cited, real-time answers from across the web.
View Perplexity →