Perplexity and MarkTechPost announce a partnership
Partnership Disputed 10% confidence first seen
Decision brief
- What changed
- Perplexity open sourced Lily, a single-process Rust + Metal inference engine for running Qwen3.6-35B-A3B locally on Apple Silicon through an OpenAI-compatible chat-completions API. In the cited test on a 40-core, 128 GB M5 Max, Perplexity reported higher prefill and decode throughput than MLX-LM.
- Why it matters
- For leaders evaluating on-device or Apple-hardware AI deployment, this suggests a more specialized inference path that could improve performance for a specific model and hardware stack without relying on PyTorch or MLX at runtime. That can affect infrastructure choices, developer tooling, and cost/performance tradeoffs for teams standardizing on Apple Silicon, although the reported gains are currently tied to one model, one runtime design, and vendor-reported benchmarks.
- Evidence
- The only provided coverage is a MarkTechPost article reporting that Perplexity open sourced Lily and citing Perplexity's benchmark figures against MLX-LM. No additional independent reporting or corroborating benchmark validation was included in the provided materials.
- What remains uncertain
- The provided coverage does not substantiate the stated 'Perplexity and MarkTechPost announce a partnership' event; it supports an open-source release covered by MarkTechPost, not a verified partnership announcement. It is also unclear how Lily performs across other models, workloads, and Apple devices, and whether the reported benchmark advantages translate into production reliability, maintainability, or total cost benefits.
- Monitor next
- Watch for independent benchmark replication or a formal Perplexity announcement clarifying whether there is an actual partnership beyond media coverage of the Lily release.
Analytical support, not advice — assumptions and open questions stated above.