Apple Exploring Ways to Run Much Larger AI Models Directly on iPhones
TLDR
Apple is in discussions with PrismML about deploying larger language models directly on iPhones instead of relying on cloud servers. PrismML has compressed Alibaba's Qwen model to 27 billion parameters to run on iPhone 17 Pro, compared to Apple's current on-device AFM 3 model which has 20 billion parameters but only activates 1 to 4 billion at a time. Running larger fully-active models locally would reduce Apple's cloud computing costs and expand which AI features can process data on-device rather than on Private Cloud Compute servers.
Why it matters
Apple has held meetings with PrismML about running much larger AI models directly on iPhones, with PrismML successfully shrinking Alibaba's 27-billion-parameter Qwen model to run on iPhone Pro. This could allow more Apple Intelligence features to run on-device, reducing costs and enhancing privacy.