Apple in talks with startup that shrinks AI models to run on an iPhone
CNBC
Apple's talking to a startup, PrismML, that claims it can squeeze huge AI models down small enough to run on an iPhone. If it works, Siri gets smarter without your data ever leaving the phone.
Based on reporting by CNBC — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Babak Hassibi, CEO of PrismML, told CNBC his Caltech spinout is in early talks with Apple, which is testing the startup's compressed AI models for speed, energy use and on-device performance. Hassibi called the conversations very early but said things are progressing nicely. Apple hasn't commented.
The timing is notable. PrismML released its compressed versions of Alibaba's Qwen model just a day after Apple opened the public beta of iOS 27, which finally brings its overhauled Siri to a wide audience. PrismML says it took Qwen's 27 billion parameters, originally about 54 GB, and shrank the whole thing to under 4 GB, small enough to run on an iPhone 15 or newer. The trick, according to Hassibi, is storing each internal value with far less precision than usual, similar in spirit to how chipmakers moved from eight-bit to four-bit computing, but pushed further.
The payoff, PrismML claims, is models that use 10 to 15 times less memory, run six to eight times faster, and burn three to six times less energy than standard versions on the same hardware. There's a catch: Hassibi admits some performance is lost, with factual recall degrading faster than reasoning, math or coding skills. The company is giving away two compressed versions for free, aimed at iPhones, MacBooks and Nvidia PCs, and says Google's Gemma is next in line, followed eventually by larger frontier-lab models that currently need datacenter-grade hardware.
For Apple, the appeal is obvious. Running more AI locally cuts the lag of calling out to a server, trims cloud costs, keeps personal data on the device, and works offline. Analysts like Creative Strategies' Carolina Milanesi point to health and fitness data as exactly the kind of sensitive information people want processed without leaving their phone. Horace Dediu of Asymco frames it less as a memory problem and more as a puzzle: how big and how capable a model Apple can cram into the same hardware constraints, something its control over both chips and software should help with.
But plenty of people are waiting to see if the claims survive contact with reality. Counterpoint's Tarun Pathak wants to see how the models handle long prompts, multitasking and millions of real-world queries before calling it proven. IDC's Phil Solis worries about battery drain if these models end up running constantly in the background for agent-style tasks. And there's a bigger question hanging over the whole industry: if models need less memory, does that mean less chip demand overall? D.A. Davidson's Gil Luria doesn't think so, arguing the same chips just shift from datacenters into phones and other devices, and that efficiency gains often just spark more usage rather than less spending. Memory costs are already climbing sharply, with Morgan Stanley projecting Apple's per-bit DRAM costs could jump roughly 190% year over year by fiscal 2027, which is part of why the stakes here go well beyond one startup's demo.
My take — AI-written commentary, not fact-checked reporting
Everyone loves to say smaller models mean less chip demand, but that's wishful thinking dressed up as an efficiency story. History says cheaper, faster AI just gets used more, not less, and Apple squeezing Qwen onto an iPhone doesn't shrink the world's appetite for GPUs, it just relocates some of them into people's pockets. The real test isn't the demo, it's whether PrismML's numbers hold up across millions of messy real-world queries, and that's a much harder thing to fake than a CNBC interview.
Read more about this at: CNBC