Small models, big results: Achieving superior intent extraction through decomposition
Google Research
Google Research found a trick to make small on-device AI models understand what you're trying to do on your phone. No cloud needed, and it works almost as well as much bigger models.
Based on reporting by Google Research — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Google's latest research tackles a problem that sounds simple but isn't: getting a phone-sized AI model to figure out what you're actually trying to accomplish as you tap through apps. Big multimodal models already do this reasonably well, but shipping every screenshot and tap sequence to a server is slow, expensive, and a privacy headache nobody wants when the task involves, say, your bank app or a doctor's appointment.
The fix, described in a paper presented at EMNLP 2025, is less about brute-forcing a bigger model onto the phone and more about splitting the job into two bite-sized pieces. First, a small multimodal LLM looks at three screens at a time — the one before, the current one, and the next — and writes a short summary of what's on screen, what the user just did, and a guess at why. Then a second, fine-tuned small model reads that whole sequence of summaries and distills it into one sentence describing the overall intent. Instead of asking a tiny model to hold an entire trajectory in its head and reason about it at once, you let it focus on one screen at a time and hand off the synthesis to a specialist step.
A few engineering choices made this actually work. The team scrubbed training intents of any detail that wasn't already present in the summaries, because otherwise the model learned to invent plausible-sounding specifics that weren't there — a tidy way of training hallucination out rather than into the system. They also found that asking the first-stage model to speculate about user intent improved its screen summaries, even though that speculation gets thrown away before the second stage sees it. Counterintuitive, but it seems the act of speculating forces better attention to relevant details.
To judge whether any of this actually worked, the researchers built an evaluation method called Bi-Fact, which breaks predicted and reference intents down into atomic facts — something like "a one-way flight" versus "a flight from London to Kigali," which counts as two facts — and checks which ones show up on both sides. That gives them real precision and recall numbers instead of vague vibes about whether an answer "looks right," and it let them trace exactly where errors crept in across the two stages.
The payoff: on both mobile and web interaction data, and across Gemini and Qwen2 base models, the decomposed approach beat both chain-of-thought prompting and standard end-to-end fine-tuning. The headline number is that Gemini 1.5 Flash 8B, running the decomposed method, matched the performance of the much larger Gemini 1.5 Pro — at a fraction of the compute cost and latency. That's the kind of result that makes on-device assistive features, the ones that quietly notice you're planning a London trip and offer festival dates without ever phoning home, look a lot more plausible than they did a year ago.
My take — AI-written commentary, not fact-checked reporting
This is the unglamorous but genuinely important side of AI progress — not a bigger model, just a smarter way to use a small one, and it's exactly the kind of work that makes on-device AI viable instead of a marketing slide. I'd rather my phone quietly infer intent locally than ship my screen taps to a data center, and if Google keeps proving small models can match Pro-tier performance on tasks like this, the case for cloud-dependent assistants gets weaker every quarter.
Read more about this at: Google Research