TLDRocket
Sign in

DeepAmbigQA: Ambiguous Multi-hop Questions for Benchmarking LLM Answer Completeness

Apple Machine Learning Research

Apple researchers built a nasty new benchmark for AI search agents: DeepAmbigQA. Even GPT-5 flunks it, scoring just 0.13 to 0.21 exact match.

Based on reporting by Apple Machine Learning Research — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Ask an AI assistant which actor from the film Heat won an Academy Award and you'd expect a clean, complete answer. But there's a catch buried in that question: multiple films share the title Heat, and pulling a correct answer means first figuring out which one you mean, then combing through a whole cast list to find every actor who happened to win an Oscar. That's the exact kind of question a team of researchers, including Jiabao Ji, Min Li, Priyanshu Kumar, Shiyu Chang and Saloni Potdar, set out to stress-test large language models with.

Their argument is that current open-domain QA benchmarks rarely force models to handle two problems at once: sorting out name ambiguity and gathering evidence across a large set of candidates to produce a full, not partial, answer. So they built DEEPAMBIGQAGEN, an automatic pipeline that generates questions grounded in text corpora and linked knowledge graphs. The pipeline is designed to produce questions that read naturally, can be verified, and deliberately bake in both ambiguous names and multi-step reasoning chains.

Out of that pipeline came DEEPAMBIGQA, a dataset of 3,600 questions. Half of them contain explicit name ambiguity that a model has to resolve before it can even start answering, and every single question demands multi-hop reasoning to reach a complete response.

The results are not flattering for even the strongest models available. GPT-5, described in the paper as state-of-the-art, managed only 0.13 exact match on the ambiguous half of the dataset and 0.21 on the non-ambiguous half. Neither number suggests a system that's close to reliably gathering and integrating evidence the way the task demands.

The researchers frame this as a call to action: QA systems need to get much better at actually collecting complete information, not just retrieving a plausible-sounding fragment of it. Given how far even a top-tier model fell short here, that gap looks less like a rounding error and more like a real hole in how these systems handle ambiguity and thoroughness together.

My take — AI-written commentary, not fact-checked reporting

A model that gets a fifth of complete answers right on non-ambiguous questions, and worse when names get murky, is not ready to be anyone's research assistant for anything that matters. The demos always use the easy single-hop trivia questions; this benchmark is a reminder that real-world questions are full of shared names and messy evidence trails, and that's exactly where these systems still fall apart.

Read more about this at: Apple Machine Learning Research

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.