Why AI Needs a “Genie Coefficient”
IEEE Spectrum Bruce Schneier
Researchers propose a new metric called the Genie coefficient to measure whether AI agents do what users actually mean, not just what they literally ask for. The metric evaluates the gap between user intent and AI actions across tasks like coding, legal work, and finance, requiring domain-specific benchmarks that test whether AI takes unreasonable shortcuts. Establishing this measurement would enable policies holding AI systems accountable for misinterpreting reasonable user requests rather than blaming users for unclear instructions.
Why it matters
Major benchmarks measure what AI can do. None measure whether it does what you mean: the distance between what you ask an AI to do and the unspoken assumptions about how you want the AI to do it. We propose a new metric: the Genie coefficient.There’s often a gap between one person’s request and another’s understanding. Most of the time, we bridge it using general knowledge. For example, if you ask a friend to get you coffee, they’ll pour a cup from the pot or buy one from a coffee shop. They won’t bring you a bag of raw beans or snatch a cup from a stranger and hand it to you. You never specified any of this. You never had to.One might think the fix is just to specify tasks, questions, and intent better. But in 1987, in their seminal book on AI, Terry Winograd and Fernando Flores succinctly captured why that won’t work: “Q: Is there any water in the refrigerator? A: Yes. Q: Where? I don’t see it. A: In the cells of the eggplant.” In human language, wants and desires are always underspe