Percept-Lens: A Deep Dive into AI-Generated Image Detection
Sakana AI
Sakana AI says a frozen vision model plus a simple rule can spot AI-made images. It beat the strongest released detector they tested, even without task-specific training.
Based on reporting by Sakana AI — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Sakana AI is pushing a deceptively simple idea: maybe the detector is not the hard part. Maybe the vision model already knows enough, and the real problem is how you read out that information. The company says its upcoming ECCV 2026 work shows a general-purpose model can separate real from AI-generated images with a basic decision rule on frozen features.
That matters because image detectors have a habit of looking good until the inputs move. A detector trained on one generator, one prompt style, or one image domain can fall apart when any of those change. To measure that problem, Sakana AI introduced Percept-Lens, a shared evaluation setup for these shifts, and used it to show that released detectors can degrade sharply outside familiar data.
The new question was more interesting than “can we build a better detector?” It was: when a detector fails, did the underlying vision model lose the signal, or did the classifier on top of it simply miss what was already there? Sakana AI’s answer was to freeze a general-purpose vision model and fit a simple Gaussian decision rule on top of its representations.
That rule looks at how labeled real and AI-generated images sit in feature space, then assigns a new image to the group it most closely resembles. On the same broad evaluation suite, Sakana AI says this setup beat the strongest released AI-generated image detector it tested, despite the base vision model never being trained specifically for detection.
The takeaway is not that bigger models magically solve the problem. It is that some of the useful structure is already sitting in a general-purpose vision model, waiting for a cleaner way to be used. Better vision models may help, but this result says the readout matters too.
My take — AI-written commentary, not fact-checked reporting
This is the sort of result that should annoy both camps in the usual AI argument. The model people get to say representation matters; the detector people get reminded that a fancy head can still be dumb. And yes, the industry’s favorite move remains building a complicated patch over a problem a simpler layer might already half-solve.
Read more about this at: Sakana AI