Consensus accelerates research with GPT-5 and Responses API
OpenAI
Consensus, a research search tool, rebuilt its backend on GPT-5 and OpenAI's Responses API to run multi-agent literature reviews in minutes. It now serves over 8 million researchers who'd otherwise spend days digging through papers by hand.
Based on reporting by OpenAI — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Academic research has always had a bottleneck problem: the papers pile up faster than any human can read them. Consensus, the search engine built specifically for scientific literature, is trying to close that gap by handing the grunt work to a fleet of GPT-5 agents.
The company switched its infrastructure to OpenAI's Responses API, which lets multiple specialized agents pass tasks back and forth instead of relying on one model to do everything. In practice, that means one agent might hunt down relevant studies, another extracts the key findings, and a third stitches the results into something resembling a literature review. What used to take a grad student an afternoon of skimming abstracts now happens in a few minutes.
Consensus says it's now working with more than 8 million researchers, a number that includes academics, clinicians, and policy analysts who need to synthesize evidence quickly rather than wade through PDFs one at a time. The pitch isn't that AI replaces the researcher's judgment call on what matters — it's that the tedious part, finding and summarizing dozens of papers, gets compressed into something you can review over coffee.
This is also a quiet signal about where OpenAI wants its enterprise tools to live. The Responses API is built for exactly this kind of multi-step, multi-agent workflow, and Consensus is essentially a showcase for what happens when you give GPT-5 a narrow, high-stakes job — reading science — instead of asking it to be a generalist chatbot. Specialized use cases like this tend to age better than flashy demos, because researchers will tolerate a tool that's a little rough if it saves them real hours.
My take — AI-written commentary, not fact-checked reporting
I'll say it plainly: this is the kind of AI application that actually deserves the hype, because it's narrow, measurable, and solves a problem researchers have complained about for decades. The bigger story here isn't Consensus, it's that multi-agent orchestration via APIs like this is quietly becoming the default architecture for serious enterprise AI, while the chatbot wars get all the headlines.
Read more about this at: OpenAI