AI Pulse by Inblix

Consensus built a multi-agent GPT‑5 researcher that does weeks of work in minutes

OpenAI Blog · Jul 12, 2026 · 2 min read · Read original article →

Curated by the Inblix editorial team


Featured image for article: Consensus built a multi-agent GPT‑5 researcher that does weeks of work in minutes

Eight million people already use Consensus to cut through the noise of 220 million academic papers. But the company’s newest feature, Scholar Agent, isn’t just a better search bar. It’s a multi-agent system built on GPT‑5 and the Responses API that mimics how a human researcher actually thinks: planning a search strategy, reading papers in batches, and synthesizing findings into structured answers. The goal is to compress weeks of literature review into minutes.

Christian Salem, who co-founded the company with Eric Olson, describes the shift in blunt terms. “Research isn’t just finding papers,” he says. “It’s interpreting results, comparing findings, and connecting ideas.” The new architecture breaks that labor across four specialized agents. One plans the query, another scours the paper index and citation graph, a third interprets individual studies, and a fourth synthesizes everything into a final output. The narrow scope of each agent is a deliberate hedge against the hallucinations that plague large language models. If no credible studies meet the quality threshold, the system simply says so rather than inventing a plausible-sounding answer.

Switching from Chat Completions to the Responses API was a practical decision driven by reliability and cost. The team says GPT‑5 outperformed GPT‑4.1, Sonnet 4, and Gemini 2.5 Pro on both tool-calling accuracy and planning stability. That performance headroom let Salem’s engineers spend less time on prompt engineering and more time on what they call “context engineering”: assembling a structured bundle of papers and metadata before any text is generated. Every answer comes with a “research context pack” that traces claims back to their source.

From the start, Consensus bypassed slow institutional sales to target researchers directly. That consumer-style approach shaped the product’s rapid adoption, spreading from graduate students to clinicians who use it to surface the latest evidence. “Everyone said you can’t go direct-to-consumer in academia, but AI has changed that,” Salem says. “People don’t wait for approval anymore—they use what works.” The Mayo Clinic’s medical library recently signed on, suggesting the ground-up strategy is starting to pull in the very institutions it initially ignored.

💡 Key Takeaways

  1. GPT‑5 outperformed GPT‑4.1, Sonnet 4, and Gemini 2.5 Pro on the tool-calling and planning tasks Consensus relies on, which let the team focus on agent behavior over prompt engineering.
  2. The four-agent architecture limits hallucinations by design—each agent has a narrow scope, and the system refuses to answer when it cannot find high-quality studies.
  3. By targeting individual researchers and clinicians directly instead of selling through institutions, Consensus built a user base of 8 million and later attracted enterprise clients like the Mayo Clinic.

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

Learn more

Glossary terms

← Back to all articles