Even the most expensive search agents leave out about a third of the correct answers on hard research tasks, according to a new benchmark from Exa. Exa sells a search API for AI agents, so the company that built ATLAS, wrote its answer key, and ran the tests also profits from the verdict that search needs to get better.
That does not make the numbers false. It does mean the comparison deserves a harder look than a neutral lab’s would. In the test where the agent’s harness is held fixed and only the search backend changes, Exa says its own search sits on the best cost-to-quality frontier. The rivals in that lineup were Perplexity, Brave, and Parallel, and Exa reports a 16 percent spread in scores across the four. The post does not say whether that is points or a relative gap.
The benchmark’s central idea is the golden answer. Each of the 547 tasks asks an agent to list everything that fits a strict set of criteria, then fill in 2 to 10 facts about each item. The items are companies, people, or places. The golden answer is the complete table that a perfect response would contain, assembled in advance by the people who built the test. Scoring then asks how much of that table came back, and how much of what came back is right.
This is why the finding is hard to spot in practice. An agent that returns seven correct rows out of ten reads as confident and plausible. Nothing on the screen says three are missing. ATLAS’s headline metric, row F1, is strict about it: a row counts only if every cell in it is correct. The best system Exa tested reached 0.66. It ran for 18 minutes on each task, at a cost of $8.92. No run costing under $1 beat 0.5. The post’s prose does not name the 0.66 system; per-system scores sit in charts. The twelve systems included Exa’s agent at four effort levels, Claude Opus 5.5 and GPT-6 Astra on their native search, and GPT-6 Luna paired with each of the four backends.
Exa makes three design claims, and each has a mechanism behind it. The first is that tasks are not memorized. Exa asked three closed-book models, with no search access, to recall full answer sets, and they managed it for 9 percent of ATLAS tasks, against 48 percent for BrowseComp, 59 percent for WideSearch, and 61 percent for DeepSearchQA. Topics come from Exa’s anonymized search demand across more than 300 topics, and a cheap screening agent discards tasks that look memorized. As a sanity check, Exa hid the top seven of every ten results: ATLAS scores roughly halved, while the older benchmarks kept 81 to 89 percent.
The second claim is wide, multi-domain search. Of the tasks, 74 percent ask for 10 or more entities. From its construction logs, Exa estimates that finding the entities takes a median of 4 domains per task and the full table takes 18. A strong agent on WideSearch cited a single website for 48 percent of tasks, against 0.2 percent on ATLAS.
The third claim, cost-effective grading, partly means a model does the grading. The grader matches rows by entity name and aliases first, then falls back to GPT-6 Luna when that fails. Luna also judges the 4.5 percent of cells that rules cannot settle. Scoring one system across all 547 tasks costs about $0.24, and regrading 7,658 answers produced the same row F1 on about 7,550 of them, or 98.6 percent. Repeatable is not the same as correct, though. The answer key itself was built by Codex on GPT-6 Astra and Claude Code on Opus 5.5, checked by model verifiers, and audited by model judges with Gemini 3.1 Pro breaking ties. Exa engineers settled what remained. A sampled audit put wrong gold values at 0.9 percent, with a 95 percent confidence interval of 0.6 to 1.4. Another 7.1 percent of cells were left ungraded as unsettleable.
Exa did take steps against home-field bias. Its construction agents searched through a tool that queried Exa, Brave, and Perplexity and hid the provider names, and cited pages came from each at near-equal rates. The ATLAS tasks, answer tables, and grader are not public yet. Exa promises them in the coming weeks, so nobody outside the company can rerun any of this today.
If your team shipped a research or deep-search feature, you have probably tested whether its answers are right. Few teams test how many answers are missing. Hand-build a complete list for ten real queries your users run and count how many rows your agent returns against it.
Based on “ATLAS: Evaluating Agents on Search-Intensive Tasks,” published by Exa on 8 October 2026 and written by Alexander Goldberg, Joshua Ahn, and Scott Langille.