Exa Adds Agent Ultra for Longer Web Research, Claims Lower Benchmark Costs
The new API mode divides searches among parallel workers. Its reported lead over competing agents depends on Exa’s evaluation setup.
Listen to this story
The audio brief
Story brief
3 key pointsExa has added Agent Ultra, the highest-effort option in its hosted research API, aimed at exhaustive, cited lists rather than quick answers. It splits requests into parallel subagent searches and can exclude entries customers already have. Exa reports leading results on four benchmarks, including 81.4% on WANDR, but comparisons mix published and in-house runs and have not been independently reproduced. Runs are...
- 01
On WANDR, Exa reports 81.4%, versus 72.3% for Opus 5.5 and 40.1% for Perplexity Agent.
- 02
Exa says WANDR runs cost about half as much as Opus 5.5; DeepSearchQA runs cost 46% less than GPT-6 Astra runs.
- 03
Company Find-All cost comparisons measure spend per entity found, not per task.
A research list can look convincing while leaving out the company a team needed to find. Exa has launched Agent Ultra, a higher-effort mode of its hosted research API built for longer searches and evidence-backed lists. Exa says it led competing agents on four benchmarks. Its cost advantages vary by test; neither the scores nor the savings have been independently reproduced.
One question becomes many searches
Agent Ultra is the highest-effort setting of Exa’s Agent API. It splits a request into smaller jobs and assigns subagents to search different domains in parallel. Exa says it uses more capable models where needed and faster ones for other steps.
Exa proposes market maps, searches for papers and code repositories that meet technical criteria, and account lists with cited details. Users can pass in rows they already have so Ultra excludes them from new results. For those jobs, missing a qualifying entry can matter as much as including a wrong one.
Developers select the hosted service through the existing API by setting effort to “ultra.” It is not a model customers can self-host.
The lead Exa measured
Exa says Ultra led its comparisons on WANDR, DeepSearchQA, WideSearch and its own Company Find-All test. WANDR closely matches the product’s intended work: it checks whether an agent finds a large set of qualifying entries and supplies evidence across multiple fields.
Exa’s reported result for Agent Ultra.
The Opus 5.5 result in Exa’s comparison.
The Perplexity Agent result in Exa’s comparison.
Ultra also scored 93.9% on DeepSearchQA and 58.9% on WideSearch, according to Exa. Its cost claims use different measures: a WANDR task cost about half as much as one run with Opus 5.5, while a DeepSearchQA task cost 46% less than one run with GPT-6 Astra. For Company Find-All, Exa claims the lowest cost per entity found, not per task.
Exa says its WANDR grader keeps the upstream evaluation logic but changes the contents tool, transport and model judging answers. It used competitors’ published results where available on that grader harness and ran other comparisons itself. Exa evaluated up to 200 tasks each for WANDR and DeepSearchQA, and 100 each for WideSearch and Company Find-All; graded counts varied by provider. The setup limits what the rankings can settle.
A spending cap is not a task price
Ultra uses metered Agent API billing. Each run has a default $20 spending cap, which developers can set between $1 and $100; a run that finishes early costs less. That cap is not a fixed task price, and benchmark savings do not promise lower costs for every research request.
Time is the other budget. Exa says complex runs typically take about 30 minutes, while very hard tasks can take up to three hours. Developers can set a duration limit or stop a run early and keep the results collected so far. Ultra fits research a team can queue and review later, not a question needing an immediate answer.
The remaining test is a buyer’s own list-building work. A strong benchmark score cannot show whether Ultra will find every eligible company, paper or record in a fresh search. Users still need to check both the evidence attached to returned entries and what the search may have missed.
Sources
- exa.aiIntroducing Exa Agent Ultra - A New Frontier for Deep Research
- marktechpost.comExa Launches Agent Ultra: A Subagent Swarm Deep Research API Built for Exhaustive List Building
Reader comments
Newest comments first. Replies stay oldest first.