Exa Adds Agent Ultra for Longer Web Research, Claims Lower Benchmark Costs

The new API mode divides searches among parallel workers. Its reported lead over competing agents depends on Exa’s evaluation setup.

By 3 min read
Exa Adds Agent Ultra for Longer Web Research, Claims Lower Benchmark Costs
Exa Adds Agent Ultra for Longer Web Research, Claims Lower Benchmark Costs

Listen to this story

The audio brief

About 1:23
0:001:23
Read transcript
Exa’s new Agent Ultra turns a research request into multiple searches running in parallel, aiming to build thorough, cited lists rather than deliver a quick answer. The company says the mode led four research benchmarks, but the results depend partly on how Exa ran the comparisons and have not been independently reproduced. On WANDR, which tests whether an agent can find qualifying entries and provide evidence for them, Exa reports a score of 81.4 percent. That compares with 72.3 percent for Opus 5.5 and 40.1 percent for Perplexity Agent. Exa also says a WANDR run cost about half as much as one with Opus 5.5. On DeepSearchQA, it reports a 46 percent cost reduction versus GPT-6 Astra. Those are benchmark-specific comparisons, not guaranteed prices for a customer’s task. Ultra is the highest-effort setting in Exa’s hosted research API. It splits a request into smaller jobs for subagents to search different domains at once. Developers can set a spending cap, with a default of 20 dollars, and can stop a run early while keeping what it has found. But complex research can take around 30 minutes, and the hardest tasks up to three hours. That makes Ultra suited to work teams can queue and review, not questions that need an immediate answer. The key test for buyers is still whether it finds the entries their own lists require—and backs them with useful evidence.

Story brief

3 key points

Exa has added Agent Ultra, the highest-effort option in its hosted research API, aimed at exhaustive, cited lists rather than quick answers. It splits requests into parallel subagent searches and can exclude entries customers already have. Exa reports leading results on four benchmarks, including 81.4% on WANDR, but comparisons mix published and in-house runs and have not been independently reproduced. Runs are...

  1. 01

    On WANDR, Exa reports 81.4%, versus 72.3% for Opus 5.5 and 40.1% for Perplexity Agent.

  2. 02

    Exa says WANDR runs cost about half as much as Opus 5.5; DeepSearchQA runs cost 46% less than GPT-6 Astra runs.

  3. 03

    Company Find-All cost comparisons measure spend per entity found, not per task.

A research list can look convincing while leaving out the company a team needed to find. Exa has launched Agent Ultra, a higher-effort mode of its hosted research API built for longer searches and evidence-backed lists. Exa says it led competing agents on four benchmarks. Its cost advantages vary by test; neither the scores nor the savings have been independently reproduced.

One question becomes many searches

Agent Ultra is the highest-effort setting of Exa’s Agent API. It splits a request into smaller jobs and assigns subagents to search different domains in parallel. Exa says it uses more capable models where needed and faster ones for other steps.

Exa proposes market maps, searches for papers and code repositories that meet technical criteria, and account lists with cited details. Users can pass in rows they already have so Ultra excludes them from new results. For those jobs, missing a qualifying entry can matter as much as including a wrong one.

Developers select the hosted service through the existing API by setting effort to “ultra.” It is not a model customers can self-host.

The lead Exa measured

Exa says Ultra led its comparisons on WANDR, DeepSearchQA, WideSearch and its own Company Find-All test. WANDR closely matches the product’s intended work: it checks whether an agent finds a large set of qualifying entries and supplies evidence across multiple fields.

Exa’s reported WANDR scores
81.4%Agent Ultra

Exa’s reported result for Agent Ultra.

72.3%Opus 5.5

The Opus 5.5 result in Exa’s comparison.

40.1%Perplexity Agent

The Perplexity Agent result in Exa’s comparison.

Ultra also scored 93.9% on DeepSearchQA and 58.9% on WideSearch, according to Exa. Its cost claims use different measures: a WANDR task cost about half as much as one run with Opus 5.5, while a DeepSearchQA task cost 46% less than one run with GPT-6 Astra. For Company Find-All, Exa claims the lowest cost per entity found, not per task.

Exa says its WANDR grader keeps the upstream evaluation logic but changes the contents tool, transport and model judging answers. It used competitors’ published results where available on that grader harness and ran other comparisons itself. Exa evaluated up to 200 tasks each for WANDR and DeepSearchQA, and 100 each for WideSearch and Company Find-All; graded counts varied by provider. The setup limits what the rankings can settle.

A spending cap is not a task price

Ultra uses metered Agent API billing. Each run has a default $20 spending cap, which developers can set between $1 and $100; a run that finishes early costs less. That cap is not a fixed task price, and benchmark savings do not promise lower costs for every research request.

Time is the other budget. Exa says complex runs typically take about 30 minutes, while very hard tasks can take up to three hours. Developers can set a duration limit or stop a run early and keep the results collected so far. Ultra fits research a team can queue and review later, not a question needing an immediate answer.

The remaining test is a buyer’s own list-building work. A strong benchmark score cannot show whether Ultra will find every eligible company, paper or record in a fresh search. Users still need to check both the evidence attached to returned entries and what the search may have missed.

Sources

  1. exa.aiIntroducing Exa Agent Ultra - A New Frontier for Deep Research
  2. marktechpost.comExa Launches Agent Ultra: A Subagent Swarm Deep Research API Built for Exhaustive List Building

Loading discussion...

YOUR READING SPACE

Notifications