A web search API 1000× cheaper than Exa
Every web search API will tell you their results are the best. So we built an open, reproducible deep-research benchmark — four providers, an LLM judge, 128 real research tasks (384 games per provider) — and let you judge the results yourself. The headline: SearchSpace is competitive on quality at roughly one-thousandth the price per query and 16–230× lower latency. Here is the chart that matters, and below it every number behind it.
The benchmark plainly shows that Exa and Brave beat us on quality, but our results are more than satisfactory for the majority of use cases, and the price / latency gap is massive. Regardless, we're actively working on improving our ranking and breadth so the quality gap inverts.
How it works
An LLM judge does a pairwise round-robin: for each brief it's shown two providers' reports side by side (with each cited source, so the judge can verify grounding) and picks the winner on coverage, grounding, depth, and clarity. Everything else about the environment, agent, LLM used, etc. is the same, with the sole exception of the search provider. Every pairwise outcome feeds a per-provider Elo rating, with the full win/loss matrix published.
Each task is written from scratch as an open-ended research question with a reproducible LLM prompt and hand selected verticals. These verticals are selected from industry reports (linked below) showing where revenue is actually flowing to different agentic/LLM use cases. The tasks are also intentionally not trivia tasks: they're meant to be real-world things a paying customer would want insights on.
The judge uses a larger model than the agents, and its cost is excluded from anything we'd attribute to a provider. The agents all run on the same smaller, more cost-efficient model, so the benchmark measures the search, not the synthesizer.
What it costs, and how fast
This is the whole point. At list price, SearchSpace is about $6.90 per million queries against Exa's $7,000 — a ~1000× difference — while landing below Exa and Brave, but beating Tavily handily. To bring Deep Research to as many people as possible, the economics have to actually work: if a dossier costs dollars, it stays a specialized tool; if it costs fractions of a cent, web search becomes a commodity the way LLM tokens have, and the set of viable use cases grows with it.
Cost per research brief
agent LLM + search fees · lower is betterSecondly, for an agent doing tens or hundreds of searches, the latency a search may take matters. If a Deep Research report can be served in single digit seconds instead of minutes, that is the equivalent of going from dial-up internet to gigabit fiber.
Search latency
p50 wall-clock per search call · lower is betterSearch latency is the provider's API p50 wall-clock time for search results to come back.
Vertical breakdowns
Some providers do better in specific fields, so a single "Elo" for the provider does too much dimensionality reduction. We report full vertical breakdowns of win rate per provider. We'd also encourage you to check out the full data to see each and every dossier made from each provider, along with the judge's analysis as to why it selected the way it did.
Why these categories
A deep-research API isn't used for bar trivia; it's used on the questions AI applications actually do in prod, where the money is. So we chose the eight categories at the highest-revenue, most token-hungry regions of the AI market: healthcare & biotech, science & IP, legal & regulatory, financial markets, crypto, software & developer tools, cybersecurity, and current events.
There are copious industry reports outlining where money is being spent and where tokens are being generated. Enterprise spending on generative AI roughly tripled to about $37 billion in 2025, and beyond the obvious software engineering, use cases sit in regulated, high-stakes verticals like healthcare and biotech, legal, and finance. Importantly, these verticals require fresh, up-to-date information that knowledge cut-offs don't allow: fresh, external information the model doesn't have memorized — and that's precisely where the purpose-built AI-search vendors (Exa, Tavily, Parallel) sell: deep-research agents, competitive and market monitoring, financial research, technical-docs lookup, and news.
These market figures are 2025 - 2026 estimates from VC and analyst surveys (Menlo Ventures' enterprise survey and Sacra's private market estimates); meant to be directional, not audited, and quickly-changing.
Look at your data
Benchmarks you can't judge yourself are a marketing claim. So this entire endeavor, every task and brief each provider's agent made, the sources it used, and the judge's winner and reasoning for every matchup, is here in a single database you easily browse. Read the actual reports next to each other, see the sources each agent accessed, and read why the judge decided the way it did.
Read every brief, side by side
The interactive explorer loads the full run entirely in your browser, no account required. Open any task to compare each provider's report and cited sources, and check out the judge's verdict on every game.
Open the data explorer