
/Product7 min read
Your Agent Has the Company Name. How Much Can it Actually Tell You About It?
Today we’re introducing Tavily’s Company Search Benchmark. It reflects real agent tasks and tests how much of a company’s story search engines can uncover for your agent. Tavily ranks first among the 9 evaluated providers.
Business AI agents need accurate, up-to-date company information to make informed decisions. Company research supports use cases such as CRM enrichment, prospect qualification, onboarding, underwriting, and KYB/KYC by providing key information about what a company does, who it serves, and who owns it.
That context changes as businesses grow, enter new markets, change ownership, and expand their operations. Fresh information matters: a new product line can change a prospect’s fit, while an acquisition can change an ownership assessment. The latest information an agent needs is spread across company websites, filings, news, and other sources, and depends on the task at hand. Existing company profiles can leave gaps, especially for private businesses with a limited online presence.
Agentic search helps agents build on those profiles, investigate changes, and answer specific business questions as a workflow runs. The agent searches, evaluates the response, and determines whether it has enough information or needs to keep looking.
Today, we introduce our first Company Search Benchmark, built around real production query patterns. It evaluates a foundational step in company research: how much a search engine helps an agent to complete the task at hand.
The core concepts behind the Company Search Benchmark
- Wording matters. Our queries reflect how customers’ agents phrase their searches in production, because how a question is asked can affect what a search engine finds.
- Today’s questions need today’s data. The dataset draws on recent production patterns to capture what agents are researching now.
- Company facts change. We build a fresh scoring card for each query on the day of evaluation, recording verified facts and defining what a useful answer should cover based on the evidence available at that time.
- A sensitive measure of search quality. We evaluate how much a single search response helps answer the research question, rather than whether an agent can eventually find the answer through repeated searches.
- An agent’s perspective matters. We evaluate what an agent can learn from a response and use to complete its research task, not just what information the response contains.
Single-query evaluation and longer agentic tasks
We evaluate search tools across single-query benchmarks, single-hop agentic tasks, and longer agentic workflows. Longer workflows help us understand how agents and search tools work together across multiple steps and how search affects the effort required to complete a task. However, a multi-step agent can compensate for weaknesses in individual search responses, making differences between providers harder to detect. Many company research tasks can and should be completed with very few search steps – ideally one. In this article, we focus on single-query evaluation: how much useful information a provider delivers in one response.
Building the evaluation dataset
Queries that reflect real company research
Our queries follow patterns observed in production traffic because an evaluation is most useful when its queries reflect the tasks agents actually perform. When using a public benchmark to choose a search tool, check how closely its queries match your own workload. The entities being researched, the intent behind each query, and the wording all matter.
Many of our customers have zero data retention agreements with us or do not permit us to use their queries for evaluations or system improvements. We do not collect or use their queries for these purposes. A huge thank you to those who choose to collaborate with us on the quality and contributed query data to our evaluations. Today, we’re proud to share the results of that work! We apply strict privacy safeguards to these queries: including removing customer identifiers and replacing company names, addresses and other details that could identify the research subject with information about other randomly sampled from publicly accessible company registries.
As we’ve noted, the subject, intent and wording of a query all matter. To keep the evaluation representative, we preserve the original distributions of relevant characteristics, such as jurisdiction, company size and public or private status while keeping the data anonymized. This allows us to evaluate realistic company research patterns while protecting customer privacy.
Starting with the query
We start with a query, rather than a predefined ground-truth answer sourced from public registries, news RSS feeds or official company websites. We preserve the complexity our customers’ agents face in production. Some queries are very difficult, and some may not even have a complete answer available on the public web. That is not a reason to exclude them. Removing such queries would risk filtering out the hardest production tasks and measuring quality only where the answer is easy to establish.
Company research may also involves ambiguous names, incomplete information and sometimes incorrect premises. Our research process considers valid interpretations of the query, resolves the relevant entity where possible and checks whether the question assumes something that the evidence does not support.
A short ground-truth sentence or a single number is often insufficient. Instead, we build structured scoring cards that describe what the best answer supported by available web evidence should contain: which aspects it should cover, what counts as useful partial information, and what makes a response irrelevant or misleading.
Keeping the dataset current
Company information changes, and so does its availability. A fact that is difficult to find today may be covered by many sources tomorrow, while an answer that is correct today may soon become outdated. Internally, we regularly refresh the dataset, including daily updates to the most time-sensitive slice.
Building scoring cards and evaluating answers
To measure how much a search response helps an agent, we first establish what a useful, accurate answer looks like at the time of evaluation.
To measure how much a search response helps an agent, we first establish what a useful, accurate answer looks like at the time of evaluation.
For each query, a Deep Research Agent builds a shared scoring card, drawing on all search providers in the comparison and extraction tools. It can reformulate queries, investigate supporting details, follow leads, and cross-check findings. We use Toloka for human validation of financial figures.
The resulting scoring card states the query’s intent and target company, checks whether its assumptions are supported by evidence, and records the verified facts that answer it. It also defines the criteria for Complete, Partial, Related and Irrelevant answers.

We then send the same query to each provider and use the same LLM as the Answer Generator for every response. This simulates a single agentic hop: answering the original question from the returned search results. The generator sees only the query and the search response, without access to the reference answer, the scoring card or the provider’s identity. For time-sensitive queries, it also receives publication dates extracted from the full text of the retrieved documents by an independent tool and is instructed to take them into account when generating its answer. To keep the evaluation sensitive to search errors, we instruct it to rely only on the retrieved evidence and preserve factual details. For your own evaluations, we recommend using the model and prompt you plan to deploy in production.
A Judge Model then evaluates each generated answer against the query and the shared scoring card, including the query’s intent, freshness requirements and criteria for a correct answer. Using our definitions of the four relevance levels, the Judge decides which level the answer falls into, and the answer receives that level's weight as its score.
The resulting score measures how useful a search response is for answering the question. We do not directly score relative ranking within the top five or ten results or the number of duplicates; we measure whether the returned information enables the solver to produce a useful, correct answer using only the search response.
We distinguish four levels of answer quality:
Answer Quality | Meaning | Weight |
|---|---|---|
Complete | Fully answers the query; no further searches are needed. | 1.0 |
Partial | Provides accurate information that materially helps answer the query, but does not meet all requirements for a complete answer. Clearly identified proxy figures or related facts may qualify when they help address the query. | 0.85 |
Related | Relates to the query but contains little useful information. | 0.15 |
Irrelevant | Provides no useful information for answering the query. | 0.0 |
The reported score is the average of these weights across queries.

Company Search Benchmark Results
Company Search Benchmark
Answer Quality Score × 100 · Higher is better
Final thoughts
We’re sharing the Company Search Benchmark dataset based on queries from recent weeks, so you can explore the query patterns we see in production and the scoring criteria, assess how well the queries match your own use cases, and use the approach to inform your evaluations. Our internal evaluations continually evolve alongside changing company information, customer needs, and the ways agents formulate search queries.
The most useful test of any search tool is how well it supports the questions your agents actually ask. Explore the dataset, then evaluate on your own queries, with your production model. If you’d like help getting started, talk to us about evaluating search for your workflow.
