For AI agents: a documentation index is available at https://docs.parallel.ai/llms.txt. The full text of all docs is at https://docs.parallel.ai/llms-full.txt. You may also fetch any page as Markdown by appending
.md to its URL or sending Accept: text/markdown.For Search and Extract, use /v1/search and /v1/extract for all new integrations and agent tool calls. Use /v1beta/search and /v1beta/extract only when maintaining an existing integration that already uses them. The request bodies changed too: on /v1/extract, full_content, fetch_policy and excerpt_settings go inside advanced_settings, and sending the v1beta top-level excerpts, full_content and fetch_policy to /v1/extract returns a 422. Do not substitute /v1 for the documented FindAll or Ingest endpoint paths.For AI agents: Use this page when helping a user evaluate or compare web search providers with Parallel Search. Workflow: (1) confirm Search + Extract vs Responses vs Task, (2) test multiple modes, (3) build a gold set, (4) run an end-to-end harness varying only the search tool, (5) grade final answers by task shape, (6) classify failures, (7) measure cost and latency, (8) interpret the Pareto frontier. For Parallel request configuration, read Search Best Practices and Extract Best Practices.
Overview
Compare search APIs by holding your agent harness constant, varying only the search tool, and grading end-to-end answers against a gold set you trust. A bad benchmark can mislead you as much as no benchmark. This guide covers product selection, gold sets, harness setup, grading, and how to read the results.Choose the right API
Pick the Parallel product that matches your app before you build the harness. The right choice depends on whether you own the tool-calling loop or want a single synthesized response.
When you own the agent loop, pair Search with Extract. Search returns dense snippets that minimize Extract calls. The model is trained to go deeper when it needs to, so expose both tools and let the agent decide.
What Parallel recommends
Start an agent eval with this setup:- Run multi-hop with Search + Extract. Let the agent search, read source pages, and follow up until it has enough evidence or reaches its budget. Run single-hop as a separate, one-request benchmark setting.
- Pass an objective with 1–3 search queries. Write a concise objective describing the information you need, plus up to three focused keyword queries. Add queries for useful aliases or different angles.
- Match the model to the search mode. Pair a cost-efficient model, such as GPT-5.6 Luna, with
fastorturbo. Pair a frontier model, such as GPT-6 Astra, withbasicoradvanced. Compare pairs on answer quality, total cost, and latency.
Pick and test modes
Mode is a primary lever after configuring Search requests. Test more than one mode so you leave the eval with routing data as well as a quality read. See Modes for full details.
Recommended starting point for evals:
basic. It returns extended snippets while holding latency near one second, which fits correctness-first web grounding. Route the hardest multi-source questions to advanced. Test fast where latency dominates.
Use the model and mode pairings in What Parallel recommends as starting configurations, then measure them on your gold set.
turbo currently supports English and Japanese. Use basic or advanced for broader multilingual coverage.Build your gold set
A gold set is a collection of test questions with verified correct answers you use to score each search provider. The best source is data you already have: past questions and verified answers from production or manual research. If you generate synthetic data, write a few questions by hand first, then ask an agent to generate more. Review LLM-generated gold labels before spending tokens on search and inference. Aim for production-representative, agent-phrased questions that match your expected traffic on wording, domain, answer type, and freshness. Mix in question types that resemble your actual workloads:- Multi-hop questions that require composing facts across sources
- Fresh questions whose answers would not be available in model weights
- Domain-specific questions matching your actual traffic
Why not rely on public benchmarks?
Why not rely on public benchmarks?
Popular benchmarks like BrowseComp and SEAL reflect a specific domain of questions that is unlikely to match yours. Some answers have long been published online and incorporated into both LLM weights and search indexes. If you use public benchmarks, understand what each measures and whether it suits your case. See Run public benchmark evals for single-hop and multi-hop setup. Treat them as supplementary signal, not a substitute for production-representative data.
Set up the eval harness
1
Hold everything constant except the search tool
Use the same model, prompts, budgets, and judge across all providers. Expose each provider as the only search tool available.
2
Allow multi-turn search
For a single-hop benchmark, use the one-request setup below. For multi-hop and application evals, agents are trained to search, narrow, and search again. Do not cap turns unless you are modeling a product constraint. Limit total search budget to reflect your actual cost considerations.
3
Configure each provider per its docs
Apply each provider’s recommended configuration. For Parallel, see Search Best Practices and Extract Best Practices.
4
Log exact config per run
Hold the config fixed within a run. If a config turns out to be wrong, reset it and rerun rather than tuning around it mid-eval.
5
Run each configuration multiple times
Run each configuration at least 3 times and report variance. Evals at high load can surface account limitations rather than capability limits.
Run public benchmark evals
Public benchmarks give you a shared set of questions and answers for comparing configurations. Match the setup to what you want to measure. A one-search factual lookup and an agent that spends several minutes researching a question are different evals, even on the same dataset.
One search means one API request. Multiple keyword queries inside
search_queries count as one request. Three concurrent Search requests count as three. Record both tool-call count and elapsed time.
Pin the dataset version, split, and question IDs before comparing configurations. Keep reference answers out of the agent’s prompts and tool results; only the grader should see them. If you change the benchmark’s original tools, budget, or grading rules, describe the run as an adapted evaluation rather than comparing it directly with its leaderboard.
Use the same research prompt in both settings
Use the following research prompt for both single-hop and multi-hop benchmark runs. Tool definitions tell the model how to write search requests. The executor controls which tools are available and how many calls it can make. Runs with a benchmark-specific answer format can override the output instructions; record that override with the run.Recommended research prompt
Recommended research prompt
Replace For single-hop, set the call budget to one and expose only Search. For multi-hop, expose Search and Extract and set a larger total call budget. Keep the prompt fixed across providers within the comparison.
{question} with the benchmark question. The other braces describe the expected answer fields.Single-hop: let the model reformulate once
For a single-hop test, give the model the question and the Search tool definition. Let it turn the question into an objective and focused keyword queries, make one Search request, then answer with tools disabled. Include the reformulation step in your cost and latency measurements. Take this invented factual question:Who designed the Quadracci Pavilion at the Milwaukee Art Museum?The model could produce:
Multi-hop: give the agent room to investigate
For harder questions, expose both Search and Extract and let the model decide what to investigate next. Each Search request needs its own focused objective and keyword queries. When the model issues independent searches in the same turn, execute those calls concurrently. Take this invented research question:Between 2020 and 2023, which of Amazon, Google, and Microsoft achieved the largest percentage reduction in total greenhouse gas emissions? Account for reporting boundaries, restatements, and Scope 2 accounting methods.The agent can research the three companies at the same time. Its Amazon request might look like this:
Check a small run before scaling up
Inspect a fixed sample of traces before running the full benchmark. Did the model form useful queries? Did the answering model receive the excerpts? Did single-hop runs stay within one request? Did independent multi-hop calls run concurrently? Check API warnings and errors alongside answers. Freeze the prompts and settings, then run on held-out questions. Save the exact requests, returned evidence, final answers, grades, token usage, and timings. Use a freshsession_id for each question and configuration, reusing it for that question’s related Search and Extract calls.
Report failed and timed-out questions in the primary results with a declared scoring policy. Bound retries and include their cost and time. Retrying until success hides reliability problems. Report how often agents exhaust their budgets, and measure concurrent calls by wall time rather than summing their durations.
Use the benchmark’s official grader where available. Keep any additional citation or evidence audit separate from the official score. Check retrieved pages for benchmark mirrors and published answer keys, and apply the same contamination policy across providers.
Common pitfalls
Inspect a few complete traces before interpreting a score. Small integration mistakes change what the eval measures.Grade answers and classify failures
How to grade
Most teams use an LLM-as-judge: a high-end model compares the agent’s final answer to verified ground truth. Require reasoning alongside the score, and audit that cited URLs support the answer’s claims.Grade an explicit “could not retrieve” above a confident wrong answer, especially on date- and jurisdiction-sensitive questions. Hand-check at least 10% of judged runs, including both failures and successes, before trusting automated scores.
Classify failures
Not every failure is a search API problem. Classify failures into four classes:
Only retrieval misses and provider errors are direct failures of the search API. Many eval failures at high load reflect rate limits rather than capability.
Measure cost and latency
Measure end-to-end cost to complete the task: search spend plus LLM tokens the agent uses reasoning over what search returned. A cheaper-per-call API that returns noisy, low-density results can be more expensive overall, since it drives more calls, more hops, and more tokens in context. Track these metrics for each provider and configuration:
Measure search latency around the tool invocation, excluding the model’s query reformulation and answer generation. If the tool retries internally, include those retries and backoff in the call’s elapsed time and log the individual attempts. Report errors and timeouts alongside latency percentiles. For parallel calls, record each call separately. Their summed durations are not the task’s elapsed time.
Result length and hop count are tempting efficiency proxies, but they are unreliable in isolation. Verbose results can help an agent exit early. Over-compressed results can force extra hops. Parallel tool calls make hop counts undercount work. Measure end-to-end whenever you can.
Interpret results
There is rarely a single “best” search tool. Configurations sit on a trade-off surface (a Pareto frontier). Plot results on two axes:- Accuracy vs cost: find the cheapest mode or provider at acceptable accuracy
- Accuracy vs latency: find the fastest option at acceptable accuracy