Deep research agents verify their answers by running a loop. The agent breaks the question into sub-questions, searches at least three independent sources, reads each result in full, pulls out every factual claim with its source URL, and scores that claim on a four-tier confidence scale before the claim is allowed into the answer. When two sources disagree, both sides go into the answer with their links attached, so the reader can weigh the evidence. That loop is the difference between an answer that sounds right and one that has been checked.

OpenAI launched Deep Research mode in February 2025, Google followed with Gemini Deep Research, and Perplexity added Pro Search. Those products run the loop well for general questions, but you cannot control the loop structure, confidence scoring or source selection, so an agent you build yourself is the one that shows the verification steps in full. This matters because single-shot answers are not reliable enough for professional use. A 2025 Stanford University study of AI legal research tools found they hallucinate between 17 and 33 per cent of the time even when purpose-built for one domain.

The demand for that check is measurable. LangChain's State of Agent Engineering, published on 12 June 2026 from a survey of more than 1,300 professionals, found 57 per cent of respondents had agents running in production, while 32 per cent named quality as a top barrier to shipping them. Observability had reached 89 per cent and evaluation 52 per cent, so most teams can watch an agent run without being able to check whether its answers were right. That gap is the one a verification loop closes.

We run a version of this loop ourselves. Every post on this blog is meant to start from a research pack gathered by our own research skill, which searches through SearXNG, reads each page with Firecrawl and records every source next to the claim it supports. The blog writing agent we built in Claude Cowork works from the same kind of pack before a draft exists. On 29 September 2026 we opened all 81 packs on disk to see how well the loop holds up in practice, and the counts are further down. Here is how the loop works, how to set it up in the tools your team already uses, and what separates an answer you can trust from one that only sounds right.

How a deep research agent verifies answers: the five-step loop

Verification is a gate, and it sits in front of every claim. A claim only reaches the answer when independent sources agree on it, and the strength of that agreement sets its confidence rating. One source is a lead. Two agreeing sources make a claim usable. Three or more make it solid. When the sources disagree, the disagreement goes into the answer with both links attached, because a reader who can see the conflict can usually work out which side applies to their case.

What makes it work is the loop reading whole pages, keeping the source URL next to every claim, and refusing to stop until the confidence rating clears the threshold you set. The five steps below take that from principle to a running loop, and the build section turns the steps into code.

Step 1: Break the question into sub-questions

The first thing a deep research agent does is decompose your question. If you ask 'What is the market size for AI agents in healthcare?', the agent splits it into sub-questions before it searches anything: total AI market size, healthcare AI adoption rates, percentage of AI spend going to agents, recent funding rounds. Each sub-question gets its own search loop, and the agent synthesises the answers at the end. The more granular your decomposition, the more likely you catch conflicting data.

In practice: In Hermes, sub-question decomposition lives in a skill file. Create ~/.hermes/skills/research/deep-research/SKILL.md with a rule that says: "Before searching, split the user's question into at least 3 sub-questions. Search each one independently, then synthesise." The agent loads this skill at session start and applies the decomposition rule to every research request. In Claude Code, the same logic goes into your CLAUDE.md under a research conventions section. Keep the decomposition rule simple, about five lines, so the agent can follow it without overthinking.

Step 2: Search with multiple sources

A single search engine introduces blind spots. The agent should query at least three sources: a web search for broad results, a news search for recent developments and a specialised source like arXiv for academic research or Crunchbase for funding data. If two sources disagree, the agent flags the discrepancy and reports both positions.

Perplexity's Pro Search is a good example of multi-source searching done well, because it queries several indices at once and reports the conflict when sources disagree. In a custom agent, you replicate this by wiring multiple MCP tool servers, each connected to a different search API. Hermes supports this natively: define separate web-search, news-search and academic-search tools in your MCP config, and the agent calls them in parallel. The 10-layer AI agent stack covers how tools (layer 6) and skills (layer 5) wire together for exactly this kind of multi-source workflow.

Why it matters: Single-source answers are a leading cause of AI hallucination in enterprise settings. The Stanford study is the clearest evidence, and even purpose-built research tools hallucinate between 17 and 33 per cent of the time. When a claim matters, the agent needs independent confirmation from a second and third source before it counts.

Step 3: Read and extract claims

For each result, the agent reads the full page. It extracts specific claims, such as numbers, dates, names and statistics, and records the source URL, publication date and a confidence flag for each claim. This is what separates a deep research agent from a regular chat answer. Every claim has a backlink to its origin. You can click through and verify the agent's work.

In practice: When configuring web_extract or a similar content tool in your agent, set the character limit high enough to capture a full article. A limit of 15,000 characters covers most in-depth pieces. Tell the agent to extract claims in a structured format: claim text, source URL, publication date and a confidence flag. We use JSON lines for this, one claim per line, with the source URL as the key:

{"claim": "AI agent market projected to reach $47B by 2030", "source": "https://...", "date": "2026-03-15", "confidence": "single"}
{"claim": "AI agent market projected $42-51B by 2030", "source": "https://...", "date": "2026-04-02", "confidence": "single"}

When the agent finds two claims about the same topic, it compares them. If they agree within a reasonable margin, confidence goes up. If they conflict, the agent notes the disagreement and searches for a third source to break the tie.

Step 4: Evaluate confidence

The agent scores each claim on a four-tier scale. If the answer relies on a single source, it is low confidence. If two independent sources agree, medium. Three or more independent sources with consistent data is high confidence. And if sources disagree, the agent presents both sides with source URLs and leaves the call to the reader.

This is the scale we recommend for an agent you build yourself:

ConfidenceWhat it takes
HighThree or more independent, credible sources agree on the same claim
MediumTwo independent sources agree, or one high-quality source with supporting context
LowOne source only, or sources that partially agree with caveats
ConflictingSources disagree, so the agent presents both sides with their source URLs

This scale lives in the agent's skill file as a scoring rubric. When the agent returns a research result, it includes a confidence badge for each claim, and the reader can click through to the sources and check them. Putting the rating next to the claim shows which parts of the answer are solid and which are tentative, so the uncertainty sits on the page where someone can act on it.

Step 5: Iterate until you have enough

After each pass the agent checks the combined confidence across all sub-questions. If it misses the configured threshold, the agent writes more specific search queries and runs the loop again. A build with one search, one read and one answer is a search tool with a summary on top, and this step is what turns it into a research agent.

We could not find a credible public figure for how many iterations a typical question needs, so treat the ceiling as a budget. Our own research skill stops working a sub-question once two independent sources confirm it, or once new pages stop adding claims, and that second rule ends more loops than any fixed number. Still set a hard ceiling, with 10 iterations per question as a sensible start, so the agent cannot loop forever on a question the web cannot answer. The ceiling and the confidence threshold go in the same skill file as the decomposition rules.

How to build a deep research agent: the loop in code

A deep research agent is a loop with five stages and two numbers that decide when it stops. The stages run plan, search, read, score and decide. The numbers are your confidence threshold and your iteration ceiling. Set those two and the search API becomes a detail you can swap without touching anything else.

Here is that loop as Python you can drop into a script or wrap as a Claude Code tool:

threshold = "high"     # "medium" for business questions
max_iterations = 10    # hard stop on unknown answers
iteration = 0
claims = []

while iteration < max_iterations:
    subs = plan(question)        # 3+ sub-questions
    pages = search(subs, k=3)    # web, news, database
    claims += extract(read_full(pages))
    scored = score(claims)       # high/medium/conflict
    if enough(scored, threshold):
        break
    iteration += 1

report = write_up(scored)

Two settings carry most of the weight. Set the confidence threshold too low and the agent writes a confident report from a single vendor blog, which is the exact failure the verification loop exists to prevent. Set the iteration ceiling too high and a question the web cannot answer burns through your search credits before the agent gives up. A reasonable starting point is high for anything a client will act on, such as pricing or legal questions, and medium for background reading.

In practice: Log every iteration to a file so the loop can be audited later, with one JSON line per iteration holding the iteration number, the sub-queries issued, the number of sources read and the confidence distribution at that point. When a report later turns out to be wrong, that log shows whether the agent found contradicting evidence and ignored it, or never found it at all.

What 81 of our own research packs show

Our pipeline keeps the output of every research run as a source ledger, one file per post. Each entry holds the URL, the claim that source supports, a credibility tag of primary or secondary, and the part of the post it is meant for. A separate list of gaps records what the research could not confirm. On 29 September 2026 we counted what is in them.

What we countedResult
Research packs on disk81
Sources recorded across all packs2,533
Tagged primary (the vendor, the study or the official docs)1,269
Tagged secondary (someone reporting on the original)1,227
Median sources per pack32
Packs below the skill's own 30-source floor19
Packs with a written gap list71, holding 305 gap notes

Nineteen packs missed the floor. A ledger makes a shortfall like that visible, because the count sits in a file anyone can open, where a chat transcript would have hidden it. The skill also caps any single domain at 25 per cent of a pack, since twenty sources from the same kind of publication give you one perspective.

The gap notes are where most of the verification happens. The ones that recur are pages that block automated readers, reports locked behind a form or paywall, figures we could only find in someone else's summary of the original, and searches that turned up no credible benchmark at all. Each note ends in an instruction to the writer. In the pack behind our win-loss analysis guide for lean teams, a 47 per cent close-rate figure credited to RAIN Group appeared only in a secondary summary, and the note told the writer to leave it out unless the primary page turned up. It never did, and the figure is not in the post. That is the confidence scale doing its job, with a lone secondary source rated low and kept out.

Our skill's written rule is shorter than the four tiers in the table above. Every claim the post leans on needs two independent sources, and a conflict gets flagged with both sides kept. The primary or secondary tag does the rest, because a secondary source standing alone is exactly the single-source lead the scale marks low. If you want to see how this research layer fits with memory, skills and tools in a full agent, read the 10-layer AI agent stack.

Why deep research matters more in 2026 than it did in 2025

There are two reasons. First, the volume of AI-generated content on the web has made single-source answers less reliable. When a search engine returns ten results, some of them may be generated by another AI, which means they repeat the same incorrect claims with different wording. Multiple-source verification catches this. If five sources all say the same thing and they are all derived from the same original flawed source, they only count as one independent claim.

Second, the tooling improved and you can now measure it. Google shipped Deep Research Max in April 2026 on Gemini 3.1 Pro, with MCP support so a research agent can reach internal systems as well as the open web. Mistral released Agentic Search in August 2026, a retrieval layer built for finding, inspecting and verifying information inside long documents. Independent measurement exists now too. The DeepResearch Bench team analysed 96,147 real user queries from a search-enabled chatbot, kept the 44,019 that needed multiple rounds of research, and built a 100-task benchmark across 22 fields to score agents on the reports they produce. Academic work points the same way. The March 2026 Marco DeepResearch paper calls the lack of explicit verification mechanisms a major bottleneck in training and scaling research agents.

Setting it up in your tools

Here is exactly how to set up deep research in the three most common agent platforms in mid-2026.

In Hermes

Deep research is configured through the MCP tool server that connects to a search API like SearXNG or Firecrawl. The agent receives a skill file that defines the loop structure, sub-question decomposition rules and confidence scoring criteria. Create ~/.hermes/skills/research/deep-research/SKILL.md with the loop instructions, the four-tier confidence scale and sub-question decomposition rules. The agent loads this at session start and applies it to every research request. Configure the search tools in ~/.hermes/config.yaml under mcp_servers, pointing to your chosen search API endpoints. For a complete breakdown of how skills and tools work together in an agent stack, see our guide on the 10-layer AI agent stack.

In Claude Code

You can wire the same pattern through a custom tool definition in your CLAUDE.md. Define a deep_research tool that the agent can call, which runs a search, reads the results, evaluates confidence and returns a structured answer. The loop logic lives in a bash script or Python module that Claude Code invokes. The key is defining the iteration count and confidence threshold in the tool description so the agent knows when to stop searching and start answering.

Using OpenAI Deep Research

OpenAI's built-in Deep Research mode handles the loop for you. It runs the search-and-read loop automatically and produces a cited synthesis. The limitation is you cannot customise the confidence scoring or sub-question decomposition logic. For quick answers this is fine. For research that needs to meet an internal audit standard, the Hermes or Claude Code approach gives you more control over the verification process.

Frequently asked questions

How do deep research agents verify their answers?

A deep research agent verifies its answers by corroborating every claim across independent sources before that claim goes into the answer. The agent reads each source in full, records every claim with its source URL and publication date, and scores it on a four-tier confidence scale. High confidence needs three or more independent sources that agree, medium needs two, low means a single source, and conflicting means the sources disagree, in which case both positions go into the answer with their links. It keeps running search-and-read iterations until the combined confidence clears a set threshold, and every claim in the final answer links back to its origin so a person can check the work.

What is a deep research AI agent?

A deep research AI agent is an automated system that searches multiple sources, reads full content, extracts specific claims, evaluates confidence levels across at least three independent sources, and iterates until it finds enough corroborated evidence. Unlike a standard chatbot, it shows its work with inline citations. By mid-2026, deep research mode had become a standard feature in the major AI assistants, from OpenAI and Google to Perplexity, because single-shot answers are not reliable enough for professional use.

How do you build a deep research agent?

Set a loop that plans sub-questions, searches, reads full pages, scores every claim and repeats until confidence clears your threshold. In code that is a while loop with four jobs inside it, which are plan the sub-queries, run the searches, extract and score the claims, then check the confidence rating before deciding to iterate again. The confidence threshold decides when the agent stops searching and starts writing, and the iteration ceiling stops it burning search credits on a question the web cannot answer. The loop rules, the four-tier scale and the iteration ceiling live in a skill file at ~/.hermes/skills/research/deep-research/SKILL.md. In Claude Code, wire it through a custom tool definition in your CLAUDE.md.

What separates deep research from regular AI search?

Regular AI search runs one query and summarises the top result. Deep research runs a multi-iteration loop: search, read, evaluate, refine, search again. The agent continues until it finds enough corroborating evidence, and every claim has a backlink to its source. A 2025 Stanford University study of purpose-built AI legal research tools found they still hallucinate between 17 and 33 per cent of the time, which is why a verification loop against multiple independent sources matters. OpenAI's Deep Research mode runs multiple search-and-read rounds before it answers, and cites its sources.

Which tools can you use to build deep research agents in 2026?

Hermes supports deep research via MCP tool servers connected to search APIs like SearXNG or Firecrawl, with loop logic in a skill file at ~/.hermes/skills/research/deep-research/SKILL.md. Claude Code supports custom tool definitions for search-and-iterate loops defined in CLAUDE.md. OpenAI's Deep Research mode offers built-in multi-step research with inline source citations. Perplexity's Pro Search queries multiple indices simultaneously. The essential infrastructure is a search API, a content extraction tool and a loop controller that manages iteration logic and confidence scoring.

How much time can a deep research agent save?

We could not find a credible public benchmark for time saved. What we can report from our own pipeline is what a run leaves behind: a ledger with a median of 32 sources, each tagged primary or secondary, plus a list of gaps that tells the writer what to leave out. Those gap notes stop weak claims before a draft exists, which saves the slower job of finding and pulling them from a published page.