A web research system you can trust with a current question does seven separate jobs. It defines what a complete answer looks like, discovers candidate sources, retrieves them, extracts evidence, checks that evidence for gaps, reconciles conflicts, and writes a cited result.
Splitting the work up this way won’t make an answer correct on its own. What it does is give every failure an address, so when an answer comes back wrong you can point at the stage that broke instead of blaming “the model”.
This is for anyone running their own model inside a product, an assistant, an agent or an internal workflow, where users ask about things the training data can’t know.
The loop
flowchart LR
A[1. Define\nAnswer contract] --> B[2. Discover\nCandidate sources]
B --> C[3. Retrieve\nReadable source content]
C --> D[4. Extract\nClaims plus context]
D --> E[5. Evaluate\nCoverage and gaps]
E -->|Material gap remains| B
E -->|Evidence is sufficient| F[6. Reconcile\nConflicts and uncertainty]
F --> G[7. Synthesise\nAnswer plus source record]
The arrow from Evaluate back to Discover is the part most implementations leave out, and it’s why I call this a loop and not a pipeline. When one piece of the answer has no evidence behind it, the system has three honest options: search again, narrow the claim, or hand back an incomplete result. Writing confident prose over the hole isn’t on that list.
Write the answer contract before you search
I’ll use one question the whole way through, because it’s the sort of thing people really do ask these systems:
Which web search approaches can an Ollama-based application use today, and what does each one return?
Before any searching, write down what a usable answer has to contain:
{
"question": "Which web search approaches can an Ollama-based application use today, and what does each one return?",
"required_elements": [
"approach name",
"where the web workflow runs",
"what the application receives",
"whether the model must still build the answer",
"a current source for each material product claim"
],
"output": {
"answer_format": "markdown",
"citations": "inline numbered references",
"source_record": "machine-readable ordered list",
"incomplete_evidence": "must be stated"
}
}
This feels like overhead the first time. It’s also the only thing that lets you decide afterwards whether a fluent answer was a complete one. A lot of “relevant but useless” answers come from nobody deciding up front what useful meant.
Stage 1: Define the question and what done looks like
“Web search approaches” is ambiguous. It could mean a raw search API, search plus page fetching, a tool loop the model drives itself, or a managed research call. Each reading is legitimate and each produces a different answer, so the first stage turns the request into a research objective, a set of subquestions, the required elements above, a freshness boundary and an output shape. “Today” in the question is doing real work here: it means current documentation, not a blog post from two years ago.
The test I’d apply is whether a reviewer could mark the answer complete or incomplete using only what this stage produced. If they’d have to invent criteria after reading the answer, the stage isn’t finished.
When the ambiguity would materially change the work, stop and ask. A silent guess at this point poisons everything after it, because every later stage will carefully execute the wrong job.
Stage 2: Discover candidate sources
Discovery finds pages that might contain evidence. Nothing is established as true yet.
A multi-part question needs more than one query. For the Ollama example I’d want one aimed at Ollama’s own search and fetch documentation, one at how Open WebUI wires search into a chat, and one at a research API’s response contract. A single broad query tends to return the loudest pages, and they’re rarely the authoritative ones.
Keep the provenance as you go: which query surfaced each page, and which subquestion it’s supposed to answer. Throw that away and any later debugging turns into guesswork.
This stage fails in a few recognisable ways. A required element has no candidate at all, the only candidates are derivative (the blog post summarising the docs rather than the docs), the results are duplicates of each other, or everything sits outside the freshness boundary.
Stage 3: Retrieve readable source content
A search result is a lead, and the page it points at is the evidence. Ollama’s web search API is a tidy illustration of the split: each result is a title, a URL and a content snippet, and a separate web fetch endpoint returns the page’s title, main content and links. The interface is honest about what search gives you.
Retrieval fetches or renders the chosen pages and turns them into something the next stage can inspect. Record the status, any redirects, a timestamp, the content size, and whether anything was truncated.
Truncation is the quiet failure. Open WebUI’s web search troubleshooting guide notes that web pages often run to 4,000 to 8,000 tokens or more, and that a model running at Ollama’s 4,096-token default context misses most of that content. Nothing errors when it happens. The model answers from part of a page and sounds exactly as sure of itself as it would with the whole thing.
A successful search followed by a failed fetch is a partial run with a known boundary, and it should be reported as one.
Stage 4: Extract claims with their context
A page can be perfectly readable and still useless as evidence. This stage pulls out the statements that bear on the question and keeps the qualifications attached to them.
At minimum I’d keep this much per claim:
{
"claim": "Ollama's web search API returns a title, URL and content snippet for each result.",
"source_url": "https://docs.ollama.com/capabilities/web-search",
"source_title": "Web search",
"supporting_excerpt": "content (string): relevant content snippet from the web page",
"retrieved_at": "2026-09-21T00:00:00Z",
"scope": "documented response fields for the web_search endpoint only",
"supports_required_element": "what the application receives"
}
The scope field is the one that earns its place. “The Ollama web search API returns snippets” is true. “Ollama-based applications only get snippets” is false, because the same page documents the fetch endpoint that returns page content. Extraction is where a qualified statement gets quietly promoted to an absolute one if nothing stops it.
It’s also where citations start meaning something. A list of URLs at the bottom of an answer is a bibliography. A record of which claim came from which page is something you can check.
Stage 5: Evaluate coverage, source quality and gaps
This is the stage that separates research from summarising whatever came back first. Compare the evidence against the contract you wrote at the start:
| Required element | Evidence found | Source quality | Fresh enough | Decision |
|---|---|---|---|---|
| Approach name | Yes | Official | Yes | Keep |
| Where workflow runs | Partial | Official plus inference | Yes | Narrow wording |
| Output returned | Yes | Official docs | Yes | Keep |
| Model still builds answer | No direct source | Architectural inference | N/A | Label as analysis, or search again |
| Current source | Yes | Official | Yes | Keep |
The row worth looking at is “model still builds answer”. You’re unlikely to find a provider page that states it outright, because it’s a conclusion drawn from what each interface returns. That’s fine, provided the answer labels it as analysis. What isn’t fine is attaching a citation to a page that doesn’t say it.
If a material gap remains, go back to discovery with a narrower query. Don’t pad the context with more loosely related text and hope the volume covers the gap. It won’t, and it makes the next stage harder.
Stage 6: Reconcile conflicts and uncertainty
Sources disagree for boring reasons. They describe different versions, plans, dates, regions or definitions, and usually neither one is wrong. Reconciliation keeps the conflict visible and works out whether it can be resolved:
- Compare publication and update dates.
- Check that both sources describe the same thing at the same version.
- Prefer the source closest to the claim. For product behaviour that’s usually the current official docs.
- Separate a factual disagreement from two sources using the same word differently.
- Narrow the claim when the evidence only supports a limited version of it.
- State the uncertainty when nothing authoritative settles it.
This is the stage I’d trust automation with least. Comparing dates and preferring official documentation is mechanical. Deciding whether two pages are describing the same object, or two objects with the same name, is judgement, and a system that quietly picks the convenient source will look fine in every demo you run. If you’re only adding human review in one place, put it here.
The failure mode is exactly that quiet pick, or its cousin: two incompatible claims blended into one sentence that neither source supports.
Stage 7: Synthesise the answer and keep the source record
Only now does the system write the thing the application asked for, and the result should carry more than prose:
{
"status": "complete",
"answer": "The approaches differ mainly in what they return and who operates the work after search. ... [1]",
"sources": [
{
"id": "src_01",
"url": "https://docs.ollama.com/capabilities/web-search",
"title": "Web search",
"claims": [
"Ollama's web search API returns a title, URL and content snippet for each result."
]
}
],
"unresolved": [],
"retrieved_at": "2026-09-21T00:00:00Z"
}
Two fields matter more than they look. status needs a value other than complete, otherwise a partial run has nowhere to go but into a confident answer. And an empty unresolved array should be something the system concluded, not the default because nobody wired it up.
Synthesis fails as a fluent answer with gaps in its support, citations that point at nothing in particular, no way to express a partial result, or output your application can’t parse.
The seven contracts on one page
| Stage | Takes | Produces | Check | Failure state |
|---|---|---|---|---|
| Define | User question, application requirements | Objective, subquestions, required elements, freshness boundary, output contract | Could a reviewer judge completeness from this alone? | Question left broad, undefined or unverifiable |
| Discover | Objective and subquestions | Candidate URLs with the query and subquestion behind each | Does every required element have a plausible candidate? | No source, derivative sources only, duplicates, stale candidates |
| Retrieve | Selected URLs | Readable content plus status, redirects, timestamps, size, truncation flag | Did we get the part of the page that holds the claim? | Blocked, timed out, unrendered, redirected elsewhere, empty or silently truncated |
| Extract | Retrieved content and answer criteria | Claim-to-source records with excerpt and scope | Does the excerpt support the claim at the same scope? | Snippet-only evidence, overbroad claims, lost qualifications, no trace back to the page |
| Evaluate | Claim records and required elements | Coverage matrix, source quality, duplicates, gap list | Is each element supported, qualified or marked unavailable? | Relevant pages that don’t cover the actual question |
| Reconcile | Supported claims and conflicting evidence | Resolved claims, scoped readings, unresolved items | Can a reviewer see why one source won? | Convenient source chosen quietly, incompatible claims blended |
| Synthesise | Reconciled evidence and output contract | Answer, citations, source record, limitations | Does every material claim have inspectable support? | Unsupported fluency, orphaned citations, no partial state, unparseable output |
The handoffs are where systems break
A visible success at one boundary tells you nothing about the next one.
| Visible event | What it proves | What it doesn’t prove |
|---|---|---|
| Search returned URLs | Discovery worked | The pages were read, or reached the model |
| Pages were fetched | Retrieval worked | The relevant claims were extracted |
| Context was assembled | Data reached a context layer | The model used it correctly |
| Answer contains links | Citation strings exist | The links support the claims beside them |
| API emitted complete | The job terminated successfully | The answer is correct or complete |
Two public Open WebUI issues from this year show the gap between the first and third rows. In #25038, on v0.9.5 with SearXNG, results came back and appeared in the interface, then source processing ended with “No sources found” and the model claimed web search wasn’t available. In #25585, after an upgrade from v0.9.5 to v0.9.6, the logs recorded nine items added to the web search collection while the model answered as though no search had happened.
These are individual, version-specific reports and I’m not using them as evidence of how often this happens. What they do show is that “search worked” and “the model used the search” are two different facts, and you need a separate check for each.
What should your application log?
Enough to find the failed boundary, without logging secrets or private source content. I’d capture:
- request ID and UTC timestamps
- a hash of the question, or an approved public test question
- the research objective and required answer elements
- every search query issued
- counts of candidate, fetched, failed, distinct and cited URLs
- redirects, fetch status and truncation flags
- context size and the source IDs passed to synthesis
- terminal status, and the stage it failed at if it failed
- where the answer and source record were written
- dependency and model versions
- duration, and usage figures where the provider reports them
Leave out API keys, authorisation headers, cookies, environment dumps, model reasoning and anything confidential the user sent.
Where should the loop live?
It helps to separate the tools first, because they get lumped together. Search returns links and snippets. Fetch or extract reads a URL you already have and gives you clean content or specific fields. Research takes a question and returns a synthesised answer from sources. Automation performs interactions on a public website when the result needs someone to click through to it.
| Approach | Best when | Your application owns |
|---|---|---|
| Search API | Links or snippets are the result you want | Any reading, synthesis and citation |
| Model-operated tools | Your model drives tools reliably and you want the control | Tool orchestration, context, evaluation, retries |
| DIY research loop | The web layer is what differentiates your product | All seven stages and every boundary between them |
| Managed Research API | You need a cited result without operating the loop | Question design, acceptance, review, downstream use |
My default is not to operate all seven stages yourself unless owning the loop is the product, or you need control over a specific stage that nothing managed gives you. Moving the loop to a provider changes who operates it. Deciding whether the output is good enough for your job stays with you either way.
What a managed call shows you of the seven stages
Tabstack Research is the one I work on, so here’s how its output lines up against the stages. Per the Research guide, you send a question and it always streams Server-Sent Events: a start, a planning phase, one or more search iterations, a writing phase, then a single complete event carrying the report and its metadata.
| Stage | Where it shows up |
|---|---|
| Define | researchObjective, researchQuestions and researchPlan in metadata, always present |
| Discover | executedQueries, plus sourceQueries on each cited page |
| Retrieve | totalPagesAnalyzed, with per-page fetch time bounded by the fetch_timeout parameter |
| Extract | claims on each entry in citedPages |
| Evaluate | gapEvaluations when populated, plus isLast and stopReason on each iteration:end event |
| Reconcile | No dedicated field. You see the outcome in the report, not the working |
| Synthesise | report, as Markdown |
The reconcile row is the honest gap, and it’s the reason the review step stays with you for anything that matters. citedPages and gapEvaluations are both optional, so read them defensively and treat a missing citedPages as an empty list.
The operational details are where a first integration usually trips. HTTP errors such as a bad key or a rate limit are raised before the stream opens. Task failures arrive as an error event inside the stream, so a consumer that only listens for complete gets nothing at all from a failed run. You also need to handle a stream that closes without either. There’s no server-side timeout on the request as a whole: fast mode typically finishes in under a minute, while balanced consults more sources, needs a paid plan, and can run for up to about four minutes on the broadest questions. A fixed total timeout will kill healthy long runs, so reset a timer on every event and fail on silence instead.
Map your current research loop
Take one current question your application has to answer and fill in these seven rows before you change any tools:
| Stage | Current owner | Input | Output | Check | Failure state | Logged? |
|---|---|---|---|---|---|---|
| Define | ||||||
| Discover | ||||||
| Retrieve | ||||||
| Extract | ||||||
| Evaluate | ||||||
| Reconcile | ||||||
| Synthesise |
Any blank cell in the owner column is work somebody assumed somebody else was doing. That’s the next thing to instrument, or the next thing to hand off.
Never trained on. Private by default. Built by Mozilla.



