NEVER TRAINED ON · PRIVATE BY DEFAULT · BUILT BY MOZILLA

GITHUB

Sources to answers

Where model-operated web research breaks

Nine places web research fails after search has already succeeded, each with a signature you can observe, the evidence that identifies it, and the smallest test that catches it.

Search returned ten results, the interface listed them, and the model answered as though it had never heard of the internet. That failure has at least four possible homes, and the one thing you can be sure of is that it isn’t in the search provider.

Web research fails after search succeeds more often than it fails at search. Discovery, retrieval, extraction, context assembly, the tool loop, synthesis and citation are separate boundaries with separate failure modes, and a visible success at one of them tells you nothing about the next. What follows is nine classes, each with a signature you can observe, the evidence that separates it from its neighbours, and the smallest test that catches it.

Every public example below is a single report in a named version and configuration. They show the classes are real and distinct from each other. They say nothing about how often any of them happens, and I’m not using them to make a claim about any project’s current state.

flowchart LR
    Q[Question] --> P[Plan]
    P --> S[Search]
    S --> F[Fetch / render]
    F --> X[Extract / normalise]
    X --> C[Assemble context]
    C --> M[Model tool loop]
    M --> A[Synthesise answer]
    A --> R[Attach source record]

When you write up an incident, name two boundaries: the last one you can prove worked, and the first one you can prove didn’t.

1. Scope and planning

The answer is about the right topic and doesn’t do the job. A comparison names the products and never says what each one returns. A question about current behaviour comes back built from undated pages. A three-part question gets its easiest part answered well.

This is the only class that happens before anything runs. Nobody wrote down what a complete answer contains, so the loop stops when it finds relevant text rather than sufficient evidence, and whoever reviews the output invents the criteria after reading it. By then the answer sets the standard it’s judged against, which is exactly backwards.

Freeze the required elements first, then score each one as answered, partly answered or missing. A well-written answer earns no points for coverage it doesn’t have.

2. Discovery

Search comes back with too few candidates, the same domain five times, stale pages, or results that cover one subquestion and ignore the rest. Open WebUI discussion #18302 shows the symptom plainly: on v0.6.32 the reporter saw one or two sources fetched for most queries, and short, thin answers as a result.

The cause is usually one broad query standing in for a multi-part question, or a loop that stops after the first relevant page instead of searching again while coverage is still short.

Keep every query you issued, the ranked results with their domains, and which required element each candidate might support. Discovery passes when every required element has a plausible candidate. A non-empty result list is not the bar.

3. Retrieval

The URLs are visible and the pages never reach the next stage. The dangerous version is the one where the interface lists the URLs, everyone assumes the pages were read, and the model answers from what it already knew.

This class is quiet by design, because fetch failures are usually caught and skipped. In the thread on Open WebUI #25038, a commenter traced the “No sources found” behaviour on v0.9.5 to a keyword argument passed twice in the page loader. It raised a TypeError on every URL before any network request, and because the loader runs with continue_on_failure=True, the only trace was a warning per skipped fetch. Search worked. The UI listed results. Nothing was read.

Open WebUI’s troubleshooting guide describes a quieter relative: Wikipedia, Cloudflare-protected sites and other large publishers filtering the default Python user agent, so pages come back empty or as 403s.

Per fetch, keep the requested and final URL, the status or browser outcome, content type, size, redirect chain, duration and a truncation flag. Then test it properly: pick a public page containing a unique sentence, fetch it through your production path, and check the sentence survives into stored content. A fetch counter going up proves nothing at all.

4. Extraction and normalisation

The page was fetched and the useful part is missing, mangled, or buried under navigation.

Ollama issue #12690 shows what happens when nothing selects. On v0.12.6, a web search through Qwen3 in the desktop app returned 102,207 tokens for a simple question, because it passed the full contents of five pages, one of which was a Wikipedia article of roughly 100,000 tokens on its own. The same question through gpt-oss:20b-cloud came back at 768 tokens. For scale, Ollama’s web search docs suggest a context of at least around 32,000 tokens, so one unfiltered page can be three times the whole budget.

Keep hashes of the raw and normalised content, both lengths, the extraction method and version, and where truncation happened. Test with one expected passage and one known block of boilerplate: the normalised output keeps the passage, loses the boilerplate, and fits a budget you declared before the run.

5. Context assembly

Search and retrieval both succeed and the model behaves as though it has no evidence at all. This is where I’d put the most monitoring, because every check upstream of it passes.

Open WebUI produced two versions of it in consecutive releases. In #25038, on v0.9.5, results appeared in the interface, then source processing ended with “No sources found” and a traceback from the web search handler, where a JSONResponse was being indexed as if it were a dictionary. In #25585, after an upgrade to v0.9.6, the logs recorded nine items added to the web search collection, no errors anywhere, and a model answering as though no search had run. Commenters traced that one to source resolution: the function had no branch for web search items, and after an access-control change the fallback treated the collection name as untrusted and dropped it silently. Status history showed zero sources retrieved. A fix was proposed in pull request #25600. #25038 has since been closed, with the closing comment saying the reported problems were fixed by 0.10.2 given correct configuration, though a later comment describes similar behaviour on v0.11.3.

Both reports are tied to their versions and setups. What makes them worth reading is that a search provider health check would have passed throughout, and in #25585 the logs actively reported success.

So keep source IDs at every hop: candidates, retrieved, placed in context, present in the final model input. The test is a sentinel. Put one unique, harmless fact into a controlled source and check that the exact fact reaches the model input and can be returned with its source ID. Never infer delivery from a “search complete” event in the UI.

6. Tool loop and state

The model calls search, then loses the question, loops on the same tool, or stops with an empty reply.

Ollama issue #15895 is a clean specimen. On v0.22.0, Gemma 4 reasoned for a while, called web search, and once the tool result arrived it produced an empty reply, having lost both the prompt and the chat context. A follow-up on the thread put it down to truncation rather than compaction, with the user’s own message falling outside the window, and another comment noted Qwen3 wasn’t affected in the same setup.

Keep the ordered message sequence with roles, tool-call IDs matched to their results, context size before and after each call, the stop reason and the final message length. Then test the message adapter away from live search: replay a fixture with the question, one tool call and one small result, and check the model can restate the question, use the result and produce a non-empty answer.

7. Synthesis, coverage and uncertainty

The sources arrived and the answer still misses required elements, blends facts that conflict, or sounds more certain than the evidence allows.

Most systems have no coverage check between “context assembled” and “start writing”, and a writing prompt that rewards finishing over admitting a gap will always find a way to finish. Two sources describing different versions get treated as interchangeable, and the first plausible story wins.

Build a fixture with two sources that explicitly disagree and one required element with no evidence at all. A passing answer scopes the conflict and says the evidence is missing. A failing one resolves the disagreement by inventing a reason to prefer one side.

8. Citation integrity

There are links, there are numbered references, and a cited page doesn’t support the claim sitting next to it. The variants: important claims with no citation, URL variants of one page counted as separate sources, and a source list with no mapping back to claims.

Our own Week 1 sample is a fair example, and I’d rather use it than pretend managed output is immune. The committed run used Tabstack Research in fast mode with SDK 2.8.5 on 15 September 2026. It completed cleanly and cited seven pages. Three of those seven were the same Ollama documentation page under different URLs: once over http, once over https, once with a .md suffix. Every claims array came back empty, and the optional relevance and reliability fields weren’t there. The Research guide’s worked example shows populated claims; this run didn’t return them.

None of that means the answer was wrong. It means three separate measurements were being read as one: whether citations are present, how many distinct sources there are, and whether each claim is supported.

Check every material claim against its cited page and score the support as direct, partial, absent or inaccessible. Report correctness and completeness separately, and don’t award a pass because the links are there.

9. Terminal and operational

The interface searches forever, a stream closes before any terminal event, an HTTP failure gets handled as a task failure, or missing output becomes a successful empty answer.

Most of these live in the consumer. A handler that only listens for success turns every failed run into silence. HTTP errors and in-stream errors sharing a handler lose the information about which phase died. A fixed total timeout kills healthy long runs. Retries can repeat work and cost, and they can be on without anyone choosing them: the Tabstack Python SDK retries a request twice by default on connection errors and on 408, 409, 429 and 5xx responses.

Replay three fixtures at minimum: a successful complete, an in-stream error, and a stream that ends with neither. Each should produce a distinct result, and both failures should exit non-zero.

The nine, on one page

Class Signature Minimum evidence Smallest test
Scope and planning Relevant answer omits required parts Frozen question and required elements Score coverage against criteria written first
Discovery Too few candidates, or a narrow source set Queries and result list Run fixed queries and inspect candidates
Retrieval URLs exist, pages absent, blocked or truncated Status, redirects, size, truncation Fetch a known page through the production path
Extraction Page fetched, evidence missing or malformed Raw and normalised content hashes Compare an expected passage with the output
Context assembly Fetch succeeds, model gets no usable sources Source IDs at each hop Inject a sentinel fact, inspect model input
Tool loop and state Question lost, tools repeated, empty stop Ordered messages and tool events Replay a frozen conversation fixture
Synthesis and coverage Answer incomplete or overconfident Required elements and the answer Score coverage and uncertainty handling
Citation integrity Links don’t support claims, sources duplicated Claim-source pairs, normalised URLs Check each material claim against its citation
Terminal and operational Stream closes, times out, or errors with no result Event sequence, timestamps, terminal status Replay complete, error and premature-close fixtures

They cascade

Two of the reports above fit together into a chain worth recognising. #12690 shows a single search returning six figures of tokens. #15895 shows what truncation does to the user’s message once the context overflows. They come from different setups, so read this as how the classes connect rather than as one incident:

flowchart TD
    A["Extraction passes ~100,000 tokens of page content"] --> B["the context budget is exceeded"]
    B --> C["earlier messages are truncated, including the question"]
    C --> D["the model appears to #quot;forget#quot; what it was asked"]
    D --> E["the answer is empty or unrelated"]

The visible symptom sits four steps from the cause, and the symptom points at the model. Diagnose from the last boundary with positive evidence instead.

Triage order

  1. Did the application define the required result? If not, freeze the acceptance criteria before touching anything else.
  2. Did search return plausible candidates for every required element?
  3. Were the selected pages actually retrieved? Check status, redirects and truncation.
  4. Did normalisation keep the evidence? Compare content before and after extraction.
  5. Did that evidence reach the model input? Check source IDs and a sentinel.
  6. Did the tool conversation keep the question and the message order?
  7. Did the answer cover the required elements and state its uncertainty?
  8. Do the citations support the material claims?
  9. Did the run reach an explicit terminal state? A non-empty UI is not evidence of success.

The order exists to stop you swapping the model when the evidence never reached it. That’s the most expensive wrong move available, and it’s the first one most teams make.

Moving the loop out of the model

A managed Research call takes query planning, retrieval, gap checks, synthesis and citation assembly out of your code. Classes 2 to 6 stop being yours to instrument, which is a real reduction in surface area. Failure still happens, in fewer places you control.

Class 1 stays with you, because you still own the question. Classes 7 and 8 stay with you as review work, and the Week 1 sample above is the argument for taking that seriously. Class 9 moves into your consumer: request validation, HTTP errors, in-stream errors, a stream that closes early, optional citation metadata that isn’t there.

I’d use search or model-operated tools when you want links, want control over each step, or the loop itself is the product. I’d use a managed call when the application needs a finished cited result and you’d rather not have the model operating every intermediate stage. Either way, a complete event proves the job terminated and nothing more.

Write the next one up properly

When the next web-enabled answer fails, record the environment and exact versions, the frozen test question or a sanitised hash of it, the answer elements you expected, the last boundary known to have worked, the first known to have failed, sanitised logs or fixtures, whether it reproduces, and whether the workaround changes the architecture or just hides the symptom. Then give it one primary class and note the downstream effects.

A useful incident report lets another engineer reproduce the boundary. A bad one shows them the answer and asks them to guess.

Never trained on. Private by default. Built by Mozilla.

START FREE

Read the guide, then make the call.

Start with 10,000 free credits. No credit card required.

curl -fsSL https://tabstack.ai/install.sh | sh