NEVER TRAINED ON · PRIVATE BY DEFAULT · BUILT BY MOZILLA

GITHUB

Sources to answers

Search gives your model sources. What turns them into a cited answer?

A web search API can return useful titles, URLs, and excerpts. A cited answer requires a research loop after search: define the question, select and read sources, identify gaps, reconcile evidence, write the answer, and preserve the claim-to-source connection.

A web search API can return useful titles, URLs, and excerpts. That is enough when your application needs search results or your model can reliably do the remaining work. A cited answer requires a research loop after search: define the question, select and read sources, identify gaps, reconcile evidence, write the answer, and preserve the claim-to-source connection.

Who this is for: People running their own model inside a product, assistant, agent, or internal workflow that needs current answers from the public web.

In this guide, you will learn:

  • what search returns and what it leaves to your system;
  • the seven stages between a question and a cited answer;
  • the difference between search, fetch, extraction, research, and automation;
  • when search alone is sufficient;
  • what a research-ready output should contain.

Search is access. A cited answer is a pipeline.

Search and research solve different parts of the same job. A search endpoint usually begins with a query and returns candidates: titles, URLs, snippets, or excerpts. Those results help a person or program decide what to read next. For example, the current Ollama Web Search API documents a response containing title, url, and a relevant content snippet for each result. Brave’s Web Search API similarly returns result objects with a title, URL, description, and optional extra snippets.

That output may be exactly right. A link finder, source browser, retrieval component, or model with a dependable tool loop can use it directly. But if the application needs an answer it can show, store, or act on, somebody still owns the work after search:

  1. decide what the question actually requires;
  2. choose which results are worth reading;
  3. retrieve the relevant pages;
  4. extract the useful evidence;
  5. check whether the evidence is sufficient;
  6. resolve conflicts and qualify uncertainty;
  7. write the answer and attach the sources.

A cited answer is therefore not a formatting option on a list of search results. It is the output of a research system.

The seven-stage research loop

flowchart LR
    A[1. Define\nQuestion and success criteria] --> B[2. Discover\nCandidate sources]
    B --> C[3. Retrieve\nRead the pages]
    C --> D[4. Extract\nRelevant claims and context]
    D --> E[5. Evaluate\nCoverage, quality, and gaps]
    E -->|Evidence is incomplete| B
    E -->|Evidence is sufficient| F[6. Reconcile\nResolve conflicts and uncertainty]
    F --> G[7. Synthesize and cite\nAnswer plus source record]

The arrows matter more than the boxes. Research is not always one pass from left to right. If the evidence does not answer part of the question, the system needs another query, another source, or a narrower claim. That loop is what separates research from “summarize the first result.”

1. Define the question and success criteria

The system first turns the user’s request into an answerable research objective. Consider this question:

What are the main approaches to browser automation for AI agents, and how do they differ?

A useful answer needs more than definitions. It should establish the comparison dimensions, such as browser control method, deployment model, supported interaction, and who owns the infrastructure. It may also need to distinguish open-source libraries, hosted browser infrastructure, and higher-level task APIs. Without a clear objective, a system can collect relevant pages and still produce an incomplete answer.

2. Discover candidate sources

Search finds pages that may contain evidence. One broad query is rarely enough for a multi-part question, so the system may create narrower searches for each comparison dimension. Discovery should preserve where each candidate came from and why it may be relevant. At this stage, a search result is a lead, not yet evidence for a final claim.

3. Retrieve and read the pages

Snippets are useful for triage, but they are not a substitute for the source page. They can omit qualifications, dates, tables, footnotes, or the paragraph that changes the meaning of a claim. The system therefore fetches the selected pages and converts them into content it can inspect. Some pages can be read through a direct HTTP request. Others require rendering before the meaningful content is available. Failed, blocked, truncated, or outdated pages need to be recorded rather than silently treated as supporting evidence.

4. Extract claims with their context

Reading a page is not the same as identifying the evidence that answers the question. The system must isolate the relevant statement, retain its source URL, and preserve enough context to avoid turning a qualified claim into an absolute one. For a comparison, this may mean extracting:

  • the specific capability being documented;
  • the conditions or plan on which it is available;
  • the publication or update date;
  • the exact source page;
  • the claim the source supports.

The goal is a claim-to-source record, not a pile of page text.

5. Evaluate coverage, source quality, and gaps

The system checks the collected evidence against the original question. Which parts are answered? Which are partial? Which have no useful source yet?

This stage should also ask whether the sources are appropriate for the claim. Official documentation may be the best source for current API behavior. A maintainer issue may be better evidence for a known failure mode. A third-party comparison can provide a useful lead, but it should not quietly become proof of a competitor’s current limits.

If a material gap remains, the loop returns to discovery with a more specific query.

6. Reconcile conflicts and qualify uncertainty

Sources disagree for ordinary reasons: they cover different dates, product tiers, regions, definitions, or versions. The system should not choose whichever sentence is easiest to quote. Reconciliation means comparing the scope of each claim, preferring the most direct and current evidence, and stating uncertainty when the sources do not support a single conclusion. When the evidence remains incomplete, the answer should say so.

7. Synthesize the answer and preserve citations

Only after the evidence is collected and evaluated should the system write the final answer. The result should separate what the sources establish from the system’s interpretation. Citations should let a reader inspect the evidence behind a material claim. A bibliography at the bottom is useful, but it is weaker than a record that connects individual claims to the pages used to support them.

What goes in and what should come out?

Here is a complete input for a research system:

{
  "query": "What are the main approaches to browser automation for AI agents, and how do they differ?",
  "mode": "fast"
}

A useful research output is not just prose. It should contain the answer and enough source metadata to inspect how the answer was built. Tabstack’s current /research response represents that contract with a report plus citation metadata:

{
  "report": "# Browser automation for AI agents\n\nThe main approaches differ in who controls the browser, where it runs, and how much orchestration the application owns. [...]",
  "metadata": {
    "citedPages": [
      {
        "id": "pg_a1b2c3",
        "url": "https://source.example/official-documentation",
        "title": "Official documentation",
        "claims": [
          "A specific statement used in the report."
        ],
        "sourceQueries": [
          "browser automation approaches for AI agents"
        ]
      }
    ],
    "totalPagesAnalyzed": 12
  },
  "message": "Research complete",
  "timestamp": 1789420248496
}

Example qualification: This response is illustrative and combines the documented /research response fields with placeholder content. It is not the output of a live research run, and the page count is not a benchmark or product claim.

The output contract matters because different consumers need different artifacts.

Output taxonomy: choose the artifact your application needs

Output What the application receives What still needs to happen Good fit
Search results Ranked titles, URLs, and excerpts Select, read, compare, synthesize, and cite Discovery, link finding, model-operated research
Clean page text Readable content from one known URL Interpret the page and combine it with other sources RAG ingestion, reading a page, downstream analysis
Matching JSON Fields returned in a defined schema Validate business rules and decide how to use the record Known-page extraction, enrichment, pipelines
Cited answer Synthesized answer plus source references Review evidence and apply domain judgment Current questions that span multiple sources
Completed public web task Progress and the result of browser interaction Confirm the result and handle downstream side effects Navigation, clicks, forms, and multi-step public workflows

These outputs are not quality levels. They are different contracts. Returning a cited report when an application needs ten links adds unnecessary work. Returning ten links when the application needs an answer transfers the research loop into the model or application.

What is the difference between search, fetch, extraction, research, and automation?

Web search returns candidate sources relevant to a query. A typical result includes a title, URL, and excerpt. Search is sufficient when the list of sources is the desired output or when another component reliably owns the rest of the research loop.

What is fetch?

Fetch retrieves the content at a known URL. It answers “read this page,” not “answer this open question.” Fetching may be direct or may require rendering, but it does not by itself establish which other pages are needed or reconcile evidence across them.

What is extraction?

Extraction turns a known page into a cleaner or more structured representation. That may be readable Markdown or JSON matching a supplied schema. Extraction is the right abstraction when the source is known and the application knows what fields or content it needs.

What is research?

Research begins with a question and produces a synthesized answer from one or more sources. It includes source discovery, reading, gap detection, reconciliation, and citation. A research API is useful when the application needs the finished answer rather than material for its model to process.

What is automation?

Automation completes a task that requires interaction with a website. It applies when the system must navigate, click, fill forms, or move through several steps before the requested result exists. Reading or researching a public page does not automatically require browser automation.

When is search alone enough?

Search alone is often the simpler and better choice when:

  • the user wants links, not an answer;
  • your model already performs the tool loop reliably;
  • the question is simple and one source is enough;
  • you need full control over ranking, retrieval, or synthesis;
  • the research loop is part of your product’s core differentiation;
  • you are building a source-discovery interface rather than an answer feature.

A managed research loop is more useful when:

  • the application must return a current, cited answer;
  • the question spans several sources or sub-questions;
  • incomplete coverage needs to trigger another search;
  • citations must remain connected to specific claims;
  • you do not want the model to spend its context and tool calls operating the loop;
  • your team would otherwise maintain search, retrieval, synthesis, and citation orchestration.

Neither path removes the need to evaluate the result. A citation makes a claim inspectable. It does not guarantee that the source is authoritative, that every claim is supported, or that the answer is complete.

The architecture choice is about ownership

Approach Application receives Who owns source selection and reading? Who owns gap detection and reconciliation? Who owns citation assembly?
Search API Results and excerpts Your model or application Your model or application Your model or application
Search + fetch tools Results plus page content Your model or application Your model or application Your model or application
DIY research loop Whatever contract you build Your team’s orchestration Your team’s orchestration Your team’s orchestration
Managed Research API Synthesized answer and source metadata The API inside the call The API inside the call The API inside the call

A managed API can reduce the orchestration in your codebase, but it does not transfer responsibility for the result. Your team still owns question design, acceptance criteria, domain-specific review, and how the application uses the answer.

The decision is not “build versus buy” in the abstract. It is where this loop should live, how much control you need, and which output your application can consume safely.

A concrete managed Research call

Tabstack is one implementation of the managed path. The application sends a question to /research; the endpoint streams progress with Server-Sent Events, then emits a complete event containing the report and metadata for cited pages.

import Tabstack from "@tabstack/sdk";

const client = new Tabstack();

const stream = await client.agent.research({
  query: "What are the main approaches to browser automation for AI agents, and how do they differ?",
  mode: "fast",
});

for await (const event of stream) {
  if (event.event === "error") {
    throw new Error(event.data.error.message);
  }

  if (event.event === "complete") {
    console.log(event.data.report);

    for (const source of event.data.metadata.citedPages ?? []) {
      console.log(`${source.title ?? "Untitled source"}: ${source.url}`);
    }
  }
}

Environment: Node.js 20 or later; @tabstack/sdk 2.8.6; TABSTACK_API_KEY set in the environment.

Last technically validated: September 14, 2026, against the current TypeScript SDK source and published Research guide.

Validation boundary: The snippet’s types and API surface were checked against SDK commit 366182e084a7d5fd372cde47e8bc2e96e63f5e16. A live fast-mode request was run on September 14, 2026 against the production API; the complete event matched the shape shown above, including the cited page fields.

Failure modes a cited-answer system must expose

A relevant result is not necessarily a usable source

A search result may point to an outdated, derivative, blocked, or incomplete page. Record source failures and continue rather than treating every candidate as evidence.

Snippets can remove the qualification that matters

A snippet may mention a price, limit, feature, or policy without its date or plan. Read the source before using the statement in an answer.

More sources do not automatically improve the answer

Ten pages repeating the same secondary claim do not provide ten independent confirmations. Source authority, independence, recency, and directness matter.

Citations can be present but poorly matched

A link beside a paragraph does not prove that it supports every sentence in the paragraph. Preserve the specific claims associated with each source and test citation correctness separately from citation presence.

The evidence can remain incomplete

Some questions have no public answer. Others have conflicting or rapidly changing evidence. The system needs an explicit partial or unanswered state rather than filling gaps with plausible prose.

The stream can fail before completion

For Tabstack /research, HTTP-level failures can occur before the stream opens, while task-level failures arrive as error events inside the stream. A consumer that only listens for completeness can mistake a failed run for an empty answer.

Trust and data flow

A managed research call sends the question and relevant processing data outside your application. Treat that as an architecture decision, not a footnote. Review which providers process inputs and outputs, what telemetry is retained, whether history can be stored, and how personal or confidential information should be handled.

For Tabstack, the approved summary is: Never trained on. Private by default. Built by Mozilla.

The current Tabstack Privacy Notice explains that Mozilla operates the service, that inputs and outputs may be processed through third-party models for applicable services, and that users should be careful about placing confidential or sensitive information in inputs. Review the notice for the full data flow rather than interpreting the summary as a promise that nothing is processed or retained.

What should you evaluate before choosing a research path?

Use one real question from your product, not a generic demo. Then inspect:

  1. Answer correctness: Are the material claims accurate?
  2. Citation correctness: Does each cited page support the associated claim?
  3. Citation completeness: Are important factual claims cited?
  4. Source quality: Are the sources direct, current, and appropriate?
  5. Coverage: Did the answer address every part of the question?
  6. Failure behavior: Does the system expose missing evidence, blocked pages, and task errors?
  7. Operational fit: Can your application handle streaming, retries, latency, and cost for the job?

This is the practical difference between adding search and shipping a current-answer feature. The first gives your system access to sources. The second requires a research loop with an output contract, evidence handling, and failure states.

Evaluate a real Research question

Choose one current question your model or application must answer today. Run it through your existing search-and-fetch path and through a managed Research call. Compare the outputs at the claim level, including missing claims and failed sources, rather than choosing the answer that sounds better.

Make a Tabstack Research call or review the Research guide before you design the evaluation.

  • Next, Build: Build a cited live-web answer in Python with one Research call.
  • Next, Prove: Search + fetch + model versus a managed Research call: the evaluation plan.
  • Tabstack Research guide
  • Research API reference
  • TypeScript SDK repository

START FREE

Read the guide, then make the call.

Start with 10,000 free credits. No credit card required.

curl -fsSL https://tabstack.ai/install.sh | sh