NEVER TRAINED ON · PRIVATE BY DEFAULT · BUILT BY MOZILLA

GITHUB

Sources to answers

Search + fetch + model versus a managed Research call: the evaluation plan

We are comparing two complete ways of answering current public-web questions: a pinned search + fetch + model pipeline the application runs itself, and a single managed Research call. This article publishes the pre-registered method. It does not publish a winner.

We are comparing two complete ways of answering current public-web questions. One is a pinned search + fetch + model pipeline that the application runs itself. The other is a single managed Research call. The questions, the system configurations, the hypotheses, the metrics, the failure rules and the analysis plan all went into Git before either system produced an output anyone looked at. This article publishes that method. It does not publish a winner, because the scored run has not happened yet.

It is written for people running their own model inside a product, an assistant or an internal workflow, who have to decide where the research loop lives.

The protocol went in first

flowchart TD
    A[Define question] --> B[Freeze both systems]
    B --> C[Commit hypotheses, 12 questions, rubric, and failure rules]
    C --> D[Run a one-question harness pilot]
    D --> E[Confirm logging and blind packaging only]
    E --> F[Run the full evaluation later]
    F --> G[Score answers before revealing system identity]
    G --> H[Report each metric, every failure, and the limits]

Protocol v1 is committed at 116d90191c6d900a97779ac0fc07dc9ffcfc229a. The application-operated system was frozen at 22c1613a8e2546acb429c818d61a5327706004ed. The protocol, question set, rubric, result schema and harness are public in the repository’s evals directory. FULL_EVALUATION_RUN=false.

The question

For current public-web questions, what differences appear between a pinned search + fetch + model pipeline the application operates, and one managed Research call?

That compares two complete configured systems, and it is worth being blunt about the limit this creates. If the outputs differ, we cannot point at search, page retrieval, model choice, orchestration, source availability or the managed service and say which one caused it. Isolating a component needs a different experiment, and a controlled one.

The decision underneath is architectural. Does your application own source discovery, retrieval, context assembly, synthesis and citation formatting, or does it call a managed web API and take back the finished result?

The input and output are held constant

Every run receives one frozen question and the same output instruction:

Answer the question directly.
Support material factual claims with citations to public source URLs.
Distinguish sourced facts from interpretation.
State when evidence is incomplete or conflicting.
End with a Sources section listing only sources used in the answer.

Each run must produce a record with:

{
  "question_id": "Q01",
  "system_id": "redacted during blind review",
  "started_at_utc": "2026-09-15T20:31:39.716Z",
  "completed_at_utc": "2026-09-15T20:31:56.383Z",
  "duration_ms": 16666,
  "terminal_status": "complete",
  "answer_path": "answer.md",
  "sources_path": "sources.json",
  "event_log_path": "events.sanitized.jsonl",
  "usage_receipt_path": "usage-receipt.json",
  "pilot_only": true,
  "scoring_status": "unscored"
}

That is a pilot record, included here to show the harness contract. It is not a scored result.

Both systems were frozen before the test

We are comparing full paths, not an abstract category against a product.

System A: application-operated search + fetch + model

System A uses a pinned web-search API, HTTP page retrieval with redirects and a 30-second fetch timeout, standard-library HTML-to-text conversion, and a dated flagship general model behind a chat-completions interface. The planner and writer prompts are committed. There are no application-level retries. Its fixed resource limits are:

max_search_iterations: 3
max_queries_per_iteration: 3
max_results_per_query: 5
max_pages_fetched_total: 12
max_chars_per_page: 12000
max_total_source_chars: 60000
total_wall_clock_timeout_seconds: 300
application_retries: 0

Model choice is the part of System A most open to challenge, so here is the reasoning behind it. We picked a configuration meant to represent a serious per-query deployment. We did not pick a smaller tier that would quietly weaken the application-operated baseline, and we did not pick a higher-cost specialist tier that a team would never run on every query. The complete provider and version record stays in the committed system configuration.

System B: Managed Research

System B uses:

endpoint: /research
mode: fast
nocache: true
fetch_timeout: default
application_retries: 0
python_sdk: tabstack==2.8.5

The application makes one /research call and receives the streamed report and the cited-page metadata.

Both configurations are part of the public record. If either one changes after anyone has inspected outputs, the protocol version has to change and the full test restarts.

The question set includes work that may not need a research loop

The frozen set contains 12 questions:

Question type Count Why it is included
Simple source discovery 3 Test jobs where a list of official sources may be enough
Multi-source synthesis 5 Test questions that require combining several pieces of evidence
Freshness-sensitive 2 Test whether sources and claims meet an explicit date requirement
Conflict or incomplete evidence 2 Test how each path handles uncertainty and unresolved evidence

Each question carries two to five required answer elements, written before execution. The set deliberately includes jobs where a list of official sources might be all anyone needs. A comparison built only around the work one system prefers would not answer the architecture question honestly.

The complete set is in evals/questions.jsonl.

The hypotheses are directional, not conclusions

Five, registered before the run:

  1. The managed path will have higher median citation completeness, because citations are part of its output contract.
  2. The managed path will have higher median question coverage on multi-source and conflicting-evidence questions.
  3. The application-operated path may have lower median wall-clock time on simple questions.
  4. Cost has no pre-registered direction.
  5. On pure source-discovery questions, the application-operated path will be no worse on usefulness within this small sample.

Any of these can be wrong. Publishing them beforehand is what makes being wrong visible.

Separate metrics, no composite score

A single number would conceal the tradeoffs we are trying to see, so every metric is reported on its own.

Answer correctness. Per material factual claim: 2 supported and accurate, 1 partly supported, imprecise, or missing a necessary qualification, 0 contradicted or unsupported. If a qualified evaluator or primary-source adjudication is unavailable, correctness stays unavailable rather than being guessed.

Citation correctness. Per claim-citation pair: 2 the cited source directly supports the claim, 1 the source is relevant but support is partial or indirect, 0 the source does not support the claim or cannot be inspected.

Citation completeness. Cited material factual claims divided by material factual claims. Opinions, transitions and interpretation that is clearly marked as interpretation do not require citations.

Question coverage. Per pre-listed required element: 2 fully answered, 1 partly answered, 0 missing.

Source quality. Per cited source: 3 primary or official source for the claim, 2 strong independent secondary source, 1 weak, derivative, undated or unclear, 0 irrelevant or unusable. We also record source-domain diversity, without assuming that more domains mean better evidence.

Freshness. Where a question defines a date window: 2 source and claim meet it, 1 the date is unclear and no contradiction was found, 0 the evidence is outside the window or stale.

Uncertainty handling. Whether the answer names material gaps and conflicts and scopes its conclusion, instead of turning incomplete evidence into certainty.

Latency. Client-observed wall-clock milliseconds for every attempt, and first-event latency where the interface exposes it. Failures stay in the record. One run is one observation, not typical latency.

Cost. Authoritative usage and billing evidence only. For the application-operated path, the harness records search calls and model input and output tokens, and those records get combined with the providers’ published rates on the run date. The managed path is the awkward one, because the API does not return usage. Tabstack request telemetry recorded one action for the fast-mode calls made while validating the harness, so the cost of a managed call is recorded action count multiplied by the published action rate, and the customer-visible receipt is a console balance change. Published pricing allows a Research call to spend a variable number of actions. One action on a validation call establishes nothing about what the next question will cost.

Terminal status. Every run gets exactly one: complete, partial, unanswered, HTTP failure, transport failure, task failure, or timeout. Provider failures, empty answers and malformed citations stay in the denominator.

The committed rubric and results template define the full field list.

Scoring happens before anyone knows which system wrote what

The harness copies each final answer under a random identifier and strips provider names from the wrapper metadata. Evaluators score answer quality before they see system identity, event logs, latency or cost. The engineer who built the systems will not be the sole scorer.

The blind is incomplete, and we are not going to pretend otherwise. The two systems can produce visibly different answer structures, and an evaluator can infer identity from formatting with every label removed. That limitation ships alongside the results.

Failures stay in the denominator

The protocol does not allow a cleaner result set to be manufactured after the run.

  • Do not stop because one system appears better.
  • Stop the whole study for security, runaway spend, a provider outage that breaks comparability, corrupted logging, or a protocol violation.
  • Exclude a run only when the harness fails before the provider receives the request.
  • Keep provider errors, timeouts, empty answers, and malformed citations as results.
  • Never rerun only the system that appears to lose.
  • If the scoring rules change after outputs are visible, create protocol v2 and rerun the full set.

Operational failure is part of the architecture being evaluated. A path that times out is telling you something about itself, and deleting the row throws that away.

The pilot tested the harness, not the systems

We ran one source-discovery question through both configured paths to check that the harness could send the frozen input, run each complete system, preserve answers and sources, record timing and usage fields, mark the rows pilot_only=true, and package answer text for blind review.

Both paths reached a terminal complete state. Their outputs, logs and operational records are committed under evals/runs/20260915T203139-pilot.

Nobody has scored those answers. They are not ranked, and the rows will not enter the full evaluation denominator. The only thing the pilot concludes is that the harness runs end to end with both frozen configurations.

What this first evaluation will not establish

The planned 12-question run is exploratory. One run per system per question cannot establish universal superiority, stable provider performance, statistical significance, typical latency or cost, reliability across workloads, or which individual component caused an observed difference.

It can show the outputs, costs, timing, failures and evaluation scores that two disclosed configurations produced on a frozen question set. That is enough to identify tradeoffs and design a stronger follow-up. It is not enough to declare a category winner.

Most teams should do something smaller than this

If you only need to know whether one workflow meets a product requirement, run an acceptance test instead. Define the required claims, the citation behaviour, the latency ceiling, the cost ceiling and the failure handling, then test that one workflow against them. It answers the question you actually have, and it takes a fraction of the setup.

If you need causality, run a component-level experiment. Hold the search provider, the fetched pages, the prompt and the model constant, then change one component at a time.

If you need estimates of variability or stable provider behaviour, you need repeated runs over a larger stratified set, with the sample, the analysis and the stopping rules registered before execution.

And when the deciding factor is ownership, privacy, deployment boundary or maintainability, a benchmark is the wrong instrument. Do the qualitative architecture review and skip the scores.

Data flow

Both paths send a public-web question and retrieved content through external services in the frozen configurations. The application-operated path sends data to its chosen search and model providers. The managed path sends the question and retrieved page content to Tabstack, and third-party models may process it under Tabstack’s contracts. Review every provider’s data terms before using confidential inputs. For Tabstack, start with the Privacy Notice.

Never trained on. Private by default.

Submit a workflow or evaluation question

The evaluation is stronger if its questions reflect jobs people are actually trying to ship. Propose a current public-web workflow, a question, and the answer elements that would make the result useful.

Open the evaluation directory and propose a question

Related reading and implementation:

  • Search gives your model sources. What turns them into a cited answer?
  • Build a cited live-web answer in Python with one Research call
  • Research guide
  • Research API reference
  • Public evaluation protocol

START FREE

Read the guide, then make the call.

Start with 10,000 free credits. No credit card required.

curl -fsSL https://tabstack.ai/install.sh | sh