Clone the example, set one environment variable, and run one command to send a current question to Tabstack Research. The call streams progress, then writes a Markdown report and a machine-readable list of cited pages. The committed sample shows the full path completed in a recorded Python environment, with every artifact included so you can reproduce it.
Who this is for: People running their own model inside a product, assistant, or internal workflow that needs a current, cited answer from the public web.
Run the example
Start with the complete implementation in the cited-research example repo rather than assembling a partial snippet:
git clone https://github.com/Mozilla-Ocho/tabstack-cited-research-python.git
cd tabstack-cited-research-python
uv sync --frozen
export TABSTACK_API_KEY=...
uv run cited-research \
--query "What are the current ways to add web search to an Ollama-based application, and what output does each approach return?" \
--mode fast \
--nocache \
--output artifacts/sample-run
The key comes from the environment. The CLI has no –api-key flag and does not write the key to its output files.
The repository pins tabstack==2.8.5 in uv.lock. The sample was run with Python 3.12.13 and uv 0.11.6 on macOS Darwin 25.6.0, arm64. The project declares Python 3.9 as its floor, but this sample was not executed on Python 3.9.
What the program does
flowchart TD
Q[Question] --> CLI[Python CLI]
CLI -->|one /research request over Server-Sent Events| R[Tabstack Research]
R -->|start, planning, searching, writing, complete| OUT[report.md + sources.json + run-manifest.json]
OUT --> APP[Your model, product interface, or internal workflow]
Your application sends the question to /research. The endpoint returns a Server-Sent Events stream. The program displays selected progress events, stops on completion, and saves the report and cited-page metadata as separate files.
The application remains responsible for reviewing the result and deciding how to use it. A completed Research call gives you a report and a source record to review; how you use the answer stays in your application.
The Python call
The repository contains the tested CLI, sanitization, file persistence, and exit-code handling. This standalone version shows the same core API path and can be run as a script after installing tabstack==2.8.5 and setting TABSTACK_API_KEY:
import json
import os
from pathlib import Path
import tabstack
from tabstack import Tabstack
QUERY = (
"What are the current ways to add web search to an Ollama-based "
"application, and what output does each approach return?"
)
OUTPUT = Path("artifacts/my-run")
PROGRESS_EVENTS = {
"start",
"planning:start",
"planning:end",
"iteration:start",
"iteration:end",
"searching:start",
"searching:end",
"writing:start",
"writing:end",
}
if not os.environ.get("TABSTACK_API_KEY"):
raise SystemExit("TABSTACK_API_KEY is not set")
OUTPUT.mkdir(parents=True, exist_ok=True)
final = None
with Tabstack() as client:
stream = client.agent.research(
query=QUERY,
mode="fast",
nocache=True,
)
for event in stream:
if event.event in PROGRESS_EVENTS:
print(event.event, event.data.message)
if event.event == "error":
raise RuntimeError(event.data.error.message)
if event.event == "complete":
final = event
break
if final is None:
raise RuntimeError("Research stream ended without a complete event")
(OUTPUT / "report.md").write_text(
final.data.report.rstrip() + "\n",
encoding="utf-8",
)
sources = []
for source in final.data.metadata.cited_pages or []:
sources.append(
{
"id": source.id,
"url": source.url,
"title": source.title,
"claims": source.claims,
"source_queries": source.source_queries,
"relevance": source.relevance,
"reliability": source.reliability,
}
)
(OUTPUT / "sources.json").write_text(
json.dumps(sources, indent=2, ensure_ascii=False) + "\n",
encoding="utf-8",
)
print(f"report -> {OUTPUT / 'report.md'}")
print(f"sources ({len(sources)}) -> {OUTPUT / 'sources.json'}")
print(f"tabstack SDK -> {tabstack.__version__}")
Python exposes SDK fields in snake case, including cited_pages, source_queries, and total_pages_analyzed. The JSON wire format uses camel case.
How the stream terminates
/research always returns a Server-Sent Events stream. The sample run received these events, in this order, once each:
start
planning:start
planning:end
iteration:start
searching:start
searching:end
iteration:end
writing:start
writing:end
complete
complete is the successful terminal event. Research does not send a later done event. A task-level failure arrives as an error event inside the stream. That is why the loop must handle both error and complete. HTTP rejections such as 401, 400, or 429 happen before the stream opens and surface as SDK exceptions instead.
Complete input
The publication run used this question:
What are the current ways to add web search to an Ollama-based application, and what output does each approach return?
It used mode=“fast” and nocache=True. The input, exact command, sanitized events, output, and run manifest are committed in artifacts/sample-run.
Complete output
The CLI writes a durable output bundle:
artifacts/sample-run/
├── question.txt
├── command.txt
├── events.sanitized.jsonl
├── report.md
├── sources.json
├── run-manifest.json
├── stdout.txt
├── stderr.txt
└── terminal.png
The complete returned report is in report.md. The ordered citation metadata is in sources.json. They remain separate so an application can render the answer while retaining a machine-readable source record.
The program printed this progress before writing the files:
start Starting research
planning:start Planning research strategy
planning:end Planning complete: 6 queries
iteration:start iteration 1/1 Starting iteration 1 of 1
searching:start iteration 1 Searching with 6 queries
searching:end iteration 1 8 new urls Found 8 URLs
iteration:end iteration 1 Iteration 1 complete (fast mode)
writing:start Writing report
writing:end Report draft complete
complete report -> artifacts/sample-run/report.md
sources (7) -> artifacts/sample-run/sources.json
The run in numbers
The run started at 2026-09-15 19:00:04.556 UTC and completed 20.3 seconds later, with the first progress event arriving after 1.06 seconds. The stream ended on complete and the process exited 0. Tabstack analyzed seven pages, cited seven, and returned a 1,649-character report with inline numbered citations.
Request telemetry recorded one Research action for the whole call; at the published fast-mode rate that is 250 credits for a planned, multi-source, cited answer. The API response itself carries no usage field, so treat that as the implied cost.
One run is one data point. Use it to see the shape of a fast-mode result, not to size latency or cost for your workload.
What fast mode returns, and how to read it
Fast mode is built for speed: one iteration, no judge pass, sources returned as they were found. That shows up in the output in four specific ways, and each has a straightforward handling pattern.
Cited pages are counted by URL, not by canonical page
Three of the seven entries pointed at the same Ollama documentation page under different URLs (http versus https, and a .md suffix). Normalize scheme and suffix before counting distinct sources, or display the report’s inline citation numbers, which already do that work for the reader.
Claim lists arrive empty in fast mode
Each cited page carries a claims array, and in this mode the array is []. The report’s inline [n] markers are the claim-to-source link: they map in order to cited_pages. The example keeps the claims field in sources.json because it is part of the SDK contract and deeper modes can populate it.
Optional metadata stays optional
Title was present on every source. Relevance, reliability, and summary were not; those come from the analysis phases that fast mode skips. Read them defensively and your code works unchanged across modes.
Source selection is broad by design
Two of the seven sources were tutorials about building your own page-reading tool, and the report used them to describe a fourth implementation path. Fast mode returns what the search surfaced; whether that path belongs in the answer is a judgment your application or reviewer makes. We have not scored this answer, so the article presents it as returned output.
Failure handling
The repository gives configuration, HTTP, and task failures different exit codes:
| Path | Observed or tested behavior |
|---|---|
| TABSTACK_API_KEY missing | Exit 5, message on stderr, no request made |
| Invalid key | HTTP 401 before the stream opened, exit 3 |
| Streamed error event | Synthetic fixture produced exit 2 and printed the returned message and phase |
| Stream closed without complete | Synthetic fixture produced exit 2 |
Only the invalid-key case sent a failing request to production. The streamed failure cases use synthetic fixtures rather than trying to break production.
The full implementation also sanitizes its public event log. It retains event names, timestamps, iteration counters, short status messages, and completion counts. It excludes API keys, authorization headers, cookies, full source text, stack traces, environment dumps, and model reasoning.
Timeouts, retries, and cleanup
The SDK applies a 600-second timeout to /research streams. The example adds no shorter total timeout, so it does not turn a longer healthy run into an application timeout.
The SDK retries certain transport failures twice by default, including 408, 409, 429, and server errors. The example adds zero application-level retries. Repeating a request can repeat work and billing, so retry policy should be an explicit product decision.
The Tabstack client runs inside a context manager. The response closes when the block exits, including when the stream returns an error.
A clean uv sync –frozen passed 17 tests. Ruff and Pyright also passed with no reported errors. The quality gate ran on Python 3.12.13, not every declared Python version.
When this approach is sufficient
Use one managed Research call when your application starts with a question and needs a cited answer from live sources without operating the search, retrieval, and synthesis loop itself.
Use a search interface instead when links and excerpts are the desired output, or when your own model already operates the remaining loop in a way that meets your requirements.
Build the loop in your application when source selection, retrieval policy, model behavior, or citation assembly is core to your product and you need control over each component.
Use known-page extraction when you already have the URL and need page text or fields in a defined shape. Do not use a research workflow for work that a simpler output contract already solves.
Data flow
A Research call sends your question and the pages it retrieves through Tabstack’s pipeline. The models and infrastructure behind that pipeline run under contracts we negotiated with zero data retention, terms you would otherwise need an enterprise agreement to get. Nothing you send is used to train a model. On our side, request data is kept for 90 days to support your account, then deleted. The Trust page spells out what each endpoint processes and where.
Never trained on. Private by default.
Run the Python example
The repository contains the lockfile, implementation, fixtures, failure evidence, sample output, and evaluation harness:
Run the cited Research example on GitHub
Then continue with:



