NEVER TRAINED ON · PRIVATE BY DEFAULT · BUILT BY MOZILLA

GITHUB

Integrate with your model stack

Add current, cited answers to your model application without touching the model

A second path at the application layer sends current questions to a managed Research call and returns a cited result in your own response shape, with the model left exactly as it is.

Your application already runs a model. Someone asks it about something that happened last month, and it answers confidently from training data that ended long before. The fix people reach for first is a bigger model, or a search tool bolted onto the one they have. Both of those change the thing that currently works.

There’s a cheaper move. Add a second path at the application layer, send the current question to a managed Research call, and return the result through the response shape your application already uses. The model keeps its prompt, its settings and its context window. It just stops being the component responsible for knowing what happened last month.

Here’s what that looks like in Python, with the parts that bite.

Don’t let the model decide

The first design decision is routing, and I’d make it boring on purpose. The application takes an explicit mode:

model-web-app chat --prompt "Explain our local cache policy"
model-web-app research --question "What are the current ways to add web search to an Ollama-based application?"

Letting a model classify each request as chat or research is the version everyone wants to build, and it’s the version you can’t test deterministically. Every failure then has two possible causes, and the first thing you’ll do when the integration misbehaves is argue about whether the classifier was right. Get the integration working with an explicit mode, prove the model path is untouched, then add classification as its own change with its own tests.

flowchart LR
    U[Request] --> R{mode}
    R -->|chat| M[Existing model adapter]
    R -->|research| T[Tabstack /research]
    M --> A[ApplicationResponse]
    T --> A

The contract belongs to you

The rest of your application should never see an SDK event object. Give it something you own:

from dataclasses import dataclass, field
from typing import Literal

@dataclass(frozen=True)
class SourceRecord:
    id: str
    url: str
    title: str | None = None
    claims: list[str] = field(default_factory=list)
    source_queries: list[str] = field(default_factory=list)

@dataclass(frozen=True)
class ApplicationResponse:
    kind: Literal["chat", "research"]
    content_markdown: str
    sources: list[SourceRecord]
    status: Literal["complete", "partial"]
    provider: str
    retrieved_at_utc: str | None = None

Two fields in there are doing quiet work.

sources stays ordered. Tabstack documents citedPages as ordered by first citation in the report, and the report’s inline [1], [2] markers refer to that order. Sort it or deduplicate it in the adapter and you’ve broken the only claim-to-source mapping you get in fast mode.

status has a value other than complete so a partial result has somewhere to live. This integration never produces one, because a Research call ends in either complete or error. It’s there so that when you add your own retrieval later, or when a call gives you a report with half its sources unreachable, the answer to “what do we return?” isn’t “a confident-looking answer with a gap in it”.

The adapter

This is the whole of it:

import os
from typing import Literal

import httpx
import tabstack
from tabstack import Tabstack

def run_research(
    question: str,
    mode: Literal["fast", "balanced"] = "fast",
    nocache: bool = True,
) -> ApplicationResponse:
    if not os.environ.get("TABSTACK_API_KEY"):
        raise ConfigurationError("TABSTACK_API_KEY is not set")

    with Tabstack(max_retries=0) as client:
        try:
            with client.agent.research(query=question, mode=mode, nocache=nocache) as stream:
                for event in stream:
                    if event.event == "error":
                        raise ResearchTaskError(
                            message=getattr(event.data.error, "message", None) or "unknown error",
                            activity=event.data.activity,
                        )
                    if event.event == "complete":
                        return to_application_response(event.data)
                    report_progress(event)
        except tabstack.APIStatusError as err:
            raise ResearchHTTPError(status=err.status_code) from None
        except tabstack.APIConnectionError:
            raise ResearchTransportError("could not reach Tabstack") from None
        except httpx.TransportError as err:
            raise ResearchTransportError(f"stream interrupted: {type(err).__name__}") from None

    raise ResearchProtocolError("stream ended without complete or error")

Four lines there aren’t obvious.

The environment check happens before the client is built. Tabstack() raises a TabstackError from its constructor when neither api_key nor TABSTACK_API_KEY is set, and a missing environment variable deserves its own exit code and its own message rather than arriving as a generic SDK error.

except httpx.TransportError is there because the SDK’s coverage stops at the stream boundary. Anything that fails before the stream opens comes back as an SDK exception: AuthenticationError, RateLimitError and friends under APIStatusError, or APIConnectionError and its APITimeoutError subclass. Once you’re iterating, the stream iterator in 2.8.5 passes transport failures straight through, so a read timeout or a dropped connection mid-stream arrives as a raw httpx exception. Catch only the SDK types and a stream that dies halfway takes your process with it.

The error event’s payload carries message, name and an optional stack. Keep the message and the activity, drop the stack. Provider stack traces in your logs are noise at best, and a leak at worst. The defensive read on message is there because the guide warns the error object can arrive unpopulated despite being typed as required.

from None keeps the SDK’s exception chain out of your tracebacks, which is the cheapest way to stop a key or a header ending up in a log aggregator.

The final raise is the case people forget. If the stream ends without complete and without error, the loop just finishes and the function returns None unless you make that an error. A silent None becomes an empty answer three layers up, and nobody can tell it apart from a model that had nothing to say.

Timeouts, and the one everybody gets wrong

Research has no server-side timeout on the call as a whole. A fast query usually comes back in under a minute. A broad balanced query consults more sources and can run for around four minutes. So don’t wrap it in a total deadline. A fixed 90-second timeout will kill healthy long runs and you’ll conclude the API is flaky when the problem is your own stopwatch.

What you want is silence detection, and you already have some. The SDK’s default is httpx.Timeout(connect=5.0, read=60, write=60, pool=60), and httpx applies that read timeout to each read, so on a stream it fires when nothing has arrived for sixty seconds. That’s the behaviour you want, just with a number you chose deliberately rather than inherited. If you add your own silence detection on top, reset it on every event.

One parameter that is not a total deadline, despite reading like one: fetch_timeout caps how long Research spends fetching a single page.

Retries, which are probably already on

    with Tabstack(max_retries=0) as client:

That max_retries=0 is deliberate. The SDK retries twice by default, on connection errors and on 408, 409, 429 and 5xx responses. The committed Week 1 sample run recorded application_retries: 0 in its manifest alongside sdk_max_retries: 2, which is a neat illustration of how “we don’t retry” can be untrue in a codebase nobody has lied in.

A replayed Research call can repeat billed work, and there’s no idempotency key to make it safe. Turn retries off until somebody has decided which failures are safe to replay and what the ceiling is, then turn them back on with that decision written down.

What actually comes back

The Week 1 example repository has a committed run of this shape: fast mode, SDK 2.8.5, 15 September 2026. It completed in 20.3 seconds with the first event arriving after 1.06 seconds, analysed seven pages, cited seven, and returned a 1,649-character report. Ten events, each once, in order: start, planning, one iteration with a search inside it, writing, complete.

The interesting part is what the citation metadata did and didn’t contain.

Three of the seven cited pages were the same Ollama documentation page under different URLs: once over http, once over https, and once with a .md suffix. Your distinct source count is not your cited page count, and normalising scheme, repeated slashes, trailing slashes and file suffixes before counting is a ten-line function that stops you reporting seven sources when you have five.

Every claims array came back empty. The field is typed as always present and it was present, carrying nothing, because fast mode skips the analysis phases that populate it. The guide’s worked example shows it populated. That’s the gap between a documented shape and a live response, and it’s why the adapter maps claims into the contract without ever depending on its contents. In that mode the report’s inline [n] markers, read against the source order, are the claim-to-source link.

relevance, reliability and summary were all absent, as the SDK docstrings say they are in fast mode. title was present on every page. And complete.data.timestamp arrived as epoch milliseconds while the guide’s example shows an ISO string, so set retrieved_at_utc from your own clock and store the raw value alongside it.

None of that makes the output unusable. It does mean you build against the response you get rather than the example you read, and that you read citedPages and gapEvaluations defensively, since both are optional.

What to log

Log enough to find the boundary that failed, and nothing that belongs to the person who asked the question. The event names in order, with timestamps. Time to first event and to terminal state. The terminal status and which phase it failed in. Counts of pages analysed, cited pages raw, and distinct URLs after normalising. Whether the model adapter was called, which for this path should always be false. The SDK version, its max_retries and its effective timeout.

Keep out API keys, authorisation headers, cookies, whole environment dumps, and model reasoning.

When this is the wrong shape

If what you need is links, use a search API and skip all of this. If the web layer is what makes your product different from the next one, own the loop yourself and accept the boundaries that come with it. If you need the model to make decisions mid-research based on what it finds, you want a tool loop, not one call.

This shape fits when your application needs a finished, cited answer to a current question, and the model you run is good at everything except knowing what happened last month. You get to keep the model, and the parts you still own are the question, the review and the decision about whether the answer is good enough to act on.

A complete event means the job finished. That’s all it means.

Never trained on. Private by default. Built by Mozilla.

START FREE

Read the guide, then make the call.

Start with 10,000 free credits. No credit card required.

curl -fsSL https://tabstack.ai/install.sh | sh