Ground Your Own Model in Live Sources
Build a command-line tool that answers a current question with your own model, grounded in live sources through one /research call. Shows the difference between an ungrounded answer and a cited one.
Your model doesn’t come with the web. This tutorial closes that gap without making the model operate a search loop: one /research call does the web work, and your model does what you chose it for.
By the end you will have a small CLI that answers a current question two ways, so you can see the difference for yourself.
What you’ll build
Section titled “What you’ll build”A command-line tool that:
- Takes a question.
- Calls
/researchonce and gets back a cited answer from live sources. - Hands that answer, plus its sources, to a model you run.
- Prints a short brief with the citations preserved.
- On request, answers the same question with no research at all, for comparison.
What you’ll learn
Section titled “What you’ll learn”- How to consume the
/researchstream and read the report off thecompleteevent. - How to keep claim-level citations attached to the answer instead of losing them.
- How to write a prompt that keeps your model inside the evidence it was given.
- How to fail honestly when research returns no sources.
Prerequisites
Section titled “Prerequisites”- Python 3.9+ or Node.js 22+.
- A Tabstack API key.
- A model you run. This tutorial uses Ollama on
http://localhost:11434, because it needs no key and keeps everything local. Any chat model works; only theask_modelfunction changes.
Create and Setup Your API Key
Before you can start using Tabstack API, you’ll need to create an API key and set it up in your environment.
1. Create Your API Key
- Visit the Tabstack Console
- Sign in to your account (or create one if you haven’t already)
- Navigate to the API Keys section and click the “Manage API Keys”
- Once you are on the API Keys page, Click “Create New API Key”
- Give your key a descriptive name (e.g., “Development”, “Production”) and click the “Create API Key”
- Copy the generated API key and store it securely
2. Set Up Environment Variable
For security and convenience, we recommend storing your API key as an environment variable rather than hardcoding it in your scripts.
macOS/Linux
# Add to your shell profile (~/.bashrc, ~/.zshrc, or ~/.bash_profile)export TABSTACK_API_KEY="your_api_key_here"
# Or set it temporarily for the current sessionexport TABSTACK_API_KEY="your_api_key_here"
# Reload your shell or run:source ~/.bashrc # or ~/.zshrcWindows (Command Prompt)
# Set temporarily for current sessionset TABSTACK_API_KEY=your_api_key_here
# Set permanently (requires restart)setx TABSTACK_API_KEY "your_api_key_here"Windows (PowerShell)
# Set temporarily for current session$env:TABSTACK_API_KEY = "your_api_key_here"
# Set permanently for current user[Environment]::SetEnvironmentVariable("TABSTACK_API_KEY", "your_api_key_here", "User")3. Verify Your Setup
Test that your environment variable is set correctly:
macOS/Linux/Windows (Git Bash):
echo $TABSTACK_API_KEYWindows (Command Prompt):
echo %TABSTACK_API_KEY%Windows (PowerShell):
echo $env:TABSTACK_API_KEYYou should see your API key printed in the terminal.
Pull a model and confirm Ollama is up:
ollama pull llama3.2curl http://localhost:11434/api/chat -d '{ "model": "llama3.2", "messages": [{ "role": "user", "content": "reply with OK" }], "stream": false}'Install the SDK:
npm install @tabstack/sdkpip install tabstackThe Ollama calls below use the standard library (urllib) and built-in fetch, so there is nothing else to install.
Step 1: Get a cited answer
Section titled “Step 1: Get a cited answer”/research always streams. Progress events arrive while the work happens, and the complete event carries the report and the sources it cited.
import Tabstack from "@tabstack/sdk";
const client = new Tabstack();
type Cited = { url: string; title?: string | null; claims?: string[] };
/** Return the report and cited pages for a question. */async function research(question: string): Promise<{ report: string; cited: Cited[] }> { const stream = await client.agent.research({ query: question, mode: "fast" });
for await (const event of stream) { if (event.event === "iteration:start") { console.log(" researching..."); }
if (event.event === "error") { throw new Error(event.data.error?.message ?? "Research failed"); }
if (event.event === "complete") { return { report: event.data.report, cited: event.data.metadata.citedPages ?? [], }; } }
throw new Error("Stream ended before the complete event");}from tabstack import Tabstack
client = Tabstack()
def research(question: str): """Return (report, cited_pages) for a question.""" for event in client.agent.research(query=question, mode="fast"): if event.event == "iteration:start": print(" researching...", flush=True)
if event.event == "error": message = event.data.error.message if event.data.error else "Research failed" raise RuntimeError(message)
if event.event == "complete": return event.data.report, event.data.metadata.cited_pages or []
raise RuntimeError("Stream ended before the complete event")Note what is absent. No query planning, no result ranking, no page fetching, no HTML parsing, no gap check, no second search. That is the whole point: the loop ran inside the call.
Step 2: Keep the citations attached
Section titled “Step 2: Keep the citations attached”The citations are what make the answer checkable, so number the sources before the model sees them. Then a claim in the output can point at a specific URL.
/** Number the sources so the model can cite them as [1], [2], and so on. */function formatSources(cited: Cited[]): string { return cited .map((page, i) => `[${i + 1}] ${page.title ?? "(untitled)"} - ${page.url}`) .join("\n");}def format_sources(cited_pages) -> str: """Number the sources so the model can cite them as [1], [2], and so on.""" lines = [] for i, page in enumerate(cited_pages, start=1): lines.append(f"[{i}] {page.title or '(untitled)'} - {page.url}") return "\n".join(lines)Step 3: Hand it to your own model
Section titled “Step 3: Hand it to your own model”Your model’s job is now narrow: turn the research into a brief, and stay inside the evidence. The system prompt is doing the real work here.
const OLLAMA_URL = "http://localhost:11434/api/chat";const MODEL = "llama3.2";
const SYSTEM = [ "You write short briefs from research you are given.", "Use only the research provided. If the research does not cover part of the", "question, say so plainly instead of filling the gap.", "Cite sources inline as [1], [2] using the numbered list provided.",].join(" ");
type Message = { role: "system" | "user" | "assistant"; content: string };
async function askModel(messages: Message[]): Promise<string> { const response = await fetch(OLLAMA_URL, { method: "POST", headers: { "Content-Type": "application/json" }, // stream defaults to true, so turn it off for a single response. body: JSON.stringify({ model: MODEL, messages, stream: false }), });
if (!response.ok) { throw new Error(`Ollama returned ${response.status}`); }
const data = await response.json(); return data.message.content;}
async function groundedAnswer(question: string): Promise<string> { const { report, cited } = await research(question);
if (cited.length === 0) { // No sources is a result. Do not let the model paper over it. return "Unverified: research returned no cited sources for this question."; }
const prompt = [ `Question: ${question}`, "", `Research:\n${report}`, "", `Sources:\n${formatSources(cited)}`, ].join("\n");
return askModel([ { role: "system", content: SYSTEM }, { role: "user", content: prompt }, ]);}import jsonimport urllib.request
OLLAMA_URL = "http://localhost:11434/api/chat"MODEL = "llama3.2"
SYSTEM = ( "You write short briefs from research you are given. " "Use only the research provided. If the research does not cover part of the " "question, say so plainly instead of filling the gap. " "Cite sources inline as [1], [2] using the numbered list provided.")
def ask_model(messages: list[dict]) -> str: payload = json.dumps( # stream defaults to true, so turn it off for a single response. {"model": MODEL, "messages": messages, "stream": False} ).encode()
request = urllib.request.Request( OLLAMA_URL, data=payload, headers={"Content-Type": "application/json"} )
with urllib.request.urlopen(request) as response: return json.load(response)["message"]["content"]
def grounded_answer(question: str) -> str: report, cited_pages = research(question)
if not cited_pages: # No sources is a result. Do not let the model paper over it. return "Unverified: research returned no cited sources for this question."
prompt = ( f"Question: {question}\n\n" f"Research:\n{report}\n\n" f"Sources:\n{format_sources(cited_pages)}" )
return ask_model( [{"role": "system", "content": SYSTEM}, {"role": "user", "content": prompt}] )Two decisions in that code:
- An empty citation list short-circuits the model. If research found nothing, saying so beats asking a model to improvise. This is the difference between a grounded system and a confident one.
- The model never sees a URL it has to fetch. It receives finished text and a numbered source list. Nothing in this design depends on the model calling tools correctly.
Step 4: See the difference
Section titled “Step 4: See the difference”Add the ungrounded path so you can compare answers to the same question.
/** The same question, with no research. For comparison only. */async function ungroundedAnswer(question: string): Promise<string> { return askModel([{ role: "user", content: question }]);}def ungrounded_answer(question: str) -> str: """The same question, with no research. For comparison only.""" return ask_model([{"role": "user", "content": question}])Ask something that changed after your model was trained: a current version number, a recent price, this month’s release notes. The ungrounded answer will usually be fluent, specific, and out of date, with nothing in the output telling you which.
Complete code
Section titled “Complete code”import Tabstack from "@tabstack/sdk";
const client = new Tabstack();
const OLLAMA_URL = "http://localhost:11434/api/chat";const MODEL = "llama3.2";
const SYSTEM = [ "You write short briefs from research you are given.", "Use only the research provided. If the research does not cover part of the", "question, say so plainly instead of filling the gap.", "Cite sources inline as [1], [2] using the numbered list provided.",].join(" ");
type Cited = { url: string; title?: string | null; claims?: string[] };type Message = { role: "system" | "user" | "assistant"; content: string };
async function research(question: string): Promise<{ report: string; cited: Cited[] }> { const stream = await client.agent.research({ query: question, mode: "fast" });
for await (const event of stream) { if (event.event === "iteration:start") { console.log(" researching..."); }
if (event.event === "error") { throw new Error(event.data.error?.message ?? "Research failed"); }
if (event.event === "complete") { return { report: event.data.report, cited: event.data.metadata.citedPages ?? [] }; } }
throw new Error("Stream ended before the complete event");}
function formatSources(cited: Cited[]): string { return cited .map((page, i) => `[${i + 1}] ${page.title ?? "(untitled)"} - ${page.url}`) .join("\n");}
async function askModel(messages: Message[]): Promise<string> { const response = await fetch(OLLAMA_URL, { method: "POST", headers: { "Content-Type": "application/json" }, body: JSON.stringify({ model: MODEL, messages, stream: false }), });
if (!response.ok) { throw new Error(`Ollama returned ${response.status}`); }
const data = await response.json(); return data.message.content;}
async function groundedAnswer(question: string): Promise<string> { const { report, cited } = await research(question);
if (cited.length === 0) { return "Unverified: research returned no cited sources for this question."; }
const prompt = [ `Question: ${question}`, "", `Research:\n${report}`, "", `Sources:\n${formatSources(cited)}`, ].join("\n");
const answer = await askModel([ { role: "system", content: SYSTEM }, { role: "user", content: prompt }, ]);
return `${answer}\n\nSources:\n${formatSources(cited)}`;}
const args = process.argv.slice(2);
if (args.length === 0) { console.log('Usage: npx tsx ground.ts "your question" [--no-research]'); process.exit(1);}
const noResearch = args.includes("--no-research");const question = args.filter((a) => a !== "--no-research").join(" ");
if (noResearch) { console.log("\n--- model alone ---"); console.log(await askModel([{ role: "user", content: question }]));} else { console.log("\n--- grounded in live sources ---"); console.log(await groundedAnswer(question));}import jsonimport sysimport urllib.request
from tabstack import Tabstack
client = Tabstack()
OLLAMA_URL = "http://localhost:11434/api/chat"MODEL = "llama3.2"
SYSTEM = ( "You write short briefs from research you are given. " "Use only the research provided. If the research does not cover part of the " "question, say so plainly instead of filling the gap. " "Cite sources inline as [1], [2] using the numbered list provided.")
def research(question: str): for event in client.agent.research(query=question, mode="fast"): if event.event == "iteration:start": print(" researching...", flush=True)
if event.event == "error": message = event.data.error.message if event.data.error else "Research failed" raise RuntimeError(message)
if event.event == "complete": return event.data.report, event.data.metadata.cited_pages or []
raise RuntimeError("Stream ended before the complete event")
def format_sources(cited_pages) -> str: return "\n".join( f"[{i}] {page.title or '(untitled)'} - {page.url}" for i, page in enumerate(cited_pages, start=1) )
def ask_model(messages: list[dict]) -> str: payload = json.dumps({"model": MODEL, "messages": messages, "stream": False}).encode() request = urllib.request.Request( OLLAMA_URL, data=payload, headers={"Content-Type": "application/json"} )
with urllib.request.urlopen(request) as response: return json.load(response)["message"]["content"]
def grounded_answer(question: str) -> str: report, cited_pages = research(question)
if not cited_pages: return "Unverified: research returned no cited sources for this question."
prompt = ( f"Question: {question}\n\n" f"Research:\n{report}\n\n" f"Sources:\n{format_sources(cited_pages)}" )
answer = ask_model( [{"role": "system", "content": SYSTEM}, {"role": "user", "content": prompt}] )
return f"{answer}\n\nSources:\n{format_sources(cited_pages)}"
def main() -> None: args = sys.argv[1:]
if not args: print('Usage: python ground.py "your question" [--no-research]') raise SystemExit(1)
no_research = "--no-research" in args question = " ".join(a for a in args if a != "--no-research")
if no_research: print("\n--- model alone ---") print(ask_model([{"role": "user", "content": question}])) else: print("\n--- grounded in live sources ---") print(grounded_answer(question))
if __name__ == "__main__": main()Run it
Section titled “Run it”npx tsx ground.ts "What is the current long-term support release of Node.js?"npx tsx ground.ts "What is the current long-term support release of Node.js?" --no-researchpython ground.py "What is the current long-term support release of Node.js?"python ground.py "What is the current long-term support release of Node.js?" --no-researchGrounded, the output ends with sources you can open:
--- grounded in live sources --- researching...Node.js 24 is the current long-term support release [1]. It entered LTS inOctober 2025, replacing Node.js 22 as the recommended line for production [1].
Sources:[1] Node.js Releases - https://nodejs.org/en/blog/releaseUngrounded, it ends with whatever the model remembers:
--- model alone ---The current long-term support release of Node.js is version 20, which enteredLTS in October 2023 and is recommended for production use.Run both against something that changed recently. The second answer is fluent, specific, and has nothing in it that tells you it is out of date.
What to try next
Section titled “What to try next”- Swap the model. Point
ask_modelat any chat API. Nothing else in the design changes, because the model was never doing the web work. - Add a known URL. When you already have the page,
/extract/markdownis 10 credits and deterministic, against a/researchcall that runs several actions. Reach for the cheap path when you can. - Ask for typed output instead of prose.
/extract/jsonreturns JSON matching a schema you supply. Start with Schema design. - Skip the hand-rolled tools. If your model runs inside a framework, the maintained packages expose the same calls as native tools. See Hermes, LangChain (Python), or the full list.
- Go deeper on the endpoint. Autonomous Research covers modes, every progress event, citation handling, and client-side timeouts.
- Check a claim instead of answering a question. Verify a Model Answer Against Live Sources uses the same call for a narrower job.