Keep a Knowledge Base Current
Re-read a defined set of public pages with /extract/markdown, detect what actually changed with a content hash, and reprocess only those pages in your own pipeline.
A retrieval system is only as current as its last ingest. The usual failure is not that re-reading is hard, it is that re-embedding everything on every run is expensive, so it happens rarely, so the index goes stale.
This example re-reads a fixed source list, compares each page against the copy you stored last time, and hands back only the pages that changed. Embedding, chunking, and scheduling stay in your pipeline, where they belong.
import { createHash } from "node:crypto";import { readFileSync, writeFileSync, existsSync } from "node:fs";
import Tabstack from "@tabstack/sdk";
const client = new Tabstack();
const STATE = "kb-state.json";
const SOURCES = [ "https://docs.tabstack.ai/guides/research", "https://docs.tabstack.ai/guides/how-to-extract-json", "https://docs.tabstack.ai/pricing",];
function fingerprint(content: string): string { return createHash("sha256").update(content).digest("hex");}
function loadState(): Record<string, string> { return existsSync(STATE) ? JSON.parse(readFileSync(STATE, "utf8")) : {};}
/** Re-read each source and return only the ones whose content changed. */async function refresh(sources: string[]) { const state = loadState(); const changed: { url: string; content: string }[] = [];
for (const url of sources) { // nocache matters here: the shared content cache is keyed by URL, // effort and region, and a cached copy defeats the whole job. const result = await client.extract.markdown({ url, nocache: true }); const digest = fingerprint(result.content);
if (state[url] === digest) { console.log(`unchanged ${url}`); continue; }
console.log(`changed ${url}`); changed.push({ url, content: result.content }); state[url] = digest; }
writeFileSync(STATE, JSON.stringify(state, null, 2)); return changed;}
for (const page of await refresh(SOURCES)) { // Your pipeline owns what happens next: chunk, embed, upsert, delete. console.log(`-> reprocess ${page.url} (${page.content.length} chars)`);}import hashlibimport jsonimport pathlib
from tabstack import Tabstack
client = Tabstack()
STATE = pathlib.Path("kb-state.json")
SOURCES = [ "https://docs.tabstack.ai/guides/research", "https://docs.tabstack.ai/guides/how-to-extract-json", "https://docs.tabstack.ai/pricing",]
def fingerprint(content: str) -> str: return hashlib.sha256(content.encode()).hexdigest()
def load_state() -> dict: return json.loads(STATE.read_text()) if STATE.exists() else {}
def refresh(sources: list[str]) -> list[dict]: """Re-read each source and return only the ones whose content changed.""" state = load_state() changed = []
for url in sources: # nocache matters here: the shared content cache is keyed by URL, # effort and region, and a cached copy defeats the whole job. result = client.extract.markdown(url=url, nocache=True) digest = fingerprint(result.content)
if state.get(url) == digest: print(f"unchanged {url}") continue
print(f"changed {url}") changed.append({"url": url, "content": result.content}) state[url] = digest
STATE.write_text(json.dumps(state, indent=2)) return changed
for page in refresh(SOURCES): # Your pipeline owns what happens next: chunk, embed, upsert, delete. print(f"-> reprocess {page['url']} ({len(page['content'])} chars)")# A shell version of the same idea, one page at a timetabstack extract markdown https://docs.tabstack.ai/pricing --nocache > new.md
if ! diff -q old.md new.md > /dev/null 2>&1; then echo "changed, reprocessing" mv new.md old.md # your ingest step hereelse echo "unchanged" rm new.mdfiA run prints one line per source, then the pages your pipeline needs to reprocess:
unchanged https://docs.tabstack.ai/guides/researchchanged https://docs.tabstack.ai/guides/how-to-extract-jsonunchanged https://docs.tabstack.ai/pricing-> reprocess https://docs.tabstack.ai/guides/how-to-extract-json (18422 chars)How it works
Section titled “How it works”- One call per source, and nothing else.
/extract/markdownis 10 credits, deterministic, and returns prose rather than markup. Re-reading a 40-page source set is a predictable cost you can budget. nocache: trueis required, not optional. Page content is cached by URL, effort, and region rather than by account, so without it a refresh can hand you the copy you already have. See Data Handling.- Hash the extracted markdown, not the HTML. A page whose nav, ads, or build hash changed will differ byte-for-byte in HTML while its content is identical. Extraction strips that, so the hash only moves when the prose moves.
- Store the digest, not the document. The state file holds one hash per URL, so change detection costs nothing to keep around even for a large source set.
Scheduling it
Section titled “Scheduling it”There is no scheduler in Tabstack. Use whatever already runs in your stack:
# Every night at 03:000 3 * * * cd /srv/kb && python refresh.py >> refresh.log 2>&1A GitHub Actions cron, a Cloud Scheduler job, or a queue worker all work the same way. The only thing Tabstack provides is the current content of the pages you name.
Extending it
Section titled “Extending it”- Need typed fields rather than prose? Swap
/extract/markdownfor/extract/jsonwith a schema, and hash the serialized object instead. Useful when you are tracking specific values rather than indexing text. - Source list that changes? Keep it in the same state file, or generate it from your own database. Tabstack does not crawl, so discovering new URLs is your step.
- Pages that need a browser? Pass
effort: 'max'for JavaScript-heavy sources. See Effort levels. - Want to know what changed, not just that it changed? Keep the previous markdown alongside the hash and diff the two. Prose diffs are readable, which is part of why markdown is the right storage format here.
Installation
Section titled “Installation”npm install @tabstack/sdkpip install tabstackSet your API key before running:
export TABSTACK_API_KEY=your_api_key