---
title: Keep a Knowledge Base Current | Tabstack
description: Re-read a defined set of public pages with /extract/markdown, detect what actually changed with a content hash, and reprocess only those pages in your own pipeline.
---

A retrieval system is only as current as its last ingest. The usual failure is not that re-reading is hard, it is that re-embedding everything on every run is expensive, so it happens rarely, so the index goes stale.

This example re-reads a fixed source list, compares each page against the copy you stored last time, and hands back only the pages that changed. Embedding, chunking, and scheduling stay in your pipeline, where they belong.

Tabstack does not schedule this, diff it, or embed anything. It returns the current content of a page you name. The change detection and the cron entry below are ordinary code in your own stack.

- [TypeScript](#tab-panel-15)
- [Python](#tab-panel-16)
- [CLI](#tab-panel-17)

```
import { createHash } from "node:crypto";
import { readFileSync, writeFileSync, existsSync } from "node:fs";


import Tabstack from "@tabstack/sdk";


const client = new Tabstack();


const STATE = "kb-state.json";


const SOURCES = [
  "https://docs.tabstack.ai/guides/research",
  "https://docs.tabstack.ai/guides/how-to-extract-json",
  "https://docs.tabstack.ai/pricing",
];


function fingerprint(content: string): string {
  return createHash("sha256").update(content).digest("hex");
}


function loadState(): Record<string, string> {
  return existsSync(STATE) ? JSON.parse(readFileSync(STATE, "utf8")) : {};
}


/** Re-read each source and return only the ones whose content changed. */
async function refresh(sources: string[]) {
  const state = loadState();
  const changed: { url: string; content: string }[] = [];


  for (const url of sources) {
    // nocache matters here: the shared content cache is keyed by URL,
    // effort and region, and a cached copy defeats the whole job.
    const result = await client.extract.markdown({ url, nocache: true });
    const digest = fingerprint(result.content);


    if (state[url] === digest) {
      console.log(`unchanged  ${url}`);
      continue;
    }


    console.log(`changed    ${url}`);
    changed.push({ url, content: result.content });
    state[url] = digest;
  }


  writeFileSync(STATE, JSON.stringify(state, null, 2));
  return changed;
}


for (const page of await refresh(SOURCES)) {
  // Your pipeline owns what happens next: chunk, embed, upsert, delete.
  console.log(`-> reprocess ${page.url} (${page.content.length} chars)`);
}
```

```
import hashlib
import json
import pathlib


from tabstack import Tabstack


client = Tabstack()


STATE = pathlib.Path("kb-state.json")


SOURCES = [
    "https://docs.tabstack.ai/guides/research",
    "https://docs.tabstack.ai/guides/how-to-extract-json",
    "https://docs.tabstack.ai/pricing",
]




def fingerprint(content: str) -> str:
    return hashlib.sha256(content.encode()).hexdigest()




def load_state() -> dict:
    return json.loads(STATE.read_text()) if STATE.exists() else {}




def refresh(sources: list[str]) -> list[dict]:
    """Re-read each source and return only the ones whose content changed."""
    state = load_state()
    changed = []


    for url in sources:
        # nocache matters here: the shared content cache is keyed by URL,
        # effort and region, and a cached copy defeats the whole job.
        result = client.extract.markdown(url=url, nocache=True)
        digest = fingerprint(result.content)


        if state.get(url) == digest:
            print(f"unchanged  {url}")
            continue


        print(f"changed    {url}")
        changed.append({"url": url, "content": result.content})
        state[url] = digest


    STATE.write_text(json.dumps(state, indent=2))
    return changed




for page in refresh(SOURCES):
    # Your pipeline owns what happens next: chunk, embed, upsert, delete.
    print(f"-> reprocess {page['url']} ({len(page['content'])} chars)")
```

Terminal window

```
# A shell version of the same idea, one page at a time
tabstack extract markdown https://docs.tabstack.ai/pricing --nocache > new.md


if ! diff -q old.md new.md > /dev/null 2>&1; then
  echo "changed, reprocessing"
  mv new.md old.md
  # your ingest step here
else
  echo "unchanged"
  rm new.md
fi
```

A run prints one line per source, then the pages your pipeline needs to reprocess:

```
unchanged  https://docs.tabstack.ai/guides/research
changed    https://docs.tabstack.ai/guides/how-to-extract-json
unchanged  https://docs.tabstack.ai/pricing
-> reprocess https://docs.tabstack.ai/guides/how-to-extract-json (18422 chars)
```

## How it works

- **One call per source, and nothing else.** `/extract/markdown` is 10 credits, deterministic, and returns prose rather than markup. Re-reading a 40-page source set is a predictable cost you can budget.
- **`nocache: true` is required, not optional.** Page content is cached by URL, effort, and region rather than by account, so without it a refresh can hand you the copy you already have. See [Data Handling](/trust/data-handling#caching/index.md).
- **Hash the extracted markdown, not the HTML.** A page whose nav, ads, or build hash changed will differ byte-for-byte in HTML while its content is identical. Extraction strips that, so the hash only moves when the prose moves.
- **Store the digest, not the document.** The state file holds one hash per URL, so change detection costs nothing to keep around even for a large source set.

## Scheduling it

There is no scheduler in Tabstack. Use whatever already runs in your stack:

Terminal window

```
# Every night at 03:00
0 3 * * * cd /srv/kb && python refresh.py >> refresh.log 2>&1
```

A GitHub Actions cron, a Cloud Scheduler job, or a queue worker all work the same way. The only thing Tabstack provides is the current content of the pages you name.

## Extending it

- **Need typed fields rather than prose?** Swap `/extract/markdown` for [`/extract/json`](/guides/how-to-extract-json/index.md) with a schema, and hash the serialized object instead. Useful when you are tracking specific values rather than indexing text.
- **Source list that changes?** Keep it in the same state file, or generate it from your own database. Tabstack does not crawl, so discovering new URLs is your step.
- **Pages that need a browser?** Pass `effort: 'max'` for JavaScript-heavy sources. See [Effort levels](/guides/effort-levels/index.md).
- **Want to know what changed, not just that it changed?** Keep the previous markdown alongside the hash and diff the two. Prose diffs are readable, which is part of why markdown is the right storage format here.

## Installation

- [TypeScript](#tab-panel-18)
- [Python](#tab-panel-19)

Terminal window

```
npm install @tabstack/sdk
```

Terminal window

```
pip install tabstack
```

Set your API key before running:

Terminal window

```
export TABSTACK_API_KEY=your_api_key
```
