KB Retention
Every project knowledge base, and every chat that has had files uploaded to it, owns a vector table in the database. Over time most of them stop being used, but their tables stay behind and make the database larger and slower to maintain.
KB Retention archives the vector table of a knowledge base that has been inactive for a configurable period. Archiving removes only the search index. The documents themselves, their text, metadata, folders, tags and uploaded files are kept, so the knowledge base can be rebuilt at any time with Restore KB. See Archived knowledge bases and Restore.
Open it from Super Admin → Platform Settings → KB Retention. It requires the Platform permission.
What is and is not removed
| Kept | Removed |
|---|---|
| Projects, chats and messages | The project's or chat's vector table |
| Every document row, its text, metadata, tags and folder | The indexes on that vector table |
| Uploaded files in storage | |
| Connector links (Confluence, Google Drive, SharePoint) and crawled URLs | |
| Saved queries |
Before a table is dropped, an immutable manifest records the table, its indexes, how many chunks it held, every document it contained and how each one will be rebuilt, when the knowledge base was last active and the settings the run used. Manifests cannot be edited or deleted. If the table changes after its manifest was written, the next run writes a new manifest; the old one stays as history.
Settings
These are changed on the KB Retention tab and take effect on the next run:
| Setting | Default | What it does |
|---|---|---|
| Mode | Report only | Off: nothing runs. Report only: the daily run produces a dry-run report and changes nothing. Canary: only projects or teams on the canary list are archived. Enforce: every eligible knowledge base is archived. |
| Project KB retention (days) | 365 | Days of inactivity before a project knowledge base is archived. 0 turns archiving off for project knowledge bases. Range 0-3650. |
| Chat KB retention (days) | 90 | Days of inactivity before a chat's uploaded-file knowledge base is archived. 0 turns archiving off for chats. Range 0-3650. |
Every change is recorded in the settings audit log and the Super Admin audit log.
Deployment settings
These are set per environment with environment variables, through the Helm values kbRetention.env, which reach the API and every worker that runs retention. The tab shows the values in effect, read-only.
| Environment variable | Default | What it does |
|---|---|---|
KB_RETENTION_BATCH_SIZE | 100 | Maximum knowledge bases archived per run (1-5000). |
KB_RETENTION_DROP_DELAY_MS | 1000 | Pause between two table drops within a run, in milliseconds (0-60000). |
KB_RETENTION_CANARY_PROJECT_IDS, KB_RETENTION_CANARY_TEAM_IDS | empty | Projects, or every project of a team, that may be archived while the mode is Canary. Comma-separated. |
KB_RETENTION_EXEMPT_PROJECT_IDS, KB_RETENTION_EXEMPT_TEAM_IDS | empty | Projects, or every project of a team, that are never archived. Their chats are exempt too. Comma-separated. |
kbRetention:
env:
KB_RETENTION_CANARY_TEAM_IDS: "team-a,team-b"
KB_RETENTION_BATCH_SIZE: "10"
An invalid number falls back to its default and is logged.
What counts as activity
A knowledge base is archived only when all of its activity is older than the retention period:
- Project knowledge base: the project's last activity (chat or agent usage), the last KB sync, the newest document added, the newest chat in the project, the last time any agent or search read its vector table, and the last restore.
- Chat knowledge base: the chat's last message, the newest file uploaded to it, the last time its vector table was read, and the last restore.
Reads of a vector table are tracked even when they come from another project, for example an agent configured to search a different project's knowledge base.
What is never archived
A knowledge base is skipped, and the report says why, when:
- documents are still queued, syncing or being deleted;
- a crawl is running, or a scheduled KB sync or recurring crawl is configured for the project;
- a full KB sync is in progress;
- the project or its team is exempt;
- any document in the vector table cannot be rebuilt: it has no retained text and is not a plain uploaded file with its original still in storage (
source_unavailable); - the owning project or chat cannot be identified exactly (
orphan,ambiguous_owner,unresolved_name), or something else in the database depends on the table (has_dependencies). These tables are listed in reports but never dropped.
Eligibility, and whether every document can still be rebuilt, is checked again under a lock immediately before each table is dropped. If anything changed since the manifest was written (changed_since_manifest) or a source has gone (source_unavailable), the knowledge base is skipped for that run.
Emergency pause
Pause cleanup stops archiving immediately: scheduled runs do nothing, Run cleanup now is refused, and a run in progress stops before its next drop. Restores are not paused, so users can still get their knowledge bases back. Resume from the same toggle.
Operators can also set the KB_RETENTION_HARD_PAUSE=true environment variable on the backend to pause cleanup without the UI.
Dry-run preview
Run dry-run inventories every vector table and shows, without changing anything:
- how many knowledge bases would be archived and the storage they use;
- why each of the others would be skipped;
- for each knowledge base: owner, last activity, table size and whether its sources can be rebuilt.
The daily run produces the same report automatically while the mode is Report only.
Jobs and history
- Cleanup jobs lists every report and cleanup run with its trigger, mode, status, knowledge bases archived, skipped and failed, storage reclaimed and duration. Open a job to see each knowledge base it evaluated.
- Run cleanup now starts a cleanup immediately using the current settings. It is only available in Canary or Enforce mode and while cleanup is not paused.
- Restore history lists every restore with the requester, progress, and the result or error.
Archive a specific knowledge base now
Archive a knowledge base now archives one project or chat knowledge base immediately, ignoring its inactivity and the canary list. Every other rule still applies, including rebuildability, exemptions and in-flight work, and the action is audited. Use it to validate restores during a canary rollout.
Rollout runbook
- Report only in production. Leave the mode at Report only for at least one daily run, or click Run dry-run.
- Review the report. Check the total reclaimable storage, and look through the skipped reasons for anything unexpected.
source_unavailableand ownership reasons are expected to be non-zero. - Canary. Set
KB_RETENTION_CANARY_TEAM_IDSorKB_RETENTION_CANARY_PROJECT_IDSto one or two low-risk teams or projects and a smallKB_RETENTION_BATCH_SIZE(for example 10) in the environment's Helm values, deploy, then switch the mode to Canary. Use Archive a knowledge base now on a test project and restore it to confirm the round trip. - Measured batches. Switch to Enforce with a modest batch size and delay, watch the cleanup jobs and restore history for failures, and raise the batch size gradually.
- Steady state. Keep Enforce with the default batch size. Use Pause cleanup at any sign of trouble; restores keep working while paused.
The daily run starts at 02:30 UTC.