OPERATIONS · v0.4.2.14

Migrating an existing installation

Source: modules/anthill/docs/MIGRATION.md, a path from the root of the product's repository — versioned with the code, rendered here at every release.

Status: inventory, store lease, backup and restore, and the store split. This document describes anthill --inventory, the first slice of W3-09 (migrating existing ANTHILL / FORAGER / MICROMOUND installations into FORMICARIA); the store lease, the second: one store, one writer; anthill --backup / anthill --restore, the third: backups that have been rehearsed before anyone relies on them; and anthill --migrate-stores (v0.4.2.0, roadmap E-3), which moves each module's tables out of anthill.db into its own file. W3-09's first rule is inventory before touching anything. The split is the only command here that moves data, and it backs up and rehearses first; the tenancy retrofit (every row's organization) is not built.

anthill --inventory <installation-path> [--forager <data-dir>] [--state <micromound-state-dir>] [--out <file.json>]

It prints a human-readable summary, and with --out writes a JSON record. Both say the same things, including what could not be determined.


The constraint this whole design exists for

Reading an ANTHILL installation through the product's own types is a write.

Constructing SqliteMemory runs InitDb() — SchemaStatements, EnsureColumns, SchemaMigrationRunner.Apply, ApplyLearningReset — and AnthillRuntime.Initialize() writes .anthill/config.json when that file is absent. Both are correct for a colony that is about to run. Both are wrong for a tool whose entire job is to describe an installation as it currently is: a tool built on them would have migrated the store and created the config file before it printed a line, and would then report the config file as present, having made it so. "Let me look first" would be the thing that changed the installation.

So the inventory opens the installation as foreign files:

  • the databases through ForeignStore.Open — Microsoft.Data.Sqlite, read-only, unpooled, and immutable=1 whenever there is no live -wal/-shm pair to read through (below);
  • everything else through File and Directory, reading sizes, counts and presence.

Read-only is not enough, and record v1 got this wrong. A plain read-only SQLite open of a WAL store that was stopped cleanly — so its -wal and -shm were removed — creates both files and leaves them. Record v1 opened stores exactly that way, so "touches nothing" was false by two files on every cleanly stopped colony; its tests could not see it because their fixture store was not in WAL mode. ForeignStore reads by what the files mean:

beside the store how it is read what the record says
no -wal immutable=1 — no lock, no file created the main file is the whole store, all of it read
-wal and -shm read-only, through them something has it open (or had it open when it stopped); the committed log is included
-wal, no -shm immutable=1 the log was not read — reading it means creating its index, a write
-journal immutable=1 the rollback journal was not applied — undoing it is a write

database_read_mode carries that sentence for each store, so a reader knows what the counts cover.

It never constructs SqliteMemory, AnthillRuntime, SqliteMoundStore or InfrastructureRepository. InstallationInventorySourceTests is a source guard that fails the build if any of the four reappears in the reader, the record types or the command — the convenience is one careless edit away, and reintroducing it would be invisible, because the record would still look right and the installation would already be different.

The verb is answered in Program.cs above AnthillRuntime.Initialize(), beside --version and --help, so running it cannot leave a workspace wherever the operator happened to type it.

Two consequences worth stating, because they look like bugs and are not:

  • A store whose -wal and -shm were left by a process that died can need recovery, and then refuses to open read-only. Recovering it would be a write. The refusal is recorded as database_open_failure and is the correct outcome.
  • --out refuses to write inside the installation it just read, including inside a FORAGER data directory or a MICROMOUND state directory passed on the same command line. Writing the record into the thing you promised not to touch is the exact bug the rest of this design avoids, and --out ./inventory.json typed from the colony's own directory is the reflex that produces it.

The record

JSON, with a record_version of its own — independent of any product version, because what changes shape across releases is the record, and a reader needs to know which shape they are holding. Top level:

key what it is
record_version this record's shape. "2" today: v2 added store_lease and database_read_mode (below), and v1's FORAGER note compared against a migration count that had gone stale.
taken_at_utc, taken_by when, and by what.
constraints what this record deliberately cannot prove, carried inside it — a limit stated in a tool's help and not in its output is a limit the output does not have.
anthill the colony: config resolution, workspace paths, the store.
forager the knowledge engine, if a data directory was given or found.
micromound the device state directory, if one was given.
not_determined every section's unknowns, rolled up.

Absence is always a stated value and never an omitted key. Every section carries its own not_determined; every fact that could not be read carries a note saying why. A blank where a fact should be is the failure mode this record exists to prevent: a plan made against a field that is empty because nothing looked is a plan made against nothing. unreadable is likewise never folded into absent — a permission problem reported as a missing directory invites a migration that recreates what is already there.


ANTHILL

Config resolution

The rule the runtime would have used: ANTHILL_CONFIG_FILE when set, else <root>/.anthill/config.json, with a relative value resolving against ANTHILL_HOME when that is set and against the working directory otherwise. Both variables are reported as this process sees them; the path is resolved against the root you named.

config_file_exists is reported, never created. When it is false, the workspace paths below it are the defaults this installation would take — not values somebody wrote down, and the record says so.

legacy_config_* names data/anthill.json. The runtime no longer reads it and only warns about it. It is in this record because an inventory that omitted it mis-describes the installation's actual configuration: an operator reading that file believes the settings in it are in force, and there is a real field failure behind that — a live qualification run spent a whole mission with acting_coder_enabled true in data/anthill.json and false in the runtime, with nothing reading the file and nothing erroring.

Workspace paths

workspace_root, db_path, backup_dir, logs_dir, exports_dir, agent_workspace_dir — as the config states them (or as defaulted), each resolved, each with kind (file / directory / absent / unreadable) and a size where there is one.

The database

anthill.db with its size, plus the presence and size of -wal and -shm. Those two are part of the database: a copy of the .db alone is the store as it stood before the last committed writes.

Read from inside it, directly:

  • anthill_meta — it already carries schema_version, anthill_version, config_profile, config_path, workspace_root and db_path. Values are reported exactly as stored, which is JSON-encoded.
  • anthill_schema_ledger — the migrations that ran. Each row was written in the same transaction as the change it names, so this ledger cannot record a change that did not happen.
  • the table list from sqlite_master, and a row count per table.

store_lease — the lease's state, looked at and never taken (see The store lease, below): absent, free, held or unknown, the holder the store recorded, and the holder the lease file names. It is read before the store is, so the record says whether something was holding the installation while it was being inventoried.

module_stores (v0.4.2.0) — each module's own file beside anthill.db (infrastructure.db, micromound.db), read the same way: whether it exists with its -wal and -shm, its lease looked at, its tables and row counts, and where the module's tables are — own, legacy (still inside anthill.db: anthill --migrate-stores has not run) or interrupted (in both files: a split that stopped, which the colony will not start on). The manifest comparison and the controller identity look across all of them.

schema_migrations is NOT evidence

The 22-entry schema_migrations table is reported, and labelled evidence: "none".

Its own header says why: the runner that writes it inserts a row per entry and executes no DDL, and every description in it says something was "verified" when nothing verified anything. The real schema work was always SchemaStatements and EnsureColumns, unconditionally, on every start.

An inventory that read this table to decide what an installation has would be wrong twenty-two times. So it is reported — a reader will find it and needs to know what it is — and nothing in this record is derived from it. What an installation actually has is anthill_schema_ledger plus the tables and columns the record lists.

The retrofit numbers

W1-02 §4.5 phase 5 names these as the check, so they are taken before anything moves.

  • null_org_id — per table, for the anchor and the scoped tables. Must reach zero. A row with no org_id is a row no organization owns.
  • null_project_id — for missions, conversations, objectives and device_placements (where a device belongs; until v0.4.2.0 that was two columns of micromound_mounds).

The second number is read the same way and means the opposite thing, which is why each count carries its meaning in the same object as the count. project_id IS NULL is unfiled (W1-02 §4.2): a defined placement, org-scoped and reachable by the organization's owners. It is not a defect and it is not a backlog.

This number is to be compared across a migration, never minimised. It must come out exactly as it went in. A count that fell means something invented a project for a row that had none — and which project a two-year-old mission belonged to was never written down, so no migration can recover it.

A table that is not in the store, or a column that this store predates, reports null with a note — unread, never zero.

Manifest drift, both directions

The store's tables compared against TenancyTables.All (plus the module tables), because that is the one existing, maintained inventory of what is in anthill.db. A second list written beside it would drift from it the first time a table was added, and then the drift report would itself be drifted.

  • in_store_not_classified — a table nobody has decided the tenancy of. A migration would carry it blind.
  • classified_not_in_store — ordinary for a table created lazily on first use, and reported so a reader can tell that case from a table that was expected and is not here.

excluded_name_patterns names what the comparison deliberately does not see (sqlite_%, missions_fts_%), rather than dropping them quietly.

Identity, by presence and never by value

  • whether micromound_controller_identity holds a row — asked with COUNT(*); the seed column is never selected;
  • whether .anthill/field.key exists (and its mode), and whether ANTHILL_ENCRYPTION_KEY is set — the variable's value is never read into the record;
  • provider_credentials and colony_credentials: counted, never read, nothing decrypted.

verified_to_open is the constant false, and that is the honest field. The record can say "a controller identity is present and I cannot verify that it opens" without opening the fleet.

This matters more here than anywhere else in the tree. The controller signing identity is the one identity that cannot be re-minted: ADR-009 allows no rotation, only physical re-enrolment of every mound. Its seal lives outside the database — the field key file, or the environment variable of the process that runs the colony. So "copy the database" is not "copy the installation", and a migration that moved anthill.db and left the key behind would produce a colony holding a fleet identity nobody can use.

The id namespaces, because three disagree

Each id is reported with the namespace it lives in and who minted it:

field namespace minted by
installations.id ins_ locally, by migration 0001_tenancy_spine, from RandomNumberGenerator
installations.fingerprint inf_ the same
organizations.id org_ the same — and deliberately unnamed, because a name a migration chose would be an invented value
FORAGER's organization org_local a constant, stamped by FORAGER migration 014_tenancy on every store

A copy of a store that has not yet run 0001 mints different values for the first three, while every FORAGER store shares the fourth. A record saying "this installation's organization is X" without saying which kind of X that is, is a claim it cannot support.


FORAGER

Given with --forager <data-dir>. Failing that, two places FORAGER itself documents are looked at: FORAGER_DATA_DIR in this process's environment, and <root>/data (that variable's default, relative to the process's own directory). The section's note says which of those it was, because a directory the command guessed and one an operator named are different kinds of fact.

A packaged install keeps its data in a per-user location this command cannot guess, so no FORAGER section means not looked at — never "this installation has no FORAGER". The --out refusal covers a discovered directory as well as a named one: a refusal computed from the arguments alone would let the record be written into a store the command had just read.

  • forager.db with its sidecars.
  • Its schema_migrations ledger — that one is evidence, unlike ANTHILL's frozen table: its runner executes each change and writes the row in the same transaction. The FORAGER in this tree ships 22 migrations (InstallationInventory.ForagerMigrationsShipped), so fewer means not all have run and more means a newer FORAGER wrote this store. Record v1 said 15, a day after 016 and 017 landed, and so would have called every current store "newer"; the number is now pinned by a test that counts FORAGER's own migration files, which caught it again when E-18 added 018 to 021. E-16 added 022 (the organization's cache scope) and moved the number with it.
  • instance.marker read as a file and forager_instance read as a row, reported side by side and never reconciled. Deciding whether they agree is what the engine's own loadInstance does, and it does it by minting a new generation and rewriting the marker — asking the product this question changes the thing you came to record. A human reads both: a marker naming another generation means the database file was swapped under it; a row with no marker means the database arrived in a directory it has not run in.
  • operator.key — presence and mode, never contents. uploads/, exports/, tmp/ — file counts and sizes.

The store lock: participating / empty / unprovable

  • participating — store_lock exists, so any holder is an engine that writes a row in it and the lock means what it says.
  • empty — no tables at all; nothing has ever written this store, so nothing can be holding it.
  • unprovable — tables, but no store_lock: a store written by an engine from before migration 012, which can only be held by an engine that never writes a row there.

The third word is kept. An empty lock table proves the table is empty and proves nothing about the store. Rounding unprovable to "free" is what lets two engines write one store. When the database was not opened at all the state is unknown, which is also not "free".

FORAGER's own classifier creates the lock table when it is missing — right for an engine about to serve the store, wrong for a command that promised to touch nothing. The classification is therefore re-read here rather than called.


MICROMOUND

Given with --state <dir>.

  • identity/seed — presence and mode, never contents.
  • state/ and evidence/ — file counts and total sizes.
  • /etc/micromound/manifest.json and /etc/micromound/micromound.env, the locations deploy/ installs to. Absent there means "not in the standard place", never "this mound has no manifest", and the record says so in not_determined.

If the env file is found, the record names the hazard: it carries MICROMOUND_ARGS, and --enroll-token in it is an enrolment token in plaintext. It must not travel in a migration bundle, a backup that moves, or a support archive. The token itself is never read and is not in the record.


The store lease

One store, one writer. W3-09's third rule is the one that silently corrupts: never let two processes write the same database, and never let a version-skewed process migrate a store a live one is serving (F74). The lease is how ANTHILL keeps it, and it is two things — the owner's decision on F20's open conflict, 20 September 2026:

  • The lock: an OS lock on <db>.owner, the file beside the database. flock(LOCK_EX|LOCK_NB) on Linux and macOS (W1-11 §5.3's executed mechanism), a byte-range lock (FileStream.Lock, which is LockFile) on Windows (F20's value). It is the enforcement. The OS releases it when the process dies, so a crashed holder's lease is free at once — there is no heartbeat, no staleness rule, and no comparison of two machines' clocks.
  • The record: a store_lease row in the store, and the same facts in the lease file. Pid, hostname, mode, purpose, version, when acquired, when released. Diagnosis only — a row is never evidence of holding, and one left by a process that crashed says who held the store last. On Linux a held lease file cannot be read by another .NET process at all, so the row is how a refused process names its holder.

Order is the point. The lock is taken before the store is opened, so a refused process has written nothing to the database. The holder then writes its row and only then constructs the store, whose schema pass and migration runner therefore always run under the lease.

Who takes it. The API host, for its life, and every CLI verb that opens the colony's store — including the ones that look like reads (--status, --config, --routes, --live-qualification), because opening the store runs its schema pass; there is no read through it. While a colony is running those verbs are refused, name the holder, and exit 11 — FORAGER's code for the same refusal. Lock-out recovery is therefore stop, --set-password, start. A second --api on the same store refuses to start the same way; under the LXC unit's Restart=on-failure it retries every five seconds and starts once the store is free. --inventory never takes the lease; it looks at it.

What it cannot do.

  • An advisory lock binds the processes that ask. Every ANTHILL before v0.4.0 takes no lease, and neither does a customer's sqlite3 shell. absent (no lease file) and free are therefore never proof that a store is idle, and the record keeps them apart for that reason.
  • Never delete <db>.owner. On Linux, deleting it while a holder runs lets a second process lock a new file of the same name. Nothing in ANTHILL deletes it; release stamps its facts and unlocks.
  • The file is owner-only (0600) on POSIX, because anyone who can open it can hold a shared lock on it and keep the colony from ever taking its lease.
  • A store named through a symbolic link to its file resolves to the file; a hard link cannot be seen from its path, and two names for one inode through hard links are not recognised as one store.
  • If the filesystem cannot lock (some network mounts), the lease cannot be taken and the process refuses rather than open the store unguarded.

Backup and restore

W3-09's second rule — consistent backups after quiescing writers, validated, rehearsed — is anthill --backup and anthill --restore (v0.4.1.2). Both are offline (owner's decision, 20 September 2026): they take the store's lease, so while the colony runs they are refused, name the holder and exit 11 like every other verb that opens the store. Stop, back up, start. A running colony backs itself up with Back up now instead (below); a restore is offline only.

Every store of the colony (v0.4.2.0). The colony is anthill.db and, once split, a file per module beside it. A backup takes every one that exists — each file's lease held, and all of them in SQLite's EXCLUSIVE mode together while they are copied, so the files in a backup are one moment of the colony — and its manifest (formicaria-backup-v2) lists them under stores, anthill.db first. The manifest's top-level fields still describe anthill.db and are what its copy is checked against; a manifest whose two records of it disagree was edited, and is refused. A v0.4.1.2 backup (formicaria-backup-v1, anthill.db alone) is still verified and restored.

A backup is a new directory under backup_dir (or --out), formicaria-backup-<time>-<id>, holding anthill.db, each module store that exists, and manifest.json, taken in this order:

  1. Quiesced. The lease keeps out every lease-aware ANTHILL. A writer from before v0.4.0 takes no lease, so the copy is taken through a connection in SQLite's EXCLUSIVE locking mode. A WAL store refuses that while any other process has the file open — idle, reading or writing — and the verb exits 12 naming what to look for (an older ANTHILL, a sqlite3 shell, a database browser). The lock is held until the copy is taken.
  2. Copied through SQLite's online backup API, which reads through the WAL, then switched to rollback-journal mode: one self-contained file.
  3. Checked — PRAGMA integrity_check, a row count for every table, a SHA-256, and a one-way check value of the field key.
  4. Rehearsed — restored into a scratch directory and checked again.

The manifest records all four and says "usable": false (exit 1) if any failed; such a directory is kept for inspection and a restore refuses it.

The field key is not in a backup unless --include-key is passed. It seals device seeds and the MICROMOUND controller identity (ADR-009: no rotation, only re-enrolment), so a directory holding the store and its key is a copy of everything. The check value is what lets a restore tell whether the workspace's key is the one the backup was sealed with. A key from ANTHILL_ENCRYPTION_KEY is never in a backup at all. A backup never creates a key — the cipher's own constructor mints one when none exists, which is why the backup path reads the key without constructing it.

A restore (anthill --restore <dir>) runs every check that reads only the backup and the key file before it takes the lease — taking the lease writes its holder into the live store, and a refused restore must leave the store byte-identical. It refuses:

  • a directory with no manifest, a manifest of another format, or one marked not usable;
  • a store file whose SHA-256 differs from the manifest (changed or damaged since it was taken);
  • a store that fails integrity_check, or whose counts differ from the manifest;
  • a backup sealed under a different key than this workspace's — every sealed value would be unreadable after it;
  • a backup sealed under a key when this workspace has none, unless the backup carries it;
  • a store that something else has open (exit 12).

Then, under every file's lease, it keeps each store it replaces as <file>.pre-restore-<time>, stages each file of the backup as <file>.restoring-<time>, checks the staged copies, moves them into place, checks them there, and puts the backup's key in place if it had to. A module file the backup does not hold is set aside — copied to <file>.pre-restore-<time> and removed — because the backup was taken while that module's tables were still inside anthill.db, and leaving today's file beside the restored store would put the module's tables in two files at once. anthill --migrate-stores splits it again. Restoring onto a machine with no colony is the same command; the empty store a lease's record created there is not kept. A backup from an older schema is restored as it is and migrated when the colony next opens it. --rehearse restores into a scratch directory only, takes no lease, and can run beside a live colony.

While the colony runs: Back up now (roadmap E-6). Settings → Diagnostics → Back up now, or POST /maintenance/backup (the manage_settings permission), takes the same backup from the running colony, into the same folder, in the same format; anthill --restore restores it like any other. What differs is how the copy is held still:

  • Every store's write lock, together. The colony takes BEGIN IMMEDIATE on a connection of the backup's own for each store — anthill.db first — and holds them all until the last store is copied, so the copies are one moment of the colony, as the offline backup's are. Each store is copied by the backup API through a second, reading connection; in WAL mode readers are never blocked, so the console and every read go on answering.
  • Writes wait, and not for long. A write arriving during the copies waits (every connection of the colony has a 30-second busy timeout) and goes through when they are done. The owner set the limit (24 September 2026): ten seconds. The copies go a megabyte at a time with the clock read between, so a colony too big to copy in ten seconds is found out as it is copied: the backup stops, lets every writer through, deletes what it had written and says to stop the colony and run anthill --backup, which has no limit. Nothing that is not a complete backup is ever left in the folder.
  • The field key only on an explicit opt-in (owner, 24 September 2026). Back up now leaves it out; Back up with key is its own action behind its own warning, and the endpoint takes the key only for {"include_key": true} — not "true", not 1. The manifest says so, as it does for --include-key.
  • The checks, the rehearsal and the manifest are the offline backup's, run after the locks are released. The manifest's quiesce says online: and how long the writes waited, out of how long they were allowed.

One runs at a time; a second is refused while the first is running. Each is recorded in the colony's event log (maintenance_backup, or maintenance_backup_refused with the reason). The console lists the full backups in the folder — how many, how much disk, and the newest, from their manifests. Delete all backups does not touch them: it acts on the mission path's snapshots. Compact & prune deletes the ones past the retention window, as every sweep does (below).

The mission path. Before a mission (at most once per backup_min_interval_minutes) the colony still writes backups/anthill_<time>.db. Until v0.4.1.2 that was a File.Copy of the main file of a live WAL database, which lacks every write since the last checkpoint. It now goes through the backup API and is deleted if it fails PRAGMA quick_check.

Backups age out (v0.4.2.5, roadmap E-14). Every backup in backup_dir — the full backups (formicaria-backup-*) and the mission path's snapshots (anthill_*.db) — is deleted backup_retention_days after it was taken: 30 by default, and held to 7–90 (decision F66; W1-10 §4.2). Until v0.4.2.5 nothing aged out: the snapshots were pruned by count (max_db_backups, which still applies) and a full backup was kept until somebody deleted it.

  • Dated by what the backup says about itself. A full backup by its manifest's created_at, or the time in its directory's name when the manifest cannot be read; a snapshot by the time in its name, or its last write when the name carries none. A backup exactly as old as the window is kept, and one that cannot be dated is kept and reported.
  • Nothing else in the folder. The patch tool's .bak copies have a retention of their own (W1-10's A4), an operator's own files are theirs, and a link is not followed. A backup written with anthill --backup --out <dir> is in a folder the operator chose and is never swept.
  • When. A minute after the colony starts and every six hours after that, and also after a mission's backup step, Back up now (only when the new backup is usable), Compact & prune and anthill --backup. Changing the window in Settings → Diagnostics → Storage (or /settings) takes effect at the next sweep.
  • Said on the event log. A sweep by the running colony that deleted, could not delete or could not date something records maintenance_backups_expired — the names, how many, the bytes freed, the cutoff — and says so when it has left the folder with no full backup at all. Back up now answers with the names of what it expired and Compact & prune with how many; anthill --backup, which runs with the colony stopped, prints them.

Why it is a promise and not a preference: an erasure cannot reach into a copy, so for any backup a deletion is complete only when that backup has expired (W1-10 §5.7). To keep a backup longer, copy it out of the backup folder. Not built: §5.7's tombstone log, which a restore replays so that a restored backup does not bring erased data back; what may be promised about backup expiry is a question for counsel (L-06).

The colony's own records age out (roadmap E-18; W1-10 §4.2, decision F59). What the colony said to itself while working — agent messages, artifacts, task results, events and source records (W1-10's A3) — is deleted event_retention_days after its mission ended: 180 by default. 0 keeps it all, 1 to 29 are read as 30, and a number above 3650 (ten years) is read as 0 — the colony never deletes sooner than it was asked to. Approvals and escalation decisions (A7) and message metrics and task-result summaries (A8) are deleted two years after they were written, whatever the setting says.

  • Counted from when the mission ended, which is the later of the mission's last save and its evaluation. A row whose mission has been cleared ages by its own time.
  • The colony's own log is not swept. The system_api mission's events are the record of what the operators and the API did — the operator shell's command lines (written before each command runs), memory wipes, resets, deleted backups, entitlement and directory decisions, and these sweeps' own events. They are not agent traffic, and no setting here deletes them, in any store: their retention belongs to the audit log that is to hold them (W1-05 §8.2; decision F64 keeps the shell command line with a 90-day floor), which does not exist yet. Until then they are kept, as they were before E-18. (Since E-12 a new shell command's line is written to the audit trail and not to this log; the rows written before stay here until E-12's slice 7 hands them over. See The audit trail, below.)
  • Never touched: a mission that is running or has not started, or one that can still run again — a pending request other than a patch's (approving it replays the refused step from the mission's own records) or an API job that has not finished. A pending approval is never deleted, whatever its age.
  • A mission a stopped colony left running is ended at the next start (roadmap E-18's remainder). Until this release a mission whose process died — the host killed, the machine rebooted — kept its row running for good: the next start re-queued its job as a new mission and abandoned its task attempts, but nothing closed the mission, so the sweep never touched its rows or the conversation that started it, the conversation read "running mission …" for ever, and Clear missions and Wipe memory refused while it stood. Now the colony, when it starts, ends every mission row a process that is gone left running (or created with work begun on it) as a failed mission: outcome failed_permanent, stop reason process_died, the reason on the row naming its job and what that start did with it — retried as the next attempt (the retry runs as a new mission, and the ended one is told which: mission_retried_as), orphaned for operator review, cancelled, or nothing to retry. Its events get mission_failed and mission_outcome, so the console's notifications and failures panel see it. It is counted as ended at its last sign of life — the latest moment the store saw the process working on it: its last save, event, task or task attempt, or its job's last heartbeat (renewed every 30 seconds while a job runs) — and its records age from then, as every other mission's age from its end and not from when the colony noticed; the reason on the row says the moment. A mission left running in this process's own lifetime is never touched. Tasks keep the status the dead process left them in; the attempt ledger says how far each got. What the operator sees after the upgrade's first start: one mission_failed notification per mission an earlier process ever left running, however old; each keeps its place in the mission history (its saved_at is its last sign of life, not the start), and its created_at still says when it began. The colony's own recovery record (mission_attempts, under startup-reconcile) is now written once per start that recovered something, numbered on; until this release it was written once per store, ever, so a colony that had recovered a crashed job before kept only that first record.
  • Kept until a patch is decided: a mission whose patch is proposed, or approved and not applied, keeps its task results (applying the patch reads the soldier's verdict) and its approvals; the rest of its traffic ages as usual.
  • Also kept whatever their age: the two finalization-ledger events that stop learning being counted twice; an artifact a live mission reads, read by id, or cites, and what those cite; the source and recall records of every mission a live mission recalled, as far as grading its citations reads them; and the escalation decision a conversation still acts on — past two years that one keeps the decision and loses its reason and its author's name. Missions, objectives and schedules are not this sweep's (W1-10 A5 and A6); conversations are, only when an auto-purge is chosen (below).
  • When. Two minutes after the colony starts and every six hours after that, and on Compact & prune, which compacts the store afterwards. Until the store is compacted a deleted row's bytes can stay in its free pages. The window is edited in Settings → Diagnostics → Storage (or /settings) and takes effect at the next sweep.
  • In steps. A sweep deletes at most 5,000 rows of a store per transaction and lets the colony write between them, so even the first sweep after the upgrade holds a mission's writes for one step, not for the whole sweep. A sweep that deleted anything then empties the write-ahead log (anthill.db-wal), briefly, and leaves the rest to the next checkpoint if a reader is in the way.
  • Said on the event log. A sweep that deleted or scrubbed something records maintenance_records_expired — how many rows per store, the two cutoffs, how many missions it left alone and why — and never a row's content. The first sweep a colony has is always recorded, and says it is the first.

Upgrading moves the default. Until E-18 event_retention_days applied to events alone, only from Compact & prune, and defaulted to 0 — keep everything; the settings save wrote that 0 into every config.json it touched, and no console page showed the key. So configuration schema 4 reads a 0 (or an absent key) in a file written before it as that old default and moves it to 180, and keeps any other number, which somebody chose. The first sweep after the upgrade then deletes, at once, everything already older than 180 days after its mission ended, and its event says what it removed. To keep everything through the upgrade, set event_retention_days to 3650 before upgrading — a 0 written by an older version reads as the old default — and to 0 afterwards; at schema 4 a 0 is kept as the operator's.

Conversations, customer source and device evidence (roadmap E-18, second build; W1-10 §4.2 and §4.4, decisions F59 and F60). The same sweep, on the same timer and on Compact & prune:

  • Conversations (A1, A2) are kept for the life of the conversation by default. conversation_retention_days chooses an auto-purge — 90, 180 or 365 days after the conversation's last turn, or 0, never (the default); a number between two choices reads as the longer one, and one above 365 as never. A purged conversation goes with its turns and attachments, pinned or not. Its missions stay (they are A5, "life of the project"), and so do its escalation decisions, which go two years after they were made (A7). Never purged, whatever its age: a conversation whose mission can still run or has a patch nobody has decided, one waiting on an operator decision, and one a schedule run is still running in.
  • Customer source in a patch (A4). A patch that was applied, rejected, failed, superseded or reverted loses its old and new file contents patch_retention_days after that decision — 90 by default, held to 30–365, with no "never"; 0 or less is read as the default, 90 — together with the pre-apply .bak copies the patch tool took for it and the repository index of the workspace it came from. The patch's row stays: status, file, hashes, who decided. Only copies the colony recorded taking are deleted, and only from the backup folder; nothing is looked for. A patch that can still be applied (proposed, approved) keeps everything, and so does a decided one while another patch of its set is undecided or its mission has not ended. Once the backup is gone, reverting a modify or delete is refused and the patch stays applied, and an alternative can no longer be built from it. The Patch Center says when the content was removed.
  • Device evidence and action records (section 4.4) are deleted device_evidence_retention_months after the colony received them — 24 by default; 0, or more than 120, keeps them — and never fewer than 12: that floor (decision F60) is a constant in the code, not a setting, so 1 to 11 are read as 12 and no plan or setting lowers it. An action whose mission has not reported is kept, with the evidence it cites. The device module sweeps three minutes after it starts and every six hours, and writes micromound_evidence_expired to the colony's event log (with the counts, the window and the floor) when it deleted something. POST /micromound/purge (a retired device's erasure) is unchanged.

Upgrading. A patch decided before this release has no recorded decision time; the first sweep dates it from what the store holds — applied_at for an applied or reverted patch, the rejecting approval for a rejected or superseded one, otherwise the sweep itself — so the first sweep after the upgrade removes at once the content and recorded backups of every patch decided more than 90 days before. To keep them through the upgrade, set patch_retention_days to 365 first. Device records carry the time the colony received them from this release on; a record from before is dated by the upgrade, so none of them goes for two years. Conversations are untouched unless an auto-purge is chosen.

FORAGER's windows, from the console (roadmap E-18's remainder; W1-10 §4.1, decision F59). The FORAGER engine the colony runs keeps what it derives for as long as six windows say. Until this release they could be set only in FORAGER's own environment (FORAGER_RETENTION_*), and F8's 0, which W1-10 wants to be one click, was an environment key. They are now settings, edited in Settings → Diagnostics → FORAGER's records (or /settings):

Key What goes, and from when FORAGER's default, range
forager_cache_retention_days a parser or extractor cache row, after it was last used (F5, F6) 90 days, 0–365; 0 turns both caches off
forager_model_log_retention_days a logged model call, with up to 1,000 characters of the model's answer, after the call (F8) 7 days, 0–30; 0 keeps none of the model's output, a choice of its own in the console
forager_review_event_retention_years a review event, after it was recorded (F14) 2 years, 1–7
forager_job_retention_days a finished job and its stages, after the job finished (F15); a project's latest job stays and loses only its text 90 days, 30–365
forager_stale_retention_days a statement, relationship or entity no live source supports any more, after it lost that support (F9–F11) 180 days; 0 at the next sweep, -1 never
forager_export_retention_days a built export's zip, after it was built (F16); the export's record stays a year 7 days, 1–90
  • Unset by default, and unset changes nothing. Each key is null until somebody sets it. The colony then hands the engine nothing for that window, so FORAGER's own default — the decided value — applies, or a FORAGER_RETENTION_* variable in the colony's environment if it has one, exactly as before this release.
  • Set, the colony's setting wins. A set window is held to FORAGER's range (a number past either end reads as that end) and handed to the engine as FORAGER's own variable, over whatever the colony's environment carries — the way the supervisor's other variables, such as FORAGER_DATA_DIR and FORAGER_PORT, already are. The stale window's -1 (any negative number) is handed over as the word never. Clearing the setting — null, or Not set in the console — gives the window back to the environment and FORAGER's default.
  • When a change takes effect. The engine reads its environment when it starts, and a save does not restart a running engine: a change applies at the engine's next start — the colony's next start, the next time Knowledge is turned on (a restart after a crash picks it up too), or Restart FORAGER now. The engine sweeps when it starts and every six hours after. For each window the console says what the engine running now has, and whether a change is waiting for its next start; /maintenance/stats carries the same as forager_retention — per window what the next start reads (next), what clearing the setting would leave (unset) and what the running engine started with (running), each with where it comes from (colony, environment or forager).
  • Restart FORAGER now. While a saved window differs from what the engine running now has, FORAGER's records offers the restart (POST /maintenance/restart-forager, an administrator's). It asks FORAGER first, with the host credential, what it is working on, and refuses (409 forager_busy) while any project the credential reaches has a processing job queued or running or an export being built, leaving the engine running as it was; otherwise the engine is drained, stopped and started again, and maintenance_forager_restart records who and which windows. It refuses as well when the colony runs no engine, when Knowledge is off, when nothing is waiting, when FORAGER cannot say what it is working on, and when the subscription would not let the engine start again.
  • A stop drains the engine. Before this, every stop of the engine the colony runs — Knowledge turned off, the colony shutting down — raced its own drain: the supervisor killed the process while it was still asking it to shut down. The stop now waits for the engine to drain and exit (ten seconds, then a kill, as documented), and a start asked for meanwhile waits for that exit instead of putting a second engine on the store.
  • A value outside FORAGER's range is refused by the console. The server still holds one it is sent (or finds in config.json) to the nearest end of the range; the console refuses it by name before it is sent, because the bottom of each range is the end that deletes the most. The stale window's negative is never, as before.
  • Only the engine the colony runs. A FORAGER you run yourself (attached mode) reads only its own environment, and these settings do not reach it.
  • Upgrading changes nothing. The keys are absent from an older config.json, read as unset, and written as null by the next settings save.

What backup and restore do not do. Restore into a running colony; back up FORAGER's store, which has its own; copy anything off the machine; see an idle reader of a store that is not in WAL mode (the EXCLUSIVE probe cannot, and every ANTHILL store is WAL). The mission path's backup is anthill.db alone: missions write nothing to a module's store.


The audit trail

Roadmap E-12, slice 1: W1-05's audit design as the E-12 slice plan builds it (audit-slice-plan.md, in the Formicaria repository's transition workspace, under security), and decisions F35 and F64.

What the colony keeps. An audit record is written into an outbox in anthill.db (audit_outbox and its one-row head audit_outbox_head), in the same transaction as the change it records, or, for an effect outside the store such as a shell command, committed before the effect runs. The outbox's triggers refuse every change but the drain stamp, and every deletion but retention's (F35). A single sequencer in the running colony drains it into the audit store beside anthill.db: audit.db, its index, and the records themselves in audit/<stream key>/<n>.db. Each record is filed in a stream of its scope and retention class — platform:<installation>/<class> for the installation's own records, org:<organization>@<installation>/<class> for an organization's — and chained there: a dense sequence and a SHA-256 over each record and the one before it. The store's triggers refuse any update or deletion, and no table in it has a foreign key. A record names a person by an opaque usr_ id, never by name: every account has one (users.user_id), and it never changes. Never edit, vacuum or delete the audit store by hand; Compact & prune never compacts it, and the storage footprint lists it as audit.

Upgrading is additive, at the first start of the new build. The ledger gains 0005_audit_outbox (the empty outbox, its head at a new generation, and its triggers) and 0006_user_ids (a usr_ id for every account), and audit.db and audit/ are created empty. Nothing is back-filled: a chain covers records as they are written. An older build refuses the upgraded store, naming both migrations; the way back is the backup taken before the upgrade, and this build runs both migrations again on it. An account an older build creates after a downgrade gets its id at the next start.

Two things an operator will notice:

  • Coordinators no longer see shell command lines. The operator shell records each command twice: operator_shell.attempted before it runs — the command line, cut at 4,096 characters with command_truncated saying so, whether it ran in the default directory (never the path), and the operator's usr_ id — filed in platform:<installation>/operator; and operator_shell.completed when it ends — exit code, whether it timed out, elapsed milliseconds, and the attempt's event id as causation_id — filed in platform:<installation>/config. A timeout completes unknown; a command that did not start completes failed with error_code: shell_start_failed, and the exception's text is recorded nowhere (the operator who ran it is still told why). A command whose attempt cannot be recorded is not run (500, shell_audit_unrecorded). The operator_shell_* rows in the colony's log stay, for the live console, and carry the audit record's event id instead of the command, the directory or the error. The rows an older build wrote keep theirs in the store, and every events path — /events/json, /events, the live stream and its replay, Colony Live, the reports, the memory vault and the SDK's event log — reads them without command, dir and error, and reads operator_shell_error's message as "Operator shell command failed to start." The one place a command line is read is GET /audit/platform, which only administrators can call (below). Decision F64 keeps the command line on these conditions: local only, installation operators only, 4,096 characters, one field wide.
  • anthill --restore drains the audit outbox first, and refuses when it cannot. Before it replaces any file, the restore runs the sequencer to the end of the outbox in the store it is about to replace, so nothing waiting there is lost with the replaced file. When the audit store cannot take the rows — it cannot be opened or written, with rows waiting — the restore refuses and replaces nothing. The audit store itself is never restored: a backup copies it under a manifest key of its own, audit_store, which v0.4.2.7 ignores, and a restore leaves the live one in place. With the files back, the restored outbox begins a new generation and colony.restored is recorded, naming the backup by its digest and what was restored and set aside. The restore's output says audit store : kept, not restored and how many rows it drained. A store put back by hand, without --restore, is found by the sequencer and recorded as audit.outbox.rewound.

The device module's records (E-12 slice 2). The device module keeps an outbox of its own in its store, micromound_audit_outbox and its head, created when the store opens and drained by the same sequencer beside the colony's; the module's records commit in the transaction of the effect they record, so a device effect exists with its record or not at all. Each is filed in the organization the host placed the device in (device_placements, at token mint), asked of the host before the effect, under the device class: device.enrollment.token_issued (the operator), .burned and device.enrolled (the device: its key's fingerprint, tier, protocol version and what it declared, as ids), .refused with why from a closed set and attempt_count (a token nobody minted names no mound, so that one is unattributed on the platform stream; anyone can present a token, so those are written at most once a minute, each record saying how many refusals it stands for — one standing for more than one names no device — and the host's device sweep writes a count no later refusal came to carry; a refusal of a token the colony minted is one record, attempt_count 1), device.replaced, device.retired, device.unlinked (a purge, with what each table lost — the counts are all that survives), charter.issued (a safety delegation: the limits and the evidence policy as digests, the signing key as a fingerprint, never a subscription reference), charter.superseded and charter.expired (found at the first acknowledged beat past the expiry, once per charter; one that cannot be written never fails the beat, and the next beat writes it), device.action.degraded (the colony could only record unverified: why is a code, never the gate's sentence, which names evidence the device chose), and device.evidence.expired (the sweep's counts, per device). The stop is kept too, once per change: device.stop.engaged and .cleared for the operator's stop and Resume of one mound (by the caller's id), for the colony-wide stop file at the first read that finds it present or gone (by the system, in the installation's organization, with ambiguity: true when the stop is the fail-closed answer of a workspace the colony cannot read, and recorded again with ambiguity: false, naming that one, when a later read finds the file itself there), and for a stop the device took itself (by the device); and device.safe_state.entered at the first beat that reports a stop or a lapsed lease. The operator's reason is never recorded. What a beat carries — its evidence, its action records and the records of verdicts the colony lowered, its mission reports, the records of what it reported — is kept in one unit with its acknowledgement, or none of it is and the device sends it again; a batch the colony cannot keep while a stop is in force is answered with the stop order alone, with no acknowledgement, so a failing trail never keeps a stop from a device. A stop whose record cannot be written is engaged all the same, and its answer says so (audit_unrecorded); a Resume whose record cannot be written is refused (500, audit_unrecorded), and the stop stands. The stop and Resume answers carry audit_event_id. One change an operator can meet: a workspace the colony cannot read now engages the colony-wide stop, as the stop file's documentation always said it did; until this release the check used a call that answers "no file" for a path it cannot read, so the fail-closed branch was never reached. What a device lost is kept as well: device.chain.broken when its uplink chain does not continue from the digest the colony last acknowledged — the digest expected and the one received, the refused batch's range — written once per break however often the device re-sends it (every refusal is still on the bus and in the beat history), with the refusals counted on the mound's record and device.chain.resumed carrying the count when the chain continues; and device.evidence.spilled and .evicted from an evidence bundle's own loss counts, which the colony now reads (a zero count writes nothing, and the bundle is named by its envelope's sequence number, never its id). Two columns are added to micromound_mounds when the module opens its store (chain_break_digest, chain_break_refusals); a store from before them has recorded no break, which is true. A mound placed nowhere is the installation organization's, as every access check on a device reads it. Upgrading adds the two tables and one column (micromound_mounds.charter_expiry_recorded) when the module opens its store, and back-fills nothing: the module's tables hold the state a change left behind, never who acted or when. A mint that is refused no longer leaves a placement behind it either: the mound is placed before its token is minted, so the token's record is filed where the token binds the device, and a refused mint puts the placement back.

GET /audit/platform — the permission read_audit_platform, which only the administrator role holds; coordinators, infrastructure operators, the static token and every service identity are refused (403), and it cannot be delegated. It reads the installation's platform streams, newest first: type (an event type such as operator_shell.attempted), since and until (times), and limit (1 to 1,000, 100 by default). It answers records — every field of each sealed record, with its stream_seq, prev_digest and digest — and legacy_shell_events: the operator_shell_command and operator_shell_error rows an older build wrote into the colony's log, as it wrote them, when the filter names no type or a shell type. It also answers anchored: "none": nothing is countersigned until a later slice. Every read is itself recorded, as audit.read.query, naming the caller, the filter and how many records and legacy rows it returned, and never a record; a read that cannot be recorded is not answered (500, audit_read_unrecorded). A filter it cannot read is 400 (invalid_filter); a colony with no audit store open is 503 (audit_store_unavailable).

GET /audit/status — read_status. How far each stream has verified and whether it is broken, each outbox's generation, how many rows wait above its mark and how old the oldest is, what the sequencer last did, and the trail's health; never what a record holds. The streams it lists are the caller's (E-12 slice 2): the platform's when the caller holds the administrator role, and the streams of the organization the request acts in; another organization's streams are not listed, and the health line counts broken streams without naming one (broken_streams names the caller's own). With sign-in off the one local operator, who reads everything, is listed every stream, the platform's and every organization's. The route is reached from the installation's organization only, so the organization it acts in is the installation's, whoever asks. With sign-in on, an owner there is listed what any other member there without the administrator role is (the administrator role, not ownership, adds the platform's streams), and another organization's streams, where its devices' records are filed, are on no status yet; a break in one is counted in the health line, and named only in the self-test's details.

anthill --audit-verify [--stream <id>] walks every stream — or the one named — from its first record, checking that every digest recomputes, every link holds and no sequence number is missing, and prints three lines for each stream: what verified, what is anchored (none for now), and the tail that is locally attested only. A missing record is reported as a gap, apart from an edit. It only reads: it takes no lease, so it runs beside a running colony. Exit 0 when every stream verifies, 1 on a break, 2 for a stream the store does not hold. The running colony walks each stream from where it last verified a minute after it starts and every hour after that, and from its first record once a day. A break is never repaired: the colony appends audit.chain.discontinuity to the broken stream, the next record chains from it, and the self-test's audit_trail check fails from then on. That check also fails while rows have waited in the outbox for 60 seconds or more, and while the outbox's retention step and its trigger disagree.

The outbox ages out. A row the sequencer has taken is deleted audit_outbox_retention_days after it was taken: 7 by default and never fewer, a floor the outbox's own trigger holds; 0 or less, and 1 to 6, are read as 7 (W1-10's E1, decision F59). The setting is in config.json only. A row the sequencer has not taken is never deleted. The step runs with the record sweep (above), and its count is maintenance_records_expired's audit_outbox_deleted and Compact & prune's audit_outbox_deleted, beside the sweep's own counts: its deleted list and deleted_count are the record sweep's stores, as before. The device module's outbox (micromound_audit_outbox) ages out the same way, at the same setting, with the device evidence sweep (every six hours, three minutes after the module starts): its count is said on the console and, when that sweep also deleted evidence, in micromound_evidence_expired's audit_outbox_deleted. A step the outbox's own trigger refuses is counted in audit_outbox_refused and said on the console as a fault; the self-test reads only the colony's. A store split by anthill --migrate-stores carries the device module's outbox and its head into micromound.db with the module's tables, and the module's next open puts back the triggers the split does not copy.

Not yet. FORAGER's outbox, sign-in, account and membership records, the refusal of high-authority work while the outbox is behind (F34), reading and exporting the trail in the console, per-class retention of the audit store, and checkpoints countersigned by the hosted service are later slices of E-12. Until the last of those, the whole of every stream is locally attested only.


The store split

Until v0.4.2.0 one file, anthill.db, held three data domains: the colony's own tables, the Infrastructure module's 24 and the Micromound module's 12. R-03 §13.2 decided one file per module, and F12 that a table-name prefix is not isolation for the module that can act on real hosts. So:

File Holds
anthill.db the colony — missions, projects, the tenancy spine, and device_placements: where each device belongs, the host's own record since v0.4.2.0
infrastructure.db the Infrastructure module's tables, its credentials and allowlist included
micromound.db the Micromound module's tables, the controller identity included

Each file sits beside anthill.db, is leased (<file>.owner) and hardened like it, and is backed up and restored with it.

A new installation starts split: each module creates its tables in its own file. An installation from before v0.4.2.0 keeps running as it was — each module on its tables inside anthill.db, and the console log saying so at every start — until an operator runs the split (the owner's decision, and F17's posture for anything that moves a customer's data: a person is present).

anthill --migrate-stores [--rehearse] [--backup-out <dir>] [--include-key]

With the colony stopped, in order, stopping at the first thing that is wrong:

  1. The colony's store is opened once under its lease, so its schema is this build's. Migration 0004_device_placements has already copied every device's organization and project out of the device module's table into anthill.db, so no access check depends on a file the host does not own.
  2. What would move is printed. Nothing to move ends here (exit 0).
  3. Every store's lease is taken, and a backup of every store is made and rehearsed exactly as anthill --backup makes one.
  4. The split runs against a scratch copy of that backup. A failed rehearsal ends here; so does --rehearse, having changed nothing but the backup directory.
  5. The split runs against the colony. For each module, through ONE connection that holds anthill.db in SQLite's EXCLUSIVE mode with the module's file attached: every table is copied in one transaction — its own CREATE statement, every row, its indexes and its AUTOINCREMENT counter — then every table's row count and a SHA-256 of every row's typed values are compared between the two files, the module's file must pass integrity_check, and only then are the tables dropped from anthill.db, in a second transaction that records what moved in anthill_meta (store_split.<module>). The moved micromound_mounds loses its org_id and project_id columns on the way, after the split checks device_placements carries every placed device.
  6. anthill.db is compacted (E-8): dropping a table leaves its pages free inside the file, so without this the colony's store stayed the size it was before the split. VACUUM, then a truncating checkpoint so the rewritten pages leave the write-ahead log too; the output gives the store's size on disk before and after. Still offline and under every lease. A compaction that cannot run is reported and changes nothing — the split stands, and Diagnostics → Compact & prune gives the pages back later.

A split that stops between its copy and its removal leaves a module's tables in both files. The colony then refuses to start (exit 13), naming the module and the command, and running anthill --migrate-stores again compares the two copies and finishes only when they are identical. If they differ, it refuses and changes nothing; restore the backup it took, which is named in its output.

Exit codes are the backup verbs': 0 done or nothing to do; 1 refused, with nothing moved; 11 a lease is held; 12 something that takes no lease has a store open.


What the inventory will not do

  • create, modify or delete anything in the installation — not a config file, not a lease file, not a store's -wal or -shm;
  • open any store through SqliteMemory, AnthillRuntime, SqliteMoundStore or InfrastructureRepository;
  • decrypt anything, or print secret material;
  • claim a migration is safe.

The exit code is 0 for "the inventory was taken", not for "this installation is fine". A record whose job is to state what it could not determine cannot carry a verdict, and an exit code that meant one would be the claim this command refuses to make.

Tests

tests/Anthill.Tests/InstallationInventoryTests.cs builds a fixture installation on disk — a SQLite file made directly through the provider, files in a temp tree — and asserts the record's fields, that a missing config file is reported and still missing afterwards, the unfiled counts, manifest drift in both directions, the --out refusal, and that no secret value the fixture planted appears anywhere in the JSON or the summary. The fixture is built by hand rather than through the colony's own types for the same reason the command is: a fixture that had already been migrated could not tell whether the command under test had migrated it.

tests/Anthill.Tests/InstallationInventorySourceTests.cs holds the source guards — the four types, the open through ForeignStore, the absence of any writing statement, the absence of any sealed-column selection, and the verb's position above AnthillRuntime.Initialize().

v2 adds a fixture in WAL mode, stopped cleanly, and compares the whole installation file by file before and after (AStoppedWalInstallation_IsLeftExactlyAsItWasFound); the lease's state for an absent and a held lease; and the FORAGER migration count against FORAGER's own files.

tests/Anthill.Tests/StoreSplitTests.cs holds the split: a new colony opening each module in its own file with none of their tables in anthill.db (F12's validation, as written); a legacy colony split with every row's count and content digest equal before and after, and the modules reading their data back; a split refused while another connection has the store open; a split stopped after its copy — both copies present, the colony refusing to open either, the next run finishing it; an interrupted split whose copies differ refused with both kept; a device whose placement was not carried stopping the device split; AUTOINCREMENT counters travelling with their tables; the rehearsal on a backup moving nothing live; a backup of a split colony holding and restoring every store; a pre-split backup setting the module files aside; a v1 backup still restored; the inventory's module stores; and the source guards — the host resolving and leasing each module store before composing any module, no colony code naming a module's table outside the five files that describe, move or count them, and the verb's steps in order.

tests/Anthill.Tests/StoreLeaseTests.cs holds the lease: a second acquirer refused and told whose store it is; a refused acquirer leaving the store byte-identical; release stamping the record and never deleting the file; a holder in another process named by its own pid and lease id, then killed — never released — and the lease free at once; one lease through a link to the store's file; the lease file owner-only; and source guards that every production path opens the colony's store through SqliteMemory.OpenLeased, which takes the lease before the store is constructed.

tests/Anthill.Tests/ColonyBackupTests.cs holds backup and restore, on WAL fixtures built with plain SQLite connections: the mission backup carrying rows still in the WAL where a File.Copy of the file has none; a backup checked, rehearsed and one file; the key left out unless asked for, never created, and only its check value recorded; the EXCLUSIVE probe refusing an idle second connection and holding out readers while it runs; a restore that keeps the previous store; the deliberately failed restore — a backup damaged in the middle of a page and its manifest re-hashed to match, refused by SQLite's own check with the live store byte-identical; refusals on hash, counts, usability, a different key and a missing key; the colony's own store round-tripped; and source guards that both verbs take the lease before touching the store and a restore checks the backup before taking the lease.

The audit trail's tests are AuditOutboxTests (the outbox's triggers against every write F35 refuses, and its retention step from both sides of the week), AuditRegistryTests (the colony's typed records held to the contract's registry both ways), UserIdTests, AuditStoreTests, AuditSequencerTests (exactly once across a crash on either side of the store's commit), AuditVerifierTests, AuditRestoreTests (a restore drains first, refuses with nothing replaced, and the next record is sequenced once) and OperatorShellAuditTests (the attempt committed before the command runs, the completion naming it, no exception text anywhere, and no command line on any events path), with cases in ColonyBackupTests, StoreMaintenanceTests and TenancySpineTests.