OPERATIONS · v0.4.2.14
Migrating an existing installation
Source: modules/anthill/docs/MIGRATION.md, a path from the root of the product's repository — versioned with the code, rendered here at every release.
Status: inventory, store lease, backup and restore, and the store split. This document describes
anthill --inventory, the first slice of W3-09 (migrating existing ANTHILL / FORAGER / MICROMOUND
installations into FORMICARIA); the store lease, the second: one store, one writer; anthill --backup /
anthill --restore, the third: backups that have been rehearsed before anyone relies on them; and
anthill --migrate-stores (v0.4.2.0, roadmap E-3), which moves each module's tables out of anthill.db
into its own file. W3-09's first rule is inventory before touching anything. The split is the only
command here that moves data, and it backs up and rehearses first; the tenancy retrofit (every row's
organization) is not built.
anthill --inventory <installation-path> [--forager <data-dir>] [--state <micromound-state-dir>] [--out <file.json>]
It prints a human-readable summary, and with --out writes a JSON record. Both say the same things,
including what could not be determined.
The constraint this whole design exists for
Reading an ANTHILL installation through the product's own types is a write.
Constructing SqliteMemory runs InitDb() — SchemaStatements, EnsureColumns,
SchemaMigrationRunner.Apply, ApplyLearningReset — and AnthillRuntime.Initialize() writes
.anthill/config.json when that file is absent. Both are correct for a colony that is about to run.
Both are wrong for a tool whose entire job is to describe an installation as it currently is: a
tool built on them would have migrated the store and created the config file before it printed a
line, and would then report the config file as present, having made it so. "Let me look first" would
be the thing that changed the installation.
So the inventory opens the installation as foreign files:
- the databases through
ForeignStore.Open—Microsoft.Data.Sqlite, read-only, unpooled, andimmutable=1whenever there is no live-wal/-shmpair to read through (below); - everything else through
FileandDirectory, reading sizes, counts and presence.
Read-only is not enough, and record v1 got this wrong. A plain read-only SQLite open of a WAL
store that was stopped cleanly — so its -wal and -shm were removed — creates both files and
leaves them. Record v1 opened stores exactly that way, so "touches nothing" was false by two files
on every cleanly stopped colony; its tests could not see it because their fixture store was not in
WAL mode. ForeignStore reads by what the files mean:
| beside the store | how it is read | what the record says |
|---|---|---|
no -wal |
immutable=1 — no lock, no file created |
the main file is the whole store, all of it read |
-wal and -shm |
read-only, through them | something has it open (or had it open when it stopped); the committed log is included |
-wal, no -shm |
immutable=1 |
the log was not read — reading it means creating its index, a write |
-journal |
immutable=1 |
the rollback journal was not applied — undoing it is a write |
database_read_mode carries that sentence for each store, so a reader knows what the counts cover.
It never constructs SqliteMemory, AnthillRuntime, SqliteMoundStore or
InfrastructureRepository. InstallationInventorySourceTests is a source guard that fails the build
if any of the four reappears in the reader, the record types or the command — the convenience is one
careless edit away, and reintroducing it would be invisible, because the record would still look
right and the installation would already be different.
The verb is answered in Program.cs above AnthillRuntime.Initialize(), beside --version and
--help, so running it cannot leave a workspace wherever the operator happened to type it.
Two consequences worth stating, because they look like bugs and are not:
- A store whose
-waland-shmwere left by a process that died can need recovery, and then refuses to open read-only. Recovering it would be a write. The refusal is recorded asdatabase_open_failureand is the correct outcome. --outrefuses to write inside the installation it just read, including inside a FORAGER data directory or a MICROMOUND state directory passed on the same command line. Writing the record into the thing you promised not to touch is the exact bug the rest of this design avoids, and--out ./inventory.jsontyped from the colony's own directory is the reflex that produces it.
The record
JSON, with a record_version of its own — independent of any product version, because what changes
shape across releases is the record, and a reader needs to know which shape they are holding. Top
level:
| key | what it is |
|---|---|
record_version |
this record's shape. "2" today: v2 added store_lease and database_read_mode (below), and v1's FORAGER note compared against a migration count that had gone stale. |
taken_at_utc, taken_by |
when, and by what. |
constraints |
what this record deliberately cannot prove, carried inside it — a limit stated in a tool's help and not in its output is a limit the output does not have. |
anthill |
the colony: config resolution, workspace paths, the store. |
forager |
the knowledge engine, if a data directory was given or found. |
micromound |
the device state directory, if one was given. |
not_determined |
every section's unknowns, rolled up. |
Absence is always a stated value and never an omitted key. Every section carries its own
not_determined; every fact that could not be read carries a note saying why. A blank where a
fact should be is the failure mode this record exists to prevent: a plan made against a field that
is empty because nothing looked is a plan made against nothing. unreadable is likewise never
folded into absent — a permission problem reported as a missing directory invites a migration that
recreates what is already there.
ANTHILL
Config resolution
The rule the runtime would have used: ANTHILL_CONFIG_FILE when set, else
<root>/.anthill/config.json, with a relative value resolving against ANTHILL_HOME when that is
set and against the working directory otherwise. Both variables are reported as this process sees
them; the path is resolved against the root you named.
config_file_exists is reported, never created. When it is false, the workspace paths below it
are the defaults this installation would take — not values somebody wrote down, and the record says
so.
legacy_config_* names data/anthill.json. The runtime no longer reads it and only warns about it.
It is in this record because an inventory that omitted it mis-describes the installation's actual
configuration: an operator reading that file believes the settings in it are in force, and there is
a real field failure behind that — a live qualification run spent a whole mission with
acting_coder_enabled true in data/anthill.json and false in the runtime, with nothing reading the
file and nothing erroring.
Workspace paths
workspace_root, db_path, backup_dir, logs_dir, exports_dir, agent_workspace_dir — as the
config states them (or as defaulted), each resolved, each with kind (file / directory /
absent / unreadable) and a size where there is one.
The database
anthill.db with its size, plus the presence and size of -wal and -shm. Those two are part of
the database: a copy of the .db alone is the store as it stood before the last committed writes.
Read from inside it, directly:
anthill_meta— it already carriesschema_version,anthill_version,config_profile,config_path,workspace_rootanddb_path. Values are reported exactly as stored, which is JSON-encoded.anthill_schema_ledger— the migrations that ran. Each row was written in the same transaction as the change it names, so this ledger cannot record a change that did not happen.- the table list from
sqlite_master, and a row count per table.
store_lease — the lease's state, looked at and never taken (see The store lease, below):
absent, free, held or unknown, the holder the store recorded, and the holder the lease file
names. It is read before the store is, so the record says whether something was holding the
installation while it was being inventoried.
module_stores (v0.4.2.0) — each module's own file beside anthill.db (infrastructure.db,
micromound.db), read the same way: whether it exists with its -wal and -shm, its lease looked at,
its tables and row counts, and where the module's tables are — own, legacy (still inside
anthill.db: anthill --migrate-stores has not run) or interrupted (in both files: a split that
stopped, which the colony will not start on). The manifest comparison and the controller identity look
across all of them.
schema_migrations is NOT evidence
The 22-entry schema_migrations table is reported, and labelled evidence: "none".
Its own header says why: the runner that writes it inserts a row per entry and executes no DDL,
and every description in it says something was "verified" when nothing verified anything. The real
schema work was always SchemaStatements and EnsureColumns, unconditionally, on every start.
An inventory that read this table to decide what an installation has would be wrong twenty-two
times. So it is reported — a reader will find it and needs to know what it is — and nothing in
this record is derived from it. What an installation actually has is anthill_schema_ledger plus
the tables and columns the record lists.
The retrofit numbers
W1-02 §4.5 phase 5 names these as the check, so they are taken before anything moves.
null_org_id— per table, for the anchor and the scoped tables. Must reach zero. A row with noorg_idis a row no organization owns.null_project_id— formissions,conversations,objectivesanddevice_placements(where a device belongs; until v0.4.2.0 that was two columns ofmicromound_mounds).
The second number is read the same way and means the opposite thing, which is why each count carries
its meaning in the same object as the count. project_id IS NULL is unfiled (W1-02 §4.2): a
defined placement, org-scoped and reachable by the organization's owners. It is not a defect and it
is not a backlog.
This number is to be compared across a migration, never minimised. It must come out exactly as it went in. A count that fell means something invented a project for a row that had none — and which project a two-year-old mission belonged to was never written down, so no migration can recover it.
A table that is not in the store, or a column that this store predates, reports null with a note —
unread, never zero.
Manifest drift, both directions
The store's tables compared against TenancyTables.All (plus the module tables), because that is the
one existing, maintained inventory of what is in anthill.db. A second list written beside it would
drift from it the first time a table was added, and then the drift report would itself be drifted.
in_store_not_classified— a table nobody has decided the tenancy of. A migration would carry it blind.classified_not_in_store— ordinary for a table created lazily on first use, and reported so a reader can tell that case from a table that was expected and is not here.
excluded_name_patterns names what the comparison deliberately does not see (sqlite_%,
missions_fts_%), rather than dropping them quietly.
Identity, by presence and never by value
- whether
micromound_controller_identityholds a row — asked withCOUNT(*); theseedcolumn is never selected; - whether
.anthill/field.keyexists (and its mode), and whetherANTHILL_ENCRYPTION_KEYis set — the variable's value is never read into the record; provider_credentialsandcolony_credentials: counted, never read, nothing decrypted.
verified_to_open is the constant false, and that is the honest field. The record can say "a
controller identity is present and I cannot verify that it opens" without opening the fleet.
This matters more here than anywhere else in the tree. The controller signing identity is the one
identity that cannot be re-minted: ADR-009 allows no rotation, only physical re-enrolment of every
mound. Its seal lives outside the database — the field key file, or the environment variable of
the process that runs the colony. So "copy the database" is not "copy the installation", and a
migration that moved anthill.db and left the key behind would produce a colony holding a fleet
identity nobody can use.
The id namespaces, because three disagree
Each id is reported with the namespace it lives in and who minted it:
| field | namespace | minted by |
|---|---|---|
installations.id |
ins_ |
locally, by migration 0001_tenancy_spine, from RandomNumberGenerator |
installations.fingerprint |
inf_ |
the same |
organizations.id |
org_ |
the same — and deliberately unnamed, because a name a migration chose would be an invented value |
| FORAGER's organization | org_local |
a constant, stamped by FORAGER migration 014_tenancy on every store |
A copy of a store that has not yet run 0001 mints different values for the first three, while
every FORAGER store shares the fourth. A record saying "this installation's organization is X"
without saying which kind of X that is, is a claim it cannot support.
FORAGER
Given with --forager <data-dir>. Failing that, two places FORAGER itself documents are looked at:
FORAGER_DATA_DIR in this process's environment, and <root>/data (that variable's default,
relative to the process's own directory). The section's note says which of those it was, because a
directory the command guessed and one an operator named are different kinds of fact.
A packaged install keeps its data in a per-user location this command cannot guess, so no FORAGER
section means not looked at — never "this installation has no FORAGER". The --out refusal
covers a discovered directory as well as a named one: a refusal computed from the arguments alone
would let the record be written into a store the command had just read.
forager.dbwith its sidecars.- Its
schema_migrationsledger — that one is evidence, unlike ANTHILL's frozen table: its runner executes each change and writes the row in the same transaction. The FORAGER in this tree ships 22 migrations (InstallationInventory.ForagerMigrationsShipped), so fewer means not all have run and more means a newer FORAGER wrote this store. Record v1 said 15, a day after016and017landed, and so would have called every current store "newer"; the number is now pinned by a test that counts FORAGER's own migration files, which caught it again when E-18 added018to021. E-16 added022(the organization's cache scope) and moved the number with it. instance.markerread as a file andforager_instanceread as a row, reported side by side and never reconciled. Deciding whether they agree is what the engine's ownloadInstancedoes, and it does it by minting a new generation and rewriting the marker — asking the product this question changes the thing you came to record. A human reads both: a marker naming another generation means the database file was swapped under it; a row with no marker means the database arrived in a directory it has not run in.operator.key— presence and mode, never contents.uploads/,exports/,tmp/— file counts and sizes.
The store lock: participating / empty / unprovable
participating—store_lockexists, so any holder is an engine that writes a row in it and the lock means what it says.empty— no tables at all; nothing has ever written this store, so nothing can be holding it.unprovable— tables, but nostore_lock: a store written by an engine from before migration 012, which can only be held by an engine that never writes a row there.
The third word is kept. An empty lock table proves the table is empty and proves nothing about
the store. Rounding unprovable to "free" is what lets two engines write one store. When the
database was not opened at all the state is unknown, which is also not "free".
FORAGER's own classifier creates the lock table when it is missing — right for an engine about to serve the store, wrong for a command that promised to touch nothing. The classification is therefore re-read here rather than called.
MICROMOUND
Given with --state <dir>.
identity/seed— presence and mode, never contents.state/andevidence/— file counts and total sizes./etc/micromound/manifest.jsonand/etc/micromound/micromound.env, the locationsdeploy/installs to. Absent there means "not in the standard place", never "this mound has no manifest", and the record says so innot_determined.
If the env file is found, the record names the hazard: it carries MICROMOUND_ARGS, and
--enroll-token in it is an enrolment token in plaintext. It must not travel in a migration
bundle, a backup that moves, or a support archive. The token itself is never read and is not in the
record.
The store lease
One store, one writer. W3-09's third rule is the one that silently corrupts: never let two processes write the same database, and never let a version-skewed process migrate a store a live one is serving (F74). The lease is how ANTHILL keeps it, and it is two things — the owner's decision on F20's open conflict, 20 September 2026:
- The lock: an OS lock on
<db>.owner, the file beside the database.flock(LOCK_EX|LOCK_NB)on Linux and macOS (W1-11 §5.3's executed mechanism), a byte-range lock (FileStream.Lock, which isLockFile) on Windows (F20's value). It is the enforcement. The OS releases it when the process dies, so a crashed holder's lease is free at once — there is no heartbeat, no staleness rule, and no comparison of two machines' clocks. - The record: a
store_leaserow in the store, and the same facts in the lease file. Pid, hostname, mode, purpose, version, when acquired, when released. Diagnosis only — a row is never evidence of holding, and one left by a process that crashed says who held the store last. On Linux a held lease file cannot be read by another .NET process at all, so the row is how a refused process names its holder.
Order is the point. The lock is taken before the store is opened, so a refused process has written nothing to the database. The holder then writes its row and only then constructs the store, whose schema pass and migration runner therefore always run under the lease.
Who takes it. The API host, for its life, and every CLI verb that opens the colony's store —
including the ones that look like reads (--status, --config, --routes, --live-qualification),
because opening the store runs its schema pass; there is no read through it. While a colony is running
those verbs are refused, name the holder, and exit 11 — FORAGER's code for the same refusal.
Lock-out recovery is therefore stop, --set-password, start. A second --api on the same store
refuses to start the same way; under the LXC unit's Restart=on-failure it retries every five
seconds and starts once the store is free. --inventory never takes the lease; it looks at it.
What it cannot do.
- An advisory lock binds the processes that ask. Every ANTHILL before v0.4.0 takes no lease, and
neither does a customer's
sqlite3shell.absent(no lease file) andfreeare therefore never proof that a store is idle, and the record keeps them apart for that reason. - Never delete
<db>.owner. On Linux, deleting it while a holder runs lets a second process lock a new file of the same name. Nothing in ANTHILL deletes it; release stamps its facts and unlocks. - The file is owner-only (0600) on POSIX, because anyone who can open it can hold a shared lock on it and keep the colony from ever taking its lease.
- A store named through a symbolic link to its file resolves to the file; a hard link cannot be seen from its path, and two names for one inode through hard links are not recognised as one store.
- If the filesystem cannot lock (some network mounts), the lease cannot be taken and the process refuses rather than open the store unguarded.
Backup and restore
W3-09's second rule — consistent backups after quiescing writers, validated, rehearsed — is
anthill --backup and anthill --restore (v0.4.1.2). Both are offline (owner's decision,
20 September 2026): they take the store's lease, so while the colony runs they are refused, name the
holder and exit 11 like every other verb that opens the store. Stop, back up, start. A running
colony backs itself up with Back up now instead (below); a restore is offline only.
Every store of the colony (v0.4.2.0). The colony is anthill.db and, once split, a file per module
beside it. A backup takes every one that exists — each file's lease held, and all of them in SQLite's
EXCLUSIVE mode together while they are copied, so the files in a backup are one moment of the colony —
and its manifest (formicaria-backup-v2) lists them under stores, anthill.db first. The manifest's
top-level fields still describe anthill.db and are what its copy is checked against; a manifest whose
two records of it disagree was edited, and is refused. A v0.4.1.2 backup (formicaria-backup-v1,
anthill.db alone) is still verified and restored.
A backup is a new directory under backup_dir (or --out), formicaria-backup-<time>-<id>,
holding anthill.db, each module store that exists, and manifest.json, taken in this order:
- Quiesced. The lease keeps out every lease-aware ANTHILL. A writer from before v0.4.0 takes no
lease, so the copy is taken through a connection in SQLite's EXCLUSIVE locking mode. A WAL
store refuses that while any other process has the file open — idle, reading or writing — and
the verb exits 12 naming what to look for (an older ANTHILL, a
sqlite3shell, a database browser). The lock is held until the copy is taken. - Copied through SQLite's online backup API, which reads through the WAL, then switched to rollback-journal mode: one self-contained file.
- Checked —
PRAGMA integrity_check, a row count for every table, a SHA-256, and a one-way check value of the field key. - Rehearsed — restored into a scratch directory and checked again.
The manifest records all four and says "usable": false (exit 1) if any failed; such a directory is
kept for inspection and a restore refuses it.
The field key is not in a backup unless --include-key is passed. It seals device seeds and the
MICROMOUND controller identity (ADR-009: no rotation, only re-enrolment), so a directory holding the
store and its key is a copy of everything. The check value is what lets a restore tell whether the
workspace's key is the one the backup was sealed with. A key from ANTHILL_ENCRYPTION_KEY is never
in a backup at all. A backup never creates a key — the cipher's own constructor mints one when
none exists, which is why the backup path reads the key without constructing it.
A restore (anthill --restore <dir>) runs every check that reads only the backup and the key
file before it takes the lease — taking the lease writes its holder into the live store, and a
refused restore must leave the store byte-identical. It refuses:
- a directory with no manifest, a manifest of another format, or one marked not usable;
- a store file whose SHA-256 differs from the manifest (changed or damaged since it was taken);
- a store that fails
integrity_check, or whose counts differ from the manifest; - a backup sealed under a different key than this workspace's — every sealed value would be unreadable after it;
- a backup sealed under a key when this workspace has none, unless the backup carries it;
- a store that something else has open (exit 12).
Then, under every file's lease, it keeps each store it replaces as <file>.pre-restore-<time>, stages
each file of the backup as <file>.restoring-<time>, checks the staged copies, moves them into place,
checks them there, and puts the backup's key in place if it had to. A module file the backup does not
hold is set aside — copied to <file>.pre-restore-<time> and removed — because the backup was taken
while that module's tables were still inside anthill.db, and leaving today's file beside the restored
store would put the module's tables in two files at once. anthill --migrate-stores splits it again.
Restoring onto a machine with no colony is the same command; the empty store a lease's record created
there is not kept. A backup from an older schema
is restored as it is and migrated when the colony next opens it. --rehearse restores into a
scratch directory only, takes no lease, and can run beside a live colony.
While the colony runs: Back up now (roadmap E-6). Settings → Diagnostics → Back up now, or
POST /maintenance/backup (the manage_settings permission), takes the same backup from the running
colony, into the same folder, in the same format; anthill --restore restores it like any other. What
differs is how the copy is held still:
- Every store's write lock, together. The colony takes
BEGIN IMMEDIATEon a connection of the backup's own for each store —anthill.dbfirst — and holds them all until the last store is copied, so the copies are one moment of the colony, as the offline backup's are. Each store is copied by the backup API through a second, reading connection; in WAL mode readers are never blocked, so the console and every read go on answering. - Writes wait, and not for long. A write arriving during the copies waits (every connection of the
colony has a 30-second busy timeout) and goes through when they are done. The owner set the limit
(24 September 2026): ten seconds. The copies go a megabyte at a time with the clock read between,
so a colony too big to copy in ten seconds is found out as it is copied: the backup stops, lets every
writer through, deletes what it had written and says to stop the colony and run
anthill --backup, which has no limit. Nothing that is not a complete backup is ever left in the folder. - The field key only on an explicit opt-in (owner, 24 September 2026). Back up now leaves it
out; Back up with key is its own action behind its own warning, and the endpoint takes the key
only for
{"include_key": true}— not"true", not1. The manifest says so, as it does for--include-key. - The checks, the rehearsal and the manifest are the offline backup's, run after the locks are
released. The manifest's
quiescesaysonline:and how long the writes waited, out of how long they were allowed.
One runs at a time; a second is refused while the first is running. Each is recorded in the colony's
event log (maintenance_backup, or maintenance_backup_refused with the reason). The console lists the
full backups in the folder — how many, how much disk, and the newest, from their manifests.
Delete all backups does not touch them: it acts on the mission path's snapshots. Compact & prune
deletes the ones past the retention window, as every sweep does (below).
The mission path. Before a mission (at most once per backup_min_interval_minutes) the colony
still writes backups/anthill_<time>.db. Until v0.4.1.2 that was a File.Copy of the main file of a
live WAL database, which lacks every write since the last checkpoint. It now goes through the backup
API and is deleted if it fails PRAGMA quick_check.
Backups age out (v0.4.2.5, roadmap E-14). Every backup in backup_dir — the full backups
(formicaria-backup-*) and the mission path's snapshots (anthill_*.db) — is deleted
backup_retention_days after it was taken: 30 by default, and held to 7–90 (decision F66; W1-10 §4.2).
Until v0.4.2.5 nothing aged out: the snapshots were pruned by count (max_db_backups, which still
applies) and a full backup was kept until somebody deleted it.
- Dated by what the backup says about itself. A full backup by its manifest's
created_at, or the time in its directory's name when the manifest cannot be read; a snapshot by the time in its name, or its last write when the name carries none. A backup exactly as old as the window is kept, and one that cannot be dated is kept and reported. - Nothing else in the folder. The patch tool's
.bakcopies have a retention of their own (W1-10's A4), an operator's own files are theirs, and a link is not followed. A backup written withanthill --backup --out <dir>is in a folder the operator chose and is never swept. - When. A minute after the colony starts and every six hours after that, and also after a mission's
backup step, Back up now (only when the new backup is usable), Compact & prune and
anthill --backup. Changing the window in Settings → Diagnostics → Storage (or/settings) takes effect at the next sweep. - Said on the event log. A sweep by the running colony that deleted, could not delete or could not
date something records
maintenance_backups_expired— the names, how many, the bytes freed, the cutoff — and says so when it has left the folder with no full backup at all. Back up now answers with the names of what it expired and Compact & prune with how many;anthill --backup, which runs with the colony stopped, prints them.
Why it is a promise and not a preference: an erasure cannot reach into a copy, so for any backup a deletion is complete only when that backup has expired (W1-10 §5.7). To keep a backup longer, copy it out of the backup folder. Not built: §5.7's tombstone log, which a restore replays so that a restored backup does not bring erased data back; what may be promised about backup expiry is a question for counsel (L-06).
The colony's own records age out (roadmap E-18; W1-10 §4.2, decision F59). What the colony said to
itself while working — agent messages, artifacts, task results, events and source records (W1-10's A3)
— is deleted event_retention_days after its mission ended: 180 by default. 0 keeps it all, 1 to 29
are read as 30, and a number above 3650 (ten years) is read as 0 — the colony never deletes sooner than it
was asked to. Approvals and escalation decisions (A7) and message metrics and task-result summaries (A8) are
deleted two years after they were written, whatever the setting says.
- Counted from when the mission ended, which is the later of the mission's last save and its evaluation. A row whose mission has been cleared ages by its own time.
- The colony's own log is not swept. The
system_apimission's events are the record of what the operators and the API did — the operator shell's command lines (written before each command runs), memory wipes, resets, deleted backups, entitlement and directory decisions, and these sweeps' own events. They are not agent traffic, and no setting here deletes them, in any store: their retention belongs to the audit log that is to hold them (W1-05 §8.2; decision F64 keeps the shell command line with a 90-day floor), which does not exist yet. Until then they are kept, as they were before E-18. (Since E-12 a new shell command's line is written to the audit trail and not to this log; the rows written before stay here until E-12's slice 7 hands them over. See The audit trail, below.) - Never touched: a mission that is running or has not started, or one that can still run again — a pending request other than a patch's (approving it replays the refused step from the mission's own records) or an API job that has not finished. A pending approval is never deleted, whatever its age.
- A mission a stopped colony left running is ended at the next start (roadmap E-18's remainder). Until
this release a mission whose process died — the host killed, the machine rebooted — kept its row
runningfor good: the next start re-queued its job as a new mission and abandoned its task attempts, but nothing closed the mission, so the sweep never touched its rows or the conversation that started it, the conversation read "running mission …" for ever, and Clear missions and Wipe memory refused while it stood. Now the colony, when it starts, ends every mission row a process that is gone left running (or created with work begun on it) as a failed mission: outcomefailed_permanent, stop reasonprocess_died, the reason on the row naming its job and what that start did with it — retried as the next attempt (the retry runs as a new mission, and the ended one is told which:mission_retried_as), orphaned for operator review, cancelled, or nothing to retry. Its events getmission_failedandmission_outcome, so the console's notifications and failures panel see it. It is counted as ended at its last sign of life — the latest moment the store saw the process working on it: its last save, event, task or task attempt, or its job's last heartbeat (renewed every 30 seconds while a job runs) — and its records age from then, as every other mission's age from its end and not from when the colony noticed; the reason on the row says the moment. A mission left running in this process's own lifetime is never touched. Tasks keep the status the dead process left them in; the attempt ledger says how far each got. What the operator sees after the upgrade's first start: onemission_failednotification per mission an earlier process ever left running, however old; each keeps its place in the mission history (itssaved_atis its last sign of life, not the start), and itscreated_atstill says when it began. The colony's own recovery record (mission_attempts, understartup-reconcile) is now written once per start that recovered something, numbered on; until this release it was written once per store, ever, so a colony that had recovered a crashed job before kept only that first record. - Kept until a patch is decided: a mission whose patch is proposed, or approved and not applied, keeps its task results (applying the patch reads the soldier's verdict) and its approvals; the rest of its traffic ages as usual.
- Also kept whatever their age: the two finalization-ledger events that stop learning being counted twice; an artifact a live mission reads, read by id, or cites, and what those cite; the source and recall records of every mission a live mission recalled, as far as grading its citations reads them; and the escalation decision a conversation still acts on — past two years that one keeps the decision and loses its reason and its author's name. Missions, objectives and schedules are not this sweep's (W1-10 A5 and A6); conversations are, only when an auto-purge is chosen (below).
- When. Two minutes after the colony starts and every six hours after that, and on Compact & prune,
which compacts the store afterwards. Until the store is compacted a deleted row's bytes can stay in its
free pages. The window is edited in Settings → Diagnostics → Storage (or
/settings) and takes effect at the next sweep. - In steps. A sweep deletes at most 5,000 rows of a store per transaction and lets the colony write
between them, so even the first sweep after the upgrade holds a mission's writes for one step, not for the
whole sweep. A sweep that deleted anything then empties the write-ahead log (
anthill.db-wal), briefly, and leaves the rest to the next checkpoint if a reader is in the way. - Said on the event log. A sweep that deleted or scrubbed something records
maintenance_records_expired— how many rows per store, the two cutoffs, how many missions it left alone and why — and never a row's content. The first sweep a colony has is always recorded, and says it is the first.
Upgrading moves the default. Until E-18 event_retention_days applied to events alone, only from
Compact & prune, and defaulted to 0 — keep everything; the settings save wrote that 0 into every
config.json it touched, and no console page showed the key. So configuration schema 4 reads a 0 (or an
absent key) in a file written before it as that old default and moves it to 180, and keeps any other
number, which somebody chose. The first sweep after the upgrade then deletes, at once, everything already
older than 180 days after its mission ended, and its event says what it removed. To keep everything
through the upgrade, set event_retention_days to 3650 before upgrading — a 0 written by an older
version reads as the old default — and to 0 afterwards; at schema 4 a 0 is kept as the operator's.
Conversations, customer source and device evidence (roadmap E-18, second build; W1-10 §4.2 and §4.4, decisions F59 and F60). The same sweep, on the same timer and on Compact & prune:
- Conversations (A1, A2) are kept for the life of the conversation by default.
conversation_retention_dayschooses an auto-purge — 90, 180 or 365 days after the conversation's last turn, or 0, never (the default); a number between two choices reads as the longer one, and one above 365 as never. A purged conversation goes with its turns and attachments, pinned or not. Its missions stay (they are A5, "life of the project"), and so do its escalation decisions, which go two years after they were made (A7). Never purged, whatever its age: a conversation whose mission can still run or has a patch nobody has decided, one waiting on an operator decision, and one a schedule run is still running in. - Customer source in a patch (A4). A patch that was applied, rejected, failed, superseded or reverted
loses its old and new file contents
patch_retention_daysafter that decision — 90 by default, held to 30–365, with no "never"; 0 or less is read as the default, 90 — together with the pre-apply.bakcopies the patch tool took for it and the repository index of the workspace it came from. The patch's row stays: status, file, hashes, who decided. Only copies the colony recorded taking are deleted, and only from the backup folder; nothing is looked for. A patch that can still be applied (proposed, approved) keeps everything, and so does a decided one while another patch of its set is undecided or its mission has not ended. Once the backup is gone, reverting a modify or delete is refused and the patch stays applied, and an alternative can no longer be built from it. The Patch Center says when the content was removed. - Device evidence and action records (section 4.4) are deleted
device_evidence_retention_monthsafter the colony received them — 24 by default; 0, or more than 120, keeps them — and never fewer than 12: that floor (decision F60) is a constant in the code, not a setting, so 1 to 11 are read as 12 and no plan or setting lowers it. An action whose mission has not reported is kept, with the evidence it cites. The device module sweeps three minutes after it starts and every six hours, and writesmicromound_evidence_expiredto the colony's event log (with the counts, the window and the floor) when it deleted something.POST /micromound/purge(a retired device's erasure) is unchanged.
Upgrading. A patch decided before this release has no recorded decision time; the first sweep dates it
from what the store holds — applied_at for an applied or reverted patch, the rejecting approval for a
rejected or superseded one, otherwise the sweep itself — so the first sweep after the upgrade removes at once
the content and recorded backups of every patch decided more than 90 days before. To keep them through the
upgrade, set patch_retention_days to 365 first. Device records carry the time the colony received
them from this release on; a record from before is dated by the upgrade, so none of them goes for two years.
Conversations are untouched unless an auto-purge is chosen.
FORAGER's windows, from the console (roadmap E-18's remainder; W1-10 §4.1, decision F59). The FORAGER
engine the colony runs keeps what it derives for as long as six windows say. Until this release they could be
set only in FORAGER's own environment (FORAGER_RETENTION_*), and F8's 0, which W1-10 wants to be one click,
was an environment key. They are now settings, edited in Settings → Diagnostics → FORAGER's records (or
/settings):
| Key | What goes, and from when | FORAGER's default, range |
|---|---|---|
forager_cache_retention_days |
a parser or extractor cache row, after it was last used (F5, F6) | 90 days, 0–365; 0 turns both caches off |
forager_model_log_retention_days |
a logged model call, with up to 1,000 characters of the model's answer, after the call (F8) | 7 days, 0–30; 0 keeps none of the model's output, a choice of its own in the console |
forager_review_event_retention_years |
a review event, after it was recorded (F14) | 2 years, 1–7 |
forager_job_retention_days |
a finished job and its stages, after the job finished (F15); a project's latest job stays and loses only its text | 90 days, 30–365 |
forager_stale_retention_days |
a statement, relationship or entity no live source supports any more, after it lost that support (F9–F11) | 180 days; 0 at the next sweep, -1 never |
forager_export_retention_days |
a built export's zip, after it was built (F16); the export's record stays a year | 7 days, 1–90 |
- Unset by default, and unset changes nothing. Each key is
nulluntil somebody sets it. The colony then hands the engine nothing for that window, so FORAGER's own default — the decided value — applies, or aFORAGER_RETENTION_*variable in the colony's environment if it has one, exactly as before this release. - Set, the colony's setting wins. A set window is held to FORAGER's range (a number past either end reads as
that end) and handed to the engine as FORAGER's own variable, over whatever the colony's environment carries —
the way the supervisor's other variables, such as
FORAGER_DATA_DIRandFORAGER_PORT, already are. The stale window's -1 (any negative number) is handed over as the wordnever. Clearing the setting —null, or Not set in the console — gives the window back to the environment and FORAGER's default. - When a change takes effect. The engine reads its environment when it starts, and a save does not restart
a running engine: a change applies at the engine's next start — the colony's next start, the next time
Knowledge is turned on (a restart after a crash picks it up too), or Restart FORAGER now. The engine sweeps
when it starts and every six hours after. For each window the console says what the engine running now has,
and whether a change is waiting for its next start;
/maintenance/statscarries the same asforager_retention— per window what the next start reads (next), what clearing the setting would leave (unset) and what the running engine started with (running), each with where it comes from (colony,environmentorforager). - Restart FORAGER now. While a saved window differs from what the engine running now has, FORAGER's records
offers the restart (
POST /maintenance/restart-forager, an administrator's). It asks FORAGER first, with the host credential, what it is working on, and refuses (409forager_busy) while any project the credential reaches has a processing job queued or running or an export being built, leaving the engine running as it was; otherwise the engine is drained, stopped and started again, andmaintenance_forager_restartrecords who and which windows. It refuses as well when the colony runs no engine, when Knowledge is off, when nothing is waiting, when FORAGER cannot say what it is working on, and when the subscription would not let the engine start again. - A stop drains the engine. Before this, every stop of the engine the colony runs — Knowledge turned off, the colony shutting down — raced its own drain: the supervisor killed the process while it was still asking it to shut down. The stop now waits for the engine to drain and exit (ten seconds, then a kill, as documented), and a start asked for meanwhile waits for that exit instead of putting a second engine on the store.
- A value outside FORAGER's range is refused by the console. The server still holds one it is sent (or finds
in
config.json) to the nearest end of the range; the console refuses it by name before it is sent, because the bottom of each range is the end that deletes the most. The stale window's negative is never, as before. - Only the engine the colony runs. A FORAGER you run yourself (attached mode) reads only its own environment, and these settings do not reach it.
- Upgrading changes nothing. The keys are absent from an older
config.json, read as unset, and written asnullby the next settings save.
What backup and restore do not do. Restore into a running colony; back up FORAGER's store, which
has its own; copy anything off the machine; see an idle reader of a store that
is not in WAL mode (the EXCLUSIVE probe cannot, and every ANTHILL store is WAL). The mission path's
backup is anthill.db alone: missions write nothing to a module's store.
The audit trail
Roadmap E-12, slice 1: W1-05's audit design as the E-12 slice plan builds it (audit-slice-plan.md, in
the Formicaria repository's transition workspace, under security), and decisions F35 and F64.
What the colony keeps. An audit record is written into an outbox in anthill.db (audit_outbox
and its one-row head audit_outbox_head), in the same transaction as the change it records, or, for an
effect outside the store such as a shell command, committed before the effect runs. The outbox's triggers
refuse every change but the drain stamp, and every deletion but retention's (F35). A single sequencer
in the running colony drains it into the audit store beside anthill.db: audit.db, its index, and
the records themselves in audit/<stream key>/<n>.db. Each record is filed in a stream of its scope and
retention class — platform:<installation>/<class> for the installation's own records,
org:<organization>@<installation>/<class> for an organization's — and chained there: a dense sequence
and a SHA-256 over each record and the one before it. The store's triggers refuse any update or deletion,
and no table in it has a foreign key. A record names a person by an opaque usr_ id, never by name:
every account has one (users.user_id), and it never changes. Never edit, vacuum or delete the audit
store by hand; Compact & prune never compacts it, and the storage footprint lists it as audit.
Upgrading is additive, at the first start of the new build. The ledger gains 0005_audit_outbox (the
empty outbox, its head at a new generation, and its triggers) and 0006_user_ids (a usr_ id for every
account), and audit.db and audit/ are created empty. Nothing is back-filled: a chain covers records
as they are written. An older build refuses the upgraded store, naming both migrations; the way back is
the backup taken before the upgrade, and this build runs both migrations again on it. An account an
older build creates after a downgrade gets its id at the next start.
Two things an operator will notice:
- Coordinators no longer see shell command lines. The operator shell records each command twice:
operator_shell.attemptedbefore it runs — the command line, cut at 4,096 characters withcommand_truncatedsaying so, whether it ran in the default directory (never the path), and the operator'susr_id — filed inplatform:<installation>/operator; andoperator_shell.completedwhen it ends — exit code, whether it timed out, elapsed milliseconds, and the attempt's event id ascausation_id— filed inplatform:<installation>/config. A timeout completesunknown; a command that did not start completesfailedwitherror_code: shell_start_failed, and the exception's text is recorded nowhere (the operator who ran it is still told why). A command whose attempt cannot be recorded is not run (500,shell_audit_unrecorded). Theoperator_shell_*rows in the colony's log stay, for the live console, and carry the audit record's event id instead of the command, the directory or the error. The rows an older build wrote keep theirs in the store, and every events path —/events/json,/events, the live stream and its replay, Colony Live, the reports, the memory vault and the SDK's event log — reads them withoutcommand,diranderror, and readsoperator_shell_error's message as "Operator shell command failed to start." The one place a command line is read isGET /audit/platform, which only administrators can call (below). Decision F64 keeps the command line on these conditions: local only, installation operators only, 4,096 characters, one field wide. anthill --restoredrains the audit outbox first, and refuses when it cannot. Before it replaces any file, the restore runs the sequencer to the end of the outbox in the store it is about to replace, so nothing waiting there is lost with the replaced file. When the audit store cannot take the rows — it cannot be opened or written, with rows waiting — the restore refuses and replaces nothing. The audit store itself is never restored: a backup copies it under a manifest key of its own,audit_store, which v0.4.2.7 ignores, and a restore leaves the live one in place. With the files back, the restored outbox begins a new generation andcolony.restoredis recorded, naming the backup by its digest and what was restored and set aside. The restore's output saysaudit store : kept, not restoredand how many rows it drained. A store put back by hand, without--restore, is found by the sequencer and recorded asaudit.outbox.rewound.
The device module's records (E-12 slice 2). The device module keeps an outbox of its own in its store,
micromound_audit_outbox and its head, created when the store opens and drained by the same sequencer
beside the colony's; the module's records commit in the transaction of the effect they record, so a
device effect exists with its record or not at all. Each is filed in the organization the host placed the
device in (device_placements, at token mint), asked of the host before the effect, under the device
class: device.enrollment.token_issued (the operator), .burned and device.enrolled (the device: its key's
fingerprint, tier, protocol version and what it declared, as ids), .refused with why from a closed set and
attempt_count (a token nobody minted names no mound, so that one is unattributed on the platform stream; anyone
can present a token, so those are written at most once a minute, each record saying how many refusals it stands
for — one standing for more than one names no device — and the host's device sweep writes a count no later refusal
came to carry; a refusal of a token the colony minted is one record, attempt_count 1), device.replaced,
device.retired, device.unlinked (a purge, with what each table lost — the counts are all that survives),
charter.issued (a safety delegation: the limits and the evidence policy as digests, the signing key as a
fingerprint, never a subscription reference), charter.superseded and charter.expired (found at the
first acknowledged beat past the expiry, once per charter; one that cannot be written never fails the beat, and
the next beat writes it), device.action.degraded (the colony could only
record unverified: why is a code, never the gate's sentence, which names evidence the device chose), and
device.evidence.expired (the sweep's counts, per device). The stop is kept too, once per change:
device.stop.engaged and .cleared for the operator's stop and Resume of one mound (by the caller's id), for
the colony-wide stop file at the first read that finds it present or gone (by the system, in the installation's
organization, with ambiguity: true when the stop is the fail-closed answer of a workspace the colony cannot
read, and recorded again with ambiguity: false, naming that one, when a later read finds the file itself
there), and for a stop the device took itself (by the device); and device.safe_state.entered at the first beat
that reports a stop or a lapsed lease. The operator's reason is never recorded. What a beat carries — its
evidence, its action records and the records of verdicts the colony lowered, its mission reports, the records of
what it reported — is kept in one unit with its acknowledgement, or none of it is and the device sends it again;
a batch the colony cannot keep while a stop is in force is answered with the stop order alone, with no
acknowledgement, so a failing trail never keeps a stop from a device. A stop whose record cannot be
written is engaged all the same, and its answer says so (audit_unrecorded); a Resume whose record cannot be
written is refused (500, audit_unrecorded), and the stop stands. The stop and Resume answers carry
audit_event_id. One change an operator can meet: a workspace the colony cannot read now engages the
colony-wide stop, as the stop file's documentation always said it did; until this release the check used a call
that answers "no file" for a path it cannot read, so the fail-closed branch was never reached. What a device
lost is kept as well: device.chain.broken when its uplink chain does not continue from the digest the colony
last acknowledged — the digest expected and the one received, the refused batch's range — written once per
break however often the device re-sends it (every refusal is still on the bus and in the beat history), with the
refusals counted on the mound's record and device.chain.resumed carrying the count when the chain continues;
and device.evidence.spilled and .evicted from an evidence bundle's own loss counts, which the colony now
reads (a zero count writes nothing, and the bundle is named by its envelope's sequence number, never its id).
Two columns are added to micromound_mounds when the module opens its store (chain_break_digest,
chain_break_refusals); a store from before them has recorded no break, which is true. A mound placed nowhere is the installation
organization's, as every access check on a device reads it. Upgrading adds the two tables and one column
(micromound_mounds.charter_expiry_recorded) when the module opens its store, and back-fills nothing: the
module's tables hold the state a change left behind, never who acted or when. A mint that is refused no
longer leaves a placement behind it either: the mound is placed before its token is minted, so the token's
record is filed where the token binds the device, and a refused mint puts the placement back.
GET /audit/platform — the permission read_audit_platform, which only the administrator role holds;
coordinators, infrastructure operators, the static token and every service identity are refused (403),
and it cannot be delegated. It reads the installation's platform streams, newest first: type (an event
type such as operator_shell.attempted), since and until (times), and limit (1 to 1,000, 100 by
default). It answers records — every field of each sealed record, with its stream_seq, prev_digest
and digest — and legacy_shell_events: the operator_shell_command and operator_shell_error rows an
older build wrote into the colony's log, as it wrote them, when the filter names no type or a shell type.
It also answers anchored: "none": nothing is countersigned until a later slice. Every read is itself
recorded, as audit.read.query, naming the caller, the filter and how many records and legacy rows it
returned, and never a record; a read that cannot be recorded is not answered (500,
audit_read_unrecorded). A filter it cannot read is 400 (invalid_filter); a colony with no audit store
open is 503 (audit_store_unavailable).
GET /audit/status — read_status. How far each stream has verified and whether it is broken, each
outbox's generation, how many rows wait above its mark and how old the oldest is, what the sequencer
last did, and the trail's health; never what a record holds. The streams it lists are the caller's
(E-12 slice 2): the platform's when the caller holds the administrator role, and the streams of the
organization the request acts in; another organization's streams are not listed, and the health line
counts broken streams without naming one (broken_streams names the caller's own). With sign-in off the
one local operator, who reads everything, is listed every stream, the platform's and every
organization's. The route is reached from the installation's organization only, so the organization it
acts in is the installation's, whoever asks. With sign-in on, an owner there is listed what any other
member there without the administrator role is (the administrator role, not ownership, adds the
platform's streams), and another organization's streams, where its devices' records are filed, are on
no status yet; a break in one is counted in the health line, and named only in the self-test's details.
anthill --audit-verify [--stream <id>] walks every stream — or the one named — from its first record,
checking that every digest recomputes, every link holds and no sequence number is missing, and prints three
lines for each stream: what verified, what is anchored (none for now), and the tail that is locally
attested only. A missing record is reported as a gap, apart from an edit. It only reads: it takes no lease,
so it runs beside a running colony. Exit 0 when every stream verifies, 1 on a break, 2 for a stream the
store does not hold. The running colony walks each stream from where it last verified a minute after it
starts and every hour after that, and from its first record once a day. A break is never repaired: the colony appends audit.chain.discontinuity to the broken stream, the next record chains from
it, and the self-test's audit_trail check fails from then on. That check also fails while rows have
waited in the outbox for 60 seconds or more, and while the outbox's retention step and its trigger
disagree.
The outbox ages out. A row the sequencer has taken is deleted audit_outbox_retention_days after it
was taken: 7 by default and never fewer, a floor the outbox's own trigger holds; 0 or less, and 1 to 6, are
read as 7 (W1-10's E1, decision F59). The setting is in config.json only. A row the sequencer has not
taken is never deleted. The step runs with the record sweep (above), and its count is
maintenance_records_expired's audit_outbox_deleted and Compact & prune's audit_outbox_deleted, beside
the sweep's own counts: its deleted list and deleted_count are the record sweep's stores, as before.
The device module's outbox (micromound_audit_outbox) ages out the same way, at the same setting, with the
device evidence sweep (every six hours, three minutes after the module starts): its count is said on the
console and, when that sweep also deleted evidence, in micromound_evidence_expired's
audit_outbox_deleted. A step the outbox's own trigger refuses is counted in audit_outbox_refused and
said on the console as a fault; the self-test reads only the colony's. A store split by
anthill --migrate-stores carries the device module's outbox and its head into micromound.db with the
module's tables, and the module's next open puts back the triggers the split does not copy.
Not yet. FORAGER's outbox, sign-in, account and membership records, the refusal of high-authority work while the outbox is behind (F34), reading and exporting the trail in the console, per-class retention of the audit store, and checkpoints countersigned by the hosted service are later slices of E-12. Until the last of those, the whole of every stream is locally attested only.
The store split
Until v0.4.2.0 one file, anthill.db, held three data domains: the colony's own tables, the
Infrastructure module's 24 and the Micromound module's 12. R-03 §13.2 decided one file per module, and
F12 that a table-name prefix is not isolation for the module that can act on real hosts. So:
| File | Holds |
|---|---|
anthill.db |
the colony — missions, projects, the tenancy spine, and device_placements: where each device belongs, the host's own record since v0.4.2.0 |
infrastructure.db |
the Infrastructure module's tables, its credentials and allowlist included |
micromound.db |
the Micromound module's tables, the controller identity included |
Each file sits beside anthill.db, is leased (<file>.owner) and hardened like it, and is backed up
and restored with it.
A new installation starts split: each module creates its tables in its own file. An installation
from before v0.4.2.0 keeps running as it was — each module on its tables inside anthill.db, and the
console log saying so at every start — until an operator runs the split (the owner's decision, and
F17's posture for anything that moves a customer's data: a person is present).
anthill --migrate-stores [--rehearse] [--backup-out <dir>] [--include-key]
With the colony stopped, in order, stopping at the first thing that is wrong:
- The colony's store is opened once under its lease, so its schema is this build's. Migration
0004_device_placementshas already copied every device's organization and project out of the device module's table intoanthill.db, so no access check depends on a file the host does not own. - What would move is printed. Nothing to move ends here (exit 0).
- Every store's lease is taken, and a backup of every store is made and rehearsed exactly as
anthill --backupmakes one. - The split runs against a scratch copy of that backup. A failed rehearsal ends here; so does
--rehearse, having changed nothing but the backup directory. - The split runs against the colony. For each module, through ONE connection that holds
anthill.dbin SQLite's EXCLUSIVE mode with the module's file attached: every table is copied in one transaction — its ownCREATEstatement, every row, its indexes and itsAUTOINCREMENTcounter — then every table's row count and a SHA-256 of every row's typed values are compared between the two files, the module's file must passintegrity_check, and only then are the tables dropped fromanthill.db, in a second transaction that records what moved inanthill_meta(store_split.<module>). The movedmicromound_moundsloses itsorg_idandproject_idcolumns on the way, after the split checksdevice_placementscarries every placed device. anthill.dbis compacted (E-8): dropping a table leaves its pages free inside the file, so without this the colony's store stayed the size it was before the split.VACUUM, then a truncating checkpoint so the rewritten pages leave the write-ahead log too; the output gives the store's size on disk before and after. Still offline and under every lease. A compaction that cannot run is reported and changes nothing — the split stands, and Diagnostics → Compact & prune gives the pages back later.
A split that stops between its copy and its removal leaves a module's tables in both files. The
colony then refuses to start (exit 13), naming the module and the command, and running
anthill --migrate-stores again compares the two copies and finishes only when they are identical. If
they differ, it refuses and changes nothing; restore the backup it took, which is named in its output.
Exit codes are the backup verbs': 0 done or nothing to do; 1 refused, with nothing moved; 11 a lease is held; 12 something that takes no lease has a store open.
What the inventory will not do
- create, modify or delete anything in the installation — not a config file, not a lease file, not
a store's
-walor-shm; - open any store through
SqliteMemory,AnthillRuntime,SqliteMoundStoreorInfrastructureRepository; - decrypt anything, or print secret material;
- claim a migration is safe.
The exit code is 0 for "the inventory was taken", not for "this installation is fine". A record
whose job is to state what it could not determine cannot carry a verdict, and an exit code that meant
one would be the claim this command refuses to make.
Tests
tests/Anthill.Tests/InstallationInventoryTests.cs builds a fixture installation on disk — a SQLite
file made directly through the provider, files in a temp tree — and asserts the record's fields, that
a missing config file is reported and still missing afterwards, the unfiled counts, manifest drift
in both directions, the --out refusal, and that no secret value the fixture planted appears
anywhere in the JSON or the summary. The fixture is built by hand rather than through the colony's
own types for the same reason the command is: a fixture that had already been migrated could not tell
whether the command under test had migrated it.
tests/Anthill.Tests/InstallationInventorySourceTests.cs holds the source guards — the four types,
the open through ForeignStore, the absence of any writing statement, the absence of any
sealed-column selection, and the verb's position above AnthillRuntime.Initialize().
v2 adds a fixture in WAL mode, stopped cleanly, and compares the whole installation file by file
before and after (AStoppedWalInstallation_IsLeftExactlyAsItWasFound); the lease's state for an
absent and a held lease; and the FORAGER migration count against FORAGER's own files.
tests/Anthill.Tests/StoreSplitTests.cs holds the split: a new colony opening each module in its own
file with none of their tables in anthill.db (F12's validation, as written); a legacy colony split with
every row's count and content digest equal before and after, and the modules reading their data back; a
split refused while another connection has the store open; a split stopped after its copy — both
copies present, the colony refusing to open either, the next run finishing it; an interrupted split
whose copies differ refused with both kept; a device whose placement was not carried stopping the
device split; AUTOINCREMENT counters travelling with their tables; the rehearsal on a backup moving
nothing live; a backup of a split colony holding and restoring every store; a pre-split backup setting
the module files aside; a v1 backup still restored; the inventory's module stores; and the source
guards — the host resolving and leasing each module store before composing any module, no colony code
naming a module's table outside the five files that describe, move or count them, and the verb's steps
in order.
tests/Anthill.Tests/StoreLeaseTests.cs holds the lease: a second acquirer refused and told whose
store it is; a refused acquirer leaving the store byte-identical; release stamping the record and never
deleting the file; a holder in another process named by its own pid and lease id, then killed —
never released — and the lease free at once; one lease through a link to the store's file; the lease
file owner-only; and source guards that every production path opens the colony's store through
SqliteMemory.OpenLeased, which takes the lease before the store is constructed.
tests/Anthill.Tests/ColonyBackupTests.cs holds backup and restore, on WAL fixtures built with plain
SQLite connections: the mission backup carrying rows still in the WAL where a File.Copy of the file
has none; a backup checked, rehearsed and one file; the key left out unless asked for, never created,
and only its check value recorded; the EXCLUSIVE probe refusing an idle second connection and holding
out readers while it runs; a restore that keeps the previous store; the deliberately failed
restore — a backup damaged in the middle of a page and its manifest re-hashed to match, refused by
SQLite's own check with the live store byte-identical; refusals on hash, counts, usability, a
different key and a missing key; the colony's own store round-tripped; and source guards that both
verbs take the lease before touching the store and a restore checks the backup before taking the
lease.
The audit trail's tests are AuditOutboxTests (the outbox's triggers against every write F35 refuses,
and its retention step from both sides of the week), AuditRegistryTests (the colony's typed records held
to the contract's registry both ways), UserIdTests, AuditStoreTests, AuditSequencerTests (exactly
once across a crash on either side of the store's commit), AuditVerifierTests, AuditRestoreTests (a
restore drains first, refuses with nothing replaced, and the next record is sequenced once) and
OperatorShellAuditTests (the attempt committed before the command runs, the completion naming it, no
exception text anywhere, and no command line on any events path), with cases in ColonyBackupTests,
StoreMaintenanceTests and TenancySpineTests.
