STATE & ROADMAP · v0.4.2.14
PLAN — measured current state
Source: modules/anthill/docs/PLAN.md, a path from the root of the product's repository — versioned with the code, rendered here at every release.
The single forward document, and the ONLY one. What is done, what is left, and the order to do
it in. AUTONOMY-10.md folded into this file; role mechanics live in
ANT_EXECUTION.md; the qualification protocol and its measured evidence live in
QUALIFICATION.md; the permanent architectural contract lives in
ADR-008.
Document authority — one responsibility each, so no two documents answer the same question:
| Document | Answers | Never |
|---|---|---|
README.md |
what Anthill does today, for a user | roadmap, or status beyond the capability table |
docs/PLAN.md |
the forward roadmap, release gates, measured status | historical release narrative |
docs/HANDOFF.md |
current session/operational handoff | the roadmap (points here) |
docs/QUALIFICATION.md |
measured qualification evidence and open gaps | forward plans |
CHANGELOG.md |
what each tagged release said and shipped — immutable | current state |
docs/adr/ |
durable architectural decisions | release status |
docs/archive/** |
historical snapshots | anything presented as current |
Shipping release: v0.4.2.14. Everything below the rule is a record, not a state.
v0.3.8.97 correction (recorded here, not by rewriting history). v0.3.8.97 is tagged and
released at a828dfe. Its own CHANGELOG entry says the tag waits for the live qualification pack;
that sentence was true when written and is not edited, because a tagged entry records what a
release said and shipped. What happened after: the pack was blocked on operator switches having no
UI control, and the operator authorized the tag on .97's own evidence. The unfinished pack items
did not lapse — they were .98 exit-gate items and, still unmet, are carried forward in §2c
rather than deleted with the shipped release's row. One .97 residual is carried openly and has
never been a gate for either release: the Windows materialized-revision dotnet_test failure, which
is a coding-lane defect and must not pull a universal-workflow release back into a coding one. It
blocks the current release only if it breaks its suite or its acceptance path.
1. Where the colony measurably is
The CODING lane is deeply built and substantially qualified. The universal workflow is not, and
begins at v0.3.8.98. Both halves of that sentence are load-bearing, and the second half is new
language: until v0.3.8.97 this document opened with "structurally complete", which was true of the
structures and false of the workflow. Measured at v0.3.8.97, only the coding mission class has a
complete execution and verification lifecycle — ObjectiveVerification recognises exactly one
deliverable kind (FileChange), worker specialization is chosen by substring match, and the
assembled answer is never compared to what the operator asked.
What v0.3.8.98 changed, and how far. ONE non-code class — system_audit — now has a complete
lifecycle: classified at intake into declared dimensions, routed by declared worker capability
rather than by substring, inspecting the repository AND the live colony state with inspection
evidence to show for it, and graded against a deliverable ledger that refuses an audit which
inspected nothing, a verifier that read nothing, or a requested deliverable nothing produced. That
is DETERMINISTICALLY qualified and NOT live-qualified: no audit has run against a real provider.
What v0.3.8.99 changes, and how far. A research answer's claims are now TRACEABLE: every url an answer cites is resolved against what the mission actually retrieved — the world's sources and the colony's own recalled missions alike — and a citation that resolves to nothing refuses the mission by name. Claims the mission could not attribute are marked and counted, never dropped and never fatal, and the marking is rendered from the claim record rather than trusted to survive synthesis. What is NOT claimed: that a source SUPPORTS the claim it was cited for (semantic, and deliberately out of reach), retrieval TIME as part of the mapping, any routing of research by declared capability, and anything live.
What v0.3.8.100 changes, and how far. A created deliverable now EXISTS AS A RECORD or the
mission does not pass: a plan that types work as document_creation or data_analysis obliges the
builder to produce a created_artifact whose content is bytes rather than a description, whose
stated requirements each trace into that content or stand visibly unmet, and whose claimed inputs
resolve — id, schema, content hash, stamped from the store's rows and never from the model's text —
to records the mission actually holds. A data analysis additionally records what it read and what
it did to it, or it is refused as a conclusion wearing an analysis's clothes. Unmet requirements
are counted and never fatal, for .99's reason: punishing the admission teaches deletion. What is
NOT claimed: that the content is GOOD or that a traced section truly SATISFIES its requirement
(semantic — the same line .99 drew), file transformation as a distinct lane (a transformation
that touches real files is still the coder's patch lane), and anything live.
What v0.3.8.101 changes, and how far. A reported symptom now reaches a diagnosis that RESTS ON
EXECUTION or the mission does not pass: intake classifies Diagnose-intent requests into the
troubleshooting class under ExecuteChecks authority — the first class to carry it — the planner
ensures a reproduction step, the tester's checks leave command_check receipts with exit statuses
at the dispatch chokepoint, the medic stamps those receipts into its diagnosis from the failed
task's own recorded evidence, and DiagnosisIntegrity refuses a mission whose diagnosis cites a
receipt nothing ran, whose checks never ran, or whose symptom nothing diagnosed. The
audit/diagnosis boundary is enforced from both sides, and a reproduced symptom grades as the
success it is. What is NOT claimed: that the root cause is CORRECT (semantic — the standing line),
"could not reproduce" as a first-class positive answer, symptom-directed check selection beyond
the workspace's allowlisted catalog, and anything live.
What v0.3.8.102 changed, and how far. An infrastructure change now happens as a RECORDED,
REVERSIBLE OPERATION or the mission does not pass: intake derives the system-action class from
Change intent plus a Service target — the first class under Modify authority, and the Service
dimension's first resolver since .98 declared it — the planner ensures the operation step, and
the TESTER's new operation lane (twelve roles is a load-bearing constant; twenty-seven guards
refused the thirteenth-role draft, and they were right about the design — §2c tells the story)
reaches the infrastructure's own approval pipeline through two SDK-named spine tools. Propose captures
the before-state from the runner's dry-run behind a mandatory rollback note; execute sits in the
escalation gate's side-effecting set, so the operator's conversation-scoped decision is demanded
at the dispatch chokepoint and stamped into the record as the approver — the lane's identity,
never the model's. OperationIntegrity refuses each absent piece by name. What is NOT claimed:
general local-system operations outside the infrastructure catalog, multi-operation missions, typed
state snapshots per runner, and anything live. Five of the seven classes are now served; external
missions are untouched, and everything unserved still resolves not_applicable at the
deliverable layer exactly as before.
What v0.3.8.103 changes, and how far. Something can now LEAVE the colony, and only to a
destination a human approved: the alias is resolved to a concrete target before approval is
offered, the resolution is recorded, and what the adapter reports it actually hit is recorded
beside it — so an approval of one destination and a send to another is refused by name rather
than being invisible with every field populated. A refused send writes its own record, and the
answer is RENDERED from that record ahead of every prose path, because a builder whose tool was
refused upstream still writes "I've posted it to the team". And MissionAuthority, declared
since .98 and read by nothing, is finally read: one table from a side-effecting action to the
authority a mission must hold, swept so no member of the escalation set can omit an entry.
What is NOT claimed: a send that reached anything. The adapter ships composed and the destination map ships EMPTY, so a fresh install refuses every send by name.
What v0.3.8.108 changes, and how far. A role can be declared without editing the core. One declaration point carries the registry entry, the runtime kind, the execution contract and the executor factory, and all four tables that decide whether a role runs read it. The layers the exit gate names — Queen, planner, scheduler, assembler — were never the obstacle and are untouched; what blocked extension was four static literals, one of them a dictionary inside the Queen's constructor. "Extensible" has been an implicit claim in every capability table since the roster was written, and it is now either true or checked.
What is NOT claimed: a MODULE still cannot contribute an ant. BaseAnt lives in Anthill.Core,
which a module may not reference — exactly where RegisterTool stood before v3.8.10, and the same
answer applies: the type moves to the SDK first, in a release of its own, and it needs this
composability underneath it either way.
What v0.3.8.109 changes, and how far. A question the colony cannot answer from itself is its own
mission class: derived from a new intent and a new outward target, planned with a retrieval step it
cannot omit, and graded on whether it went and looked. A retrieval leaves its own evidence kind — not
inspection, because an audit of the operator's repository must not be satisfiable by searching the
internet — so "did this mission retrieve anything" is answerable from the store rather than inferred
from an artifact the answer might not cite. CitationIntegrity's second trigger, open since .99
and recorded as unbuildable at .104, now fires from the contract as well as the record, which is
what makes a mission that retrieved NOTHING catchable: an empty store contradicts no citation.
What v0.3.8.110 changes, and how far. An approved decision replays the step that was refused. A mission that stopped for a side-effecting action it had no answer for is finished by answering it, not by running the whole mission again — which needed a typed mission loader, the first in this tree, and needed the approval ledger and the mission-lane gate to stop being two disjoint tables that could not see each other's answers.
What is NOT claimed: that any mission can be resumed. Only the tasks a mission's own refusal events name for the approved action are replayed, and a task that COMPLETED is never touched — its effects have already landed. A rejection replays nothing, which is the point of asking.
What v0.3.8.128 changes, and how far. The escalation policy an operator set on a conversation
now governs that conversation's MISSIONS. .102 recorded the reason it did not — "a mission does
not run inside the ambient ConversationScope, so this branch is unreachable from one" — and that
sentence describes a gate at the colony's single tool chokepoint which was silent for every dispatch
a mission ever made. .102–.110 closed it by hand for the two execute tools and the API's action
path, which covered the loudest actions and left one rule living at three call sites; apply_patch,
write_text_file, shell_command and run_allowlisted_check crossed the boundary ungated. The
chokepoint reads the durable record now, after the ambient one and only when the ambient one had
nothing to say, so a live answer still beats a stored one. A standing permission allows the action
AND logs the decision that permitted it; an absent answer files the question rather than failing the
mission.
What .128 did NOT claim, and what v0.3.8.130 then did. .128 left a mission WITHOUT a
conversation ungated and argued for it: such a mission "has no operator policy to consult", and
manufacturing Ask for it "would refuse every patch the coding lane has ever written on the grounds
that a conversation nobody started did not answer a question nobody asked." The mechanism in that
sentence was right and the conclusion was not — the objection was never to asking, it was to asking
a question nobody could answer, and the distance between the two turned out to be one config key.
autonomy_escalation_policy is the answer given once and in advance, so the scheduled, CLI and
Director lanes are governed by the same chokepoint; ask is the default and every safety profile
pins it there. Neither release widened the escalation SET: it is EscalationGate.SideEffecting,
unchanged, read rather than re-decided.
What is NOT claimed: that the sources are any good, or that they support what they are cited for.
Those are semantic judgments, and a model asserting one is the evidence v2.19.0 stopped accepting.
Traceability is checkable; support is not. Nor is a request naming both worlds admitted — its answer
rests half on an inspection this gate cannot speak for. Nothing live, still. See
ADR-008 §1 for the evidence and §2 for the contract
the .98–.107 sequence exists to satisfy.
The colony has been demonstrated live end-to-end as a coding colony (mission 3bbbde32,
completed_verified). It has not been demonstrated as a general-purpose one, and this document
will not describe it as one until §2b's exit gates are met.
Done and load-bearing:
- Twelve roles, contracted, gated, each with a real production trigger.
- Patch integrity:
addmeans create, destructive applies require a base hash, a patch set applies as a unit or not at all, and one decision function (PatchApply.Compute) answers for every applier. - The typed artifact channel: declared task inputs, a consumption ledger recording what each role actually read and at which hash, schema validation at both boundaries, provenance carrying the provider and model that served each call.
- Structural enforcement: a UI change cannot dispatch without a valid
ui_map; verification is policy-inserted and fails closed; the repair bound reads typed signatures; the scribe cannot certify unverified work;MissionReconstructionreplays a mission from artifact IDs. - Every operator message is a mission — no chat lane, no
conversationroute, no unconfined agent access. The colony dispatches a coding agent as a TOOL inside a mission that plans, reviews, tests and verifies its work (v0.3.8.58). - A goal reaches applied bytes on the operator's tree, through all nine gates (v0.3.8.75).
- The security review of §5 is closed: S1–S7, with residuals recorded there.
Counted honestly
| Deterministic qualification scenarios | 20 of 20 closed by substance (v0.3.8.79). None open, none partial |
| Acceptance gates | 12 of 12 (v0.3.8.80). Gates 1 and 2 closed with R3's cancellation matrix, which is what made the roster's declarations true enough to grade |
| Cancellation cells | 48 of 48 decided; 33 driven live, 0 cited, 15 not-applicable (v0.3.8.88 — R3's exit gate MET). Two cells moved from not-applicable to DRIVEN by looking at the role's own trigger instead of the planner's: the archivist's at v0.3.8.85 (nothing had ever stopped it at finalization) and the medic's at v0.3.8.88 (from the tester→medic handoff). The 15 that remain are points the runtime cannot produce, each with the reason recorded |
| Graduation record | Complete (v0.3.8.81). The cancellation column and the ui_cartographer fault cell were the last nulls |
| Live qualification | PARTIAL — coding lane only. Real Claude Code acting missions have run and passed (3bbbde32, completed_verified, 41.6s, plus six failing runs whose findings shipped in .95/.96). None captured the §3 telemetry table, ran with objective verification enabled, or exported a LiveQualificationRecord; no non-code class has been run live at all. QUALIFICATION.md §3 is the authority and says PARTIAL, not NEVER RUN |
| Model-calling roles | 5 roles and 8 routes, all declared, both directions asserted (v0.3.8.76). Was "7 declared, 5 of which cannot call a model" — and separately, 3 routes that do call one were declared nowhere |
| Structured output | Asked for on the wire by coder, planner, strategist (v0.3.8.76). The field existed, was plumbed and was gated since v3.4.0; no producer had ever set it |
No scenario is now weaker than the ledger says, which is a sentence this document has not been able to write before. Scenario 5's note used to admit "the composed UI-patch lifecycle is not [proved]" while the entry was not labelled partial, so the guard built for exactly that could not see it; scenario 17 was labelled honestly and stayed labelled for eleven releases. Both closed on substance rather than on wording — 5 at v0.3.8.78, 17 at v0.3.8.79.
Capability status — what is true per capability, not per component
Statuses are exact and mean only what they say. Implemented = the code exists and is on the production path. Default on = enabled without operator action. Det. qualified = proved through the real composition by the deterministic suite. Live = proved with a real provider through the real application. Ext. = requires an external adapter, connection or authority. "Supported" is not a status here, because it has been used to mean all five.
| Capability | Implemented | Default on | Det. qualified | Live | Notes |
|---|---|---|---|---|---|
| Coding: worktree execution | yes | no (acting_coder_enabled) |
yes | yes | the qualified lane |
| Coding: patch promotion and apply | yes | gates off by default | yes | yes | target identity, atomic sets (.97) |
| Repository inspection | yes | yes | yes | no | audit missions resolve repo_researcher + file inspection by declared capability and leave inspection evidence (.98); live run pending |
| Runtime/state inspection (read-only) | yes | yes | yes | no | colony_state + researcher.runtime_researcher (.98); live run pending |
| Web research: claim→source traceability | yes | yes | yes | no | every cited url resolved against what was retrieved, or the mission is refused (.99); unsourced claims marked, not dropped |
| Web research: retrieval | yes | no | yes | no | Ext. — needs a search provider. .109: a retrieval now leaves a source_retrieval evidence row, so "did this mission go and look" is answerable from the store rather than inferred from an artifact the answer might not cite |
| Internal-memory research | yes | yes | yes | no | recall leaves a recall_set, so mission:<id> is citable and held to the same standard as a url (.99) |
| Claim↔source SUPPORT | no | — | no | no | deliberately absent — semantic, the .99 line |
| Artifact/document creation | yes | yes | yes | no | a creation-typed task must leave a created_artifact whose content exists, requirements trace or stand unmet, inputs resolve (.100) |
| Data analysis | yes | yes | yes | no | input identity (id + content hash from the store) and a transformation account, or the mission is refused (.100) |
| Troubleshooting / diagnosis | yes | yes | yes | no | a symptom reproduced by executed checks, diagnosed with receipts cited by name, boundary enforced both ways (.101) |
| Local system actions | partial | no | partial | no | infrastructure-catalog operations shipped (.102); a general local-system lane outside the catalog is not claimed |
| Infrastructure/infrastructure actions | yes | no | yes | no | on the mission spine: propose → operator decision → execute → verify, recorded as a reversible operation (.102) |
| External actions (approval-gated) | yes | no destinations configured | yes | no | Ext. — the class, record, ceiling, gate AND a ConfiguredWebhookAdapter composed by the API host. external_destinations is empty by default and IS the allowlist, so a fresh install resolves nothing and refuses by name |
| Mission authority ceiling | yes | yes | yes | no | .104: read at the dispatch chokepoint, from the mission's recorded contract, for recognized classes only — General defaults to Observe and means "unclassified", not "read-only" |
| Mission contract (persisted) | yes | yes | yes | no | .104: written once at intake, read by every stage; an intake rule change cannot reclassify a mission that already ran |
| Mission preflight | yes | yes | yes | no | .104: producer, verifier, worker, dependency and orphan checks before execution; runs after class coverage so it refuses only what the runtime cannot repair |
| Live qualification export | yes | yes | yes | no | .104: anthill --live-qualification — the record type shipped at .89 and had no production caller until now |
| Dispatch-time reroute | yes | yes | yes | no | .105: a task whose worker does not declare its required capability is rerouted within the role before the durable claim, or refused as capability_unserved; covers the tasks admitted after preflight ran |
| Operator-decision pause and resume | yes | yes | yes | no | .105: an unanswered side-effecting action files a pending ToolUse approval and the mission grades waiting_for_approval instead of failed. .110: approving now REPLAYS the refused step — the mission is rehydrated, only the tasks its own refusal events name are reset, and it is re-executed and re-graded. A completed task is never replayed; a rejection replays nothing |
| Dynamic repair (Medic) | partial | yes | partial | no | bounded repair exists; at its bound the stop now NAMES a reproducible failure instead of only the spent budget (.105). The loop's generations are unchanged — a recurrence explains a stop and never causes one. Still not evidence-driven |
| Recovery decisions | yes | yes | yes | no | .105: RecoveryOrchestrator consults the FailureClass taxonomy — a denial is never retried, an unclassified failure escalates for evidence, and a typed class narrows a caller's optimism without ever widening it |
| Multi-mission continuity | yes | yes | yes | no | .106: read_artifact reads an EARLIER mission's artifact by id, refused unless that mission graded completed_verified; the consumption ledger records who read it as distinct from who produced it. READ-side only — a paused mission does not resume its refused step |
| Answer coverage | yes | yes | yes | no | .106: the answer is assembled section by section from the specification, and an unanswered request demotes the mission. Claim-and-served, never a word search — MissionDeliverable.Subject stays unread by design |
| Pheromone / skill learning | yes | yes | yes | no | positive learning restricted to completed_verified |
| Roster extensibility | yes | yes | yes | no | .108: a role is declared once — registry entry, runtime kind, execution contract, executor factory — and every table that decides whether it runs reads that declaration. A contribution cannot shadow a built-in. MODULE-contributed ants are not claimed: BaseAnt is core, and the SDK move is its own release |
| Verified route learning | yes | yes | yes | no | .107: a route that carried missions to completed_verified is preferred for a role the operator has NOT routed. Never overrides an explicit route, a priority override or compatibility; reads verified_route trails only, never the per-call model_route ones |
| Research (sourced answers) | yes | yes | yes | no | .109: a question the colony cannot answer from itself is its own class — retrieval step planned deterministically, source_retrieval evidence required, every claim resolved against what was fetched, and each requested section traced to a retrieval. A mission that retrieved nothing is refused, which the .99 gate could not see |
| Objective verification (non-code) | yes | yes — recognized classes | yes | no | .104: a recognized class (system_audit, troubleshooting, system_action, external_action, research) no longer consults objective_verification_enabled and FAILS CLOSED when its gate cannot run. The flag still governs the general and coding lanes. Six releases of gates were inert on a default install until this |
No row may claim a status stronger than QUALIFICATION.md records. That is checked, not trusted:
see DocumentationConsistencyTests.
2b. The universal-workflow program — v0.3.8.98 → v0.3.8.113 · ✅ CLOSED at v0.3.8.113
This is the current forward sequence. It supersedes the earlier framing in which R4–R10 were
the next thing to run; those items are not deleted, and each release below names the R-items it
advances. The program's permanent contract is ADR-008;
this section owns only when and what proves it.
The organizing decision, taken after external review: vertical slices, not horizontal layers. An earlier draft of this program spent its first five releases building a specification, a capability registry, a plan compiler, a typed envelope set and an adapter boundary before anything an operator could observe changed — which is the exact failure mode this repository has shipped before. Each release below instead makes one mission class work end to end on the shared spine, and introduces only the universal structure that class actually requires.
v0.3.8.97 is not release 1 of this program. It was an urgent prerequisite: it made the coding
promotion path correct (project identity through the whole transaction, set-atomic apply, faithful
capture) and made two operator switches reachable. It strengthened the lane that was already the
strongest. The program started at .98.
The table below is the REMAINING sequence, and it shrinks as releases ship. A shipped release
leaves it and its story passes to CHANGELOG.md, which owns what each release said — this document
does not keep a release narrative. What does NOT leave with the row is anything the release did not
finish: unmet items are carried explicitly into the current slice in §2c, so a row can only be
deleted when its work is either done or written down somewhere still current. The heading names the
range and DocumentationConsistencyTests checks that the rows are exactly that range beginning at
the shipping version — a program whose first entry has already shipped is a plan describing the
past.
THE TABLE IS EMPTY, AND THAT IS WHAT FINISHED LOOKS LIKE. .113 was the last row; it shipped,
and by this section's own rule it left. Five mission classes — system_audit, troubleshooting,
system_action, external_action, research — work end to end on the shared spine, each with an
integrity gate that runs whatever the operator switch says. What the program did not finish is
carried in §2c under the release that inherited it, not left here as a row nobody is working on.
Rules that bound every release in this program, kept because they still bind the work that inherited its unfinished items:
- Extend the existing consumption ledger, artifact store, evidence store, qualification matrix and outcome vocabulary. No parallel ledger, no second matrix, no duplicate schema.
- Each release adds at least one composed positive and one composed negative acceptance scenario through the real composition root.
- A release is not complete while its exit gate is unproved, and unfinished work is not converted into a "documented limitation" to close it.
- Any new operator switch must reach the editable settings surface, persist, and appear in the
settings snapshot in the same release — the
.97lesson:objective_verification_enabledbecame API-editable with no control rendering it, so an operator following the changelog looked for a switch that did not exist.
2c. v0.3.8.114 — the colony can direct a mound, and R0 closes
Delivers: R0's last item, and the MICROMOUND controller — the half of the link .60 deliberately
did not build.
R0 CLOSES. The generated configuration schema is above, under R0 itself. .112's entry called
its own work "R0's LAST ITEM"; that was wrong when written and two releases shipped past it. The
correction is recorded, never by editing the tagged entry.
THE COMMAND PATH EXISTS. .60 shipped the uplink and said what it had not built: "M1 has no
command path, so the colony can see mounds and cannot direct them." This release builds the other
half — a signing identity, charter issuance, configuration authoring, physical mission dispatch,
structured evidence, and the capability resolver — and every one of them signs an envelope into a
downlink queue, because the colony never dials a mound.
AND THE BEAT NOW ANSWERS, which is what made any of it work. Three protocol obligations were unmet, and each breaks a fleet on its own:
- The ack. PROTOCOL.md §6's retention rule is written in terms of one message: until an ack covers a sequence number, the device's uplink queue must retain the envelope and its evidence store must retain the proof. We sent none, so no mound could ever release anything.
- The lease. §5 — an acknowledged
mound_syncrenews it, and nothing on-device can. The device renews when it sees an ack covering its beat's sequence. A colony that sends no ack does not merely fail to renew: every chartered mound runs its lease down and enterssafe_stateon schedule, beating perfectly, reported online, silently refusing every actuation from then on. - The downlink. §1 — the response carries any pending downlink. A charter signed and queued is not delivered by being queued.
THE DEVICE-FACING WIRE SHAPE WAS INVENTED RATHER THAN READ. M1's two endpoints were written from
the protocol document instead of from the client that calls them, and a real MicroMound could not
have used either: /v0/enroll required a mound_id the device does not send, read the key from
public_key where the device writes device_public_key, and returned no controller_public_key —
so even a device that got past the first two could never verify a downlink envelope. /v0/sync
expected {mound_id, envelopes[]} and answered an object, where the device POSTs one raw envelope
and parses the whole response body as List<Envelope>. Nothing caught it because both ends of every
test were ours.
THE UI IS NOT IN THIS RELEASE, AND §31 IS NOT MET. The integration brief's acceptance experience — an operator adding a mound, authoring its configuration, issuing a charter and watching evidence arrive, all from the console — is not delivered and must not be described as delivered. The operator instruction that set this scope was explicit ("lets leave ui/frontend work out of this release, just everything else"). Everything here is reachable over the API and nothing here renders. Recording that is §37's requirement, and the reason it is written in the plan rather than only in a commit message is that a deferral nobody wrote down becomes a capability somebody assumes.
Exit gate — named tests, and the first is the one that matters:
SimulatedPeerTests.TheLeaseSurvivesFarPastItsTtl_BecauseEveryBeatIsAcknowledged— an hour of beats against a fifteen-minute lease, asserted from the DEVICE's side, with a positive control proving the lease does lapse when the beats stop;SimulatedPeerTests.ADeviceIsEnrolled_Chartered_AndCarriesOutWorkTheColonySent— enrol, charter, dispatch, execute, report, through the real device runtime and a JSON round trip;DeviceWireContractTests— the enrolment and sync shapes read out of the pinned checkout's own HTTP clients;ConfigCatalogTests— byte-for-byte regeneration of the example file and the reference doc;StorePersistenceTests.EveryPerMoundTable_IsSweptOnRemoveMound— derived from the schema, not from a hand-maintained list.
WHAT .114 DID NOT DO, said plainly. This release is large and none of it is hygiene: the
standing R0 items below are all untouched, and the typed-row ratchet did not move. That is a missed
beat rather than a rule broken — TheUntypedStoreSurface_OnlyShrinks requires the count to be ≤
45, not to fall every release — and it is recorded because a release that quietly skips its slice
twice is how "one slice per release" becomes a sentence nobody is keeping.
Carried out of the program, unchanged and still named:
- The remaining 45 untyped store readers — one slice per release, each lowering the ratchet in
the same commit. Still at 45:
.114spent itself on R0's last item and the Micromound controller, and typed nothing. The ratchet is enforced, so the work cannot quietly reverse; only this note stops it quietly stopping. AnalysisModebeyond the SDK default — needs its own census, because with warnings-as-errors on every CA diagnostic becomes a build failure.- Central package management — the real fix for version drift; its failure mode is "nothing
builds". The guard against drift shipped at
.112. - Four of eleven literal-only guards, each lower risk than the seven done and each named at
.112§2c with why. - The S1 filesystem TOCTOU window — needs a handle-returning path guard and P/Invoke; .NET
exposes no portable
openat. - Four of five agent CLIs declare no system-prompt channel — a data change each, gated on confirming the vendor's flag against an installed binary.
- Ollama capability discovery — R1's last item, a contested design decision rather than an oversight.
- The
.97Windowsdotnet_testresidual — needs the machine that reproduces it. - The capability-table reconciliation and
QUALIFICATION.md§3 — need the live pack's exported records, and no code change moves them.
2d. v0.3.8.115 — colony live stops inventing a colony
Delivers: the console half of the last two releases — a Colony Live that draws only what the colony recorded, and a Micromound view that shows what the fleet listing says rather than what it could plausibly say.
THE PREMISE OF THE BRIEF WAS FALSE, and that was worth finding first. The work was specified as
"preserve the existing WebGL renderer". There was no WebGL renderer: .111 shipped a canvas-2D
projection with a hand-rolled 3D transform. The decision taken was to vendor three.js 0.128.0 as an
embedded asset served from this origin — pinned at that version because later releases dropped the
UMD build that defines window.THREE under a plain <script src>, which is what keeps the CSP at
script-src 'self'. The prototype's unpkg-plus-Blob-URL fallback is not ported and a guard refuses
it.
SEVEN THINGS .111 INVENTED. Enumerated in colony-topology.js's header rather than here, and
each replaced by a fact or by an honest absence. The costliest was a hand-maintained role→sector map
resolving a miss to the Queen, which silently mis-filed every role added after it was last edited.
Membership is now the registry's, projected once, server-side.
A GUARD FOUND THE WORST OF IT. The §17 rule against Math.random in the feature failed on the
classic fallback, and following it in found demoTopology() — an invented mission, three named ants
at made-up speeds, a fabricated approval boundary — running unconditionally at startup, plus a
recordAt fabricating records for unfilled particles. That same file's setTopology also still read
five fields the new projection does not emit, so it would have rendered permanently blank while
looking wired up. Both fixed. This is the release's own evidence for writing the guards before
believing the code.
Verified by — 3,571 tests, zero failures, and specifically: ColonyLiveGuardTests (fourteen
rules, each with a vacuity floor), UiAbsenceTests extended to the two new assets and the vendored
bundle, and ConsoleAssetSplitTests covering both new console files without being edited, because it
enumerates the directory.
AND THE MICROMOUND CONSOLE SHIPPED WITH IT. .114 deferred it explicitly and recorded seven
routes in the coverage ledger as UI GAPs so the deferral could be CHECKED rather than asserted. That
block is now empty: src/Anthill.UI/micromound.js lists the fleet with the colony's own status
verdict, mints and retires devices, engages and clears the per-mound stop, issues charters and
manifests, composes and dispatches physical missions, reads one mission's two verdicts, asks the
resolver and shows the evidence feed. Only the two DEVICE endpoints remain in the ledger, because a
mound is not an operator.
Every form field was read off the declaring type — Micromound.Protocol for the wire shapes,
ApiHost.Micromound.cs for the request bodies. ConsoleVocabularyTests lives in
Anthill.Tests.Micromound and compares the console's five closed vocabularies against the protocol's
own sets at the TYPED tier, so adding an operation to the protocol fails a test rather than quietly
leaving the console a version behind. The one deliberate narrowing — ceilings stop at controlled —
is asserted as a narrowing, against the issuer's own refusal of hazardous.
MicromoundWidgets.StatusOf became public and the fleet listing carries its verdict, closing the one
thing .115 had first deferred for a good reason: the console could not show online/offline without
recomputing a rule from configuration a browser cannot see, and carrying the answer is smaller than
duplicating the rule.
WHAT .115 DID NOT DO, said plainly. Colony Live has no pheromone overlay bound to real scores —
nothing in this model carries a per-record score to bind one to. The canvas fallback plays no
transition flights; the WebGL renderer does, and the fallback has no truthful source for an ant in
transit. Growth playback cannot reach events that were never persisted, which is every Micromound
event and four others — a backend gap, named in §2e. And the standing R0 hygiene is untouched for
the second release running.
Carried forward, unchanged and still named: the list under §2c stands as written, with one correction — the typed-row ratchet is still at 45 and has now been skipped twice. The ratchet is enforced at ≤ 45, so the work cannot reverse; only this note stops it quietly stopping, and it is worth more the second time it has to be written.
2k. v0.3.8.121 — what the organization knows
Delivers: the colony can ask what the organization knows and get statements back with the source
text behind them. The knowledge lives in FORAGER — a separate local application that turns documents
into canonical, traceable records — and this release is the seam, not a second copy of it. ANTHILL
parses no documents, stores no knowledge, resolves no conflicts and ranks nothing; it asks, and
presents what comes back without degrading it. The boundary is HTTP because it had to be: FORAGER is
TypeScript on Node and there is no in-process edge to share, so the MICROMOUND ProjectReference
pattern has no analogue here.
Retrieval is evidence-first, which is not what RAG usually means. The ordinary shape — embed,
take the nearest chunks, paste them in — hands a model text with no accountability, judged only by
plausibility. This ranks CANDIDATES, then fetches what supports each one, then attaches the
disagreements. Evidence is fetched before assembly because an item whose evidence cannot be resolved
has to be LABELLED rather than dropped, and you cannot label what you have already flattened.
Conflicts are printed BEFORE the facts, with FORAGER's suggestion marked NOT APPLIED: a model that
meets the statements first has already formed an answer. There is deliberately no option to hide a
conflict, because an option to hide them is a way to hide them.
Scope is ambient, and that is a security decision. ITool.Run receives arguments and nothing
else, so a knowledge tool learns its scope from an argument or from ambient state — and tool
arguments are chosen by a MODEL. A project_id parameter would make the reach of a query something
the model selects, and the no-cross-project rule would then be enforced by its discretion.
KnowledgeScopeContext is entered by the core at intake (SINCE v0.3.8.136 — between .121 and
.135 this sentence was the design and not the tree: only the console resolved a scope, and every
knowledge tool an ant dispatched refused; MissionKnowledgeScopeTests now pins the resolution
rules), only ever narrows, and defaults to a scope
that retrieves nothing. Verified against a running FORAGER: GET /api/knowledge/{id} is NOT
project-scoped upstream and returns another project's row with HTTP 200, so the provider checks
project_id on the RESPONSE and answers NotFound — not a denial, because confirming an id exists
in a project the caller cannot see is itself a disclosure.
And a bound knowledge base is now READ by a step. v0.3.8.156. .136 entered the scope so that
a knowledge tool an ant dispatched would resolve; nothing ever planned a step that dispatched one, so
a project mapped to a FORAGER project still planned builder -> verifier and answered from the
model's weights. Planner.EnsureKnowledgeConsultation guarantees a researcher step that calls
knowledge_retrieve ahead of every synthesis whenever the mission's ambient scope is queryable —
class-independent, reported through PlanSubstitutions.KnowledgeBaseBound, and excluded for
simple_answer alone, whose promise is that the answer rests on nothing retrieved. The condition is
read from the scope the Queen entered, never from the project map a second time: two readers of
"which knowledge may this mission read" is the one failure the scope model exists to prevent.
And what it reads is citable. v0.3.8.157. .156 planned the step; the researcher is a
deterministic handler and never dispatched the tool, so the step reached no chooser. The retrieval is
now dispatched whenever the mission's scope is queryable, and every evidence-backed statement it
returns is recorded in a source_set as knowledge:<project>/<id> — the mission:<id> vocabulary
from .99, so one gate resolves all three kinds. Facts that are UNRESOLVED are deliberately NOT
citable: HasProvenance is true for them by Rule 9, and citing on it would let an answer rest on the
one statement the context says it cannot support.
And the page an operator uses was rebuilt around it. v0.3.8.157. Connect, bind, import — one
card, in that order, with sources, conflicts, review proposals, jobs and the full binding table
folded below. The arguments the page used to carry as prose moved to docs/KNOWLEDGE_ARCHITECTURE.md
§3b, which is where an argument belongs; a console says what is true now and offers the next action.
Off by default, and it adds no tables. knowledge_enabled ships false; an existing config loads
unchanged and the database is untouched in both directions, so enabling and disabling are equally
safe. The tools register unconditionally and refuse at call time — "registered and refusing" is a
different fact from "declared and absent", and only the second makes a role unqualified. That was
learned the hard way: the first implementation registered nothing when the feature was off, and
three guards refused it in one run because a researcher would never have passed readiness.
Also lands: the Mission Replay configuration contract — typed, validated settings for future Obsidian replay, with no parser, no indexing and no execution. And in FORAGER's own repository, two defects found by running the integration: directory import was open by default (the containment check skipped entirely when no roots were configured, on an API with no authentication), and FTS5's AND-joined terms meant a natural-language question returned nothing while the unranked fallback had better recall than the ranked backend.
Not done, on purpose: no vector search (FORAGER's SearchBackend seam is where it belongs and
does not exist there yet), no agent-authored knowledge, no cross-project retrieval — which is not
expressible in the scope type at all. The two applications have not yet been run together end to
end; the retrieval pipeline was verified against a live FORAGER and the C# against the suite, but no
environment in this work had both.
2j. v0.3.8.120 — the colony page, tuned by hand
Delivers: an operator's pass over .119, and one defect it uncovered that had nothing to do with
polish. A chamber's records now take seats on a fixed 96-slot Fibonacci lattice — evenly spread by
construction, stable per record — so a dominant cluster no longer reads as a lopsided clump, and the
chamber glow is an envelope that contains every seat rather than a nucleus its outer grains escaped.
Residents are drawn as halo, core and ring so an ant is never mistaken for a record, and clicking one
opens an inspector that also takes a display name and a colour. Chambers take a colour, a glow size
and a brightness; conduits take particle density, brightness and a colour; labels take All / Focused
only / None; Void is black. All of it persists in the layout the /ui/state schema already carried,
and reset returns the projection's own. Mounds is greyed until the fleet has one and + Mound
opens the console where a device is actually enrolled.
Light mode is a first-class sky rather than the dark one with a white background: the chamber glow
is a stronger tint that falls off late and closes on a rim, labels are the palette darkened further,
and the conduit strands carry roughly twice the alpha — the same weight to the eye. Auto follows
Settings › Theme.
The defect: Colony Live enables at DOMContentLoaded, which on a fresh session is the sign-in
screen. Both bounded reads were refused, nothing retried them, and the operator signed in to stars and
no chambers — a page that looked like a rendering bug and was a lifecycle one. Hydration is now
re-attempted on a trigger and never a clock (page entry, and the first event on the stream, which
connects only after auth), idempotent and non-overlapping, with a refused snapshot no longer counted
as a hydration. ColonyLiveGuardTests.Hydration_IsReAttemptedAfterSignIn_AndIsNeverPolled pins it.
And one that was not about the colony at all. ApiJobRegistry.Dispose() disposed its queue while
worker threads were blocked inside it, which throws on the worker thread — an unhandled background
exception, which is a process kill rather than a logged error. CI showed it as a test host crash
after all 1,652 tests had passed; the same shape would end the API host on shutdown. Dispose now
drains its workers first and the take is guarded, with JobRegistryShutdownTests on it.
2i. v0.3.8.119 — the Colony Live UI re-ported onto the approved read model
Delivers: the .115–.117 WebGL renderer, its HUD and the vendored three.js are removed — the
UI was rejected in review; the backend it was built against was approved and is untouched. The
.111 canvas formicarium (landing page in focus, live bar, composer → Chat, sector/record panels,
the galaxy sky) now consumes the reducer and the endpoints unchanged. A chamber's grains are its
persisted records and its orbs its residents; the chambers are the server's nine; the mound exists
only when the fleet says so; labels are the projection's unless the operator renamed them, with the
layout in /ui/state (schema 3; a schema-2 layout from the retired build migrates by offset, ×10, y
flipped). ColonyLiveGuardTests keeps every rule that protected the
contract and gains the canvas renderer's own; the §18 constant comparison went with the renderer
it compared. The design handoff stays under docs/design/colony-live-3d/ as a record, and says so.
The CHANGELOG's ## v0.3.8.119 entry carries the full account, including what the V&V pass found.
2h. v0.3.8.118 — the mission honours what was asked, or says why it cannot
Delivers: the first two items of the orchestration brief, and the correction .117 shipped
without.
ROOT CAUSE OF THE FIXED-WORKFLOW BEHAVIOUR, and it is not a dispatch bug. MissionRequest
carried { Goal, IdempotencyKey }; a repo-wide search for requested_roles, output_schema or any
equivalent returned zero matches in production or test code. There was no input contract to
ignore. Worse, the goal string was also the trigger for spec ingestion — Planner.cs:128 gates on
goal.Length > 6000 — so the more precisely an operator specified roles, ordering and output
shape, the more certain it became that the whole request would be chunked into section_analysis
tasks. Precision was punished. Full evidence in the project doc orchestration-root-cause.md.
WHAT .118 ADDS. RequestedWorkflow is the missing input contract, and it keeps three things
apart that the runtime had collapsed into one: a LABEL is what the operator called a step and is
descriptive; a TASK TYPE is what some worker contract declares it supports and is executable; a ROLE
is neither. Treating a label as executable is how an arbitrary name reached a worker.
DispatchPlanner resolves or REFUSES before dispatch — unsupported task types, unsupported output
schemas, unknown roles, role/type mismatches, unresolvable labels, dangling dependencies. Labels
match exactly and never fuzzily, because near-matching is how a request becomes something adjacent
to itself without anyone being told.
THE FAILING TEST TAUGHT THE DESIGN, TWICE. The verifier is SchedulingMode.PolicyInserted —
the registry's words are "the steps a plan must not be able to omit" — and the planner first treated
"the planner cannot pick it" as "it is unavailable", which would have refused exactly the missions
this release most wants to succeed: the ones asking to be verified. Requiring a policy-inserted role
is now SATISFIED; authoring a step for one is refused with the alternative named; requiring a
lifecycle-only role (medic, archivist) is refused because nothing can promise it. Routable and
Dispatchable became separate fields for the same reason Registered and Dispatched had to.
AND A GUARD CAUGHT THE AUTHOR. ShippedChangelogTests failed this release because the .117
entry had been edited after its tag — the logo correction was written into a shipped entry instead
of a new one. Third occurrence in this repository, first one caught by a test rather than by a
reader. The correction lives here, where it belongs.
Verified by — 3,615 tests. Nothing supplies a workflow yet (no API field, no CLI flag), so
every mission plans planner_chosen and runs exactly the path it did before; the only visible
change is one event per mission recording that the planner chose.
Carried: items 3–8 of the brief — authoritative execution records, artifact/evidence handoff,
real verification, closure enforcement, unsourced-claim rejection — all need task records that do
not exist yet. Planner.cs:128's character-count gate is still live and moving it waits on those.
The typed-row ratchet is at 45 for a fourth release.
2g. v0.3.8.117 — the colony view stops being opt-in
Delivers: the shell around the renderer .116 built, and the first deliberate divergence from
the design it was ported from.
THE VIEW WAS OPT-IN. Two releases of work sat behind a button an operator had to know existed.
Colony Live 3D is the default now; the canvas projection is the fallback and the single 2D option
(Command, Active and Chambers hidden, Expanded renamed 2D View), and the 3D HUD's control bar
renders into #colony-viewbar instead of floating a second row bottom-right. Four reset buttons
across two bars became one that resets whichever renderer is showing.
THE FIRST NAMED DIVERGENCE. Conduit grain counts drop from the reference's 60/24/40 to 36/15/24 —
eighteen roots against its sixteen, and far sparser chambers, made the streams the loudest thing in
the frame. The interesting part is the mechanism: rather than deleting that row from
ThePortedConstants_StillAgreeWithTheVendoredReference, it moved to a divergence table pinned on
BOTH sides. Dropping a row would have left the strongest guard in the file blind to exactly the
value most likely to move again; pinning both means lowering it further AND the reference changing
underneath it each fail with a reason attached. That is the pattern for every future deviation
from a ported design — record it in the comparison, never remove it from the comparison.
AND THE ANTHILL MARK, FROM THE ARTWORK — the supplied logo, keyed off its background and used as
the nav mark, the favicon and the .ico. The first attempt hand-drew an SVG in the style of it,
which is .116's lesson un-generalised: when someone hands you the artifact, use the artifact. A
redraw is a rebuild-from-description in different clothes and fails the same way — close, and not the
thing. Worth carrying, because "I'll make a cleaner vector version" is a tempting instinct every time
an icon comes up.
2f. v0.3.8.116 — what looking at it found
Delivers: the defects .115 could not have caught, because nothing it shipped ever ran.
THE CONSOLE HAD NO RUNTIME TEST, AND STILL DOES NOT IN CI. .115 was verified by source scans
and C# facts; a browser had never loaded the renderer. A headless harness — vendored three.js, the
renderer mounted against a synthetic scene in the projection's own shape, one screenshot — found six
defects in minutes, including a camera that could not be orbited at all and a chamber that drew none
of its records. It also rendered the reference package's own renderer beside it, which is how the
halo hue and cluster tightness were measured instead of guessed. Adding that harness to CI is the
highest-value item this release leaves behind, and it is now the second item in §2e.
TWO COVERAGE BUGS, ONE LESSON. ByColony had fifteen of the registry's seventeen Colony
values, and only role ids were indexed while most executable units are workers — so unassigned
filled up in a release whose premise was that membership comes from the registry. Fourteen guards all
checked the mapping's SHAPE and none checked its COVERAGE. The new fact reads the real roster at the
typed-registry tier.
THEN THE DESIGN HANDOFF ARRIVED AND THE RENDERER WAS PORTED, NOT RE-DERIVED. Everything above
was still a rebuild from a description, and the handoff's first line is "do not rebuild this from a
description — port the working code". Its failure table names, by symptom, four of the exact things
this release had already produced. colony-renderer.js is now a port of the reference: every numeric
constant, both GLSL shader pairs, the four canvas texture stop tables, the Catmull-Rom conduit
sampling on a rotation-minimising frame, the pixel-sized crew orbs, the screen-space hit test and the
per-frame easing factors, unchanged. The only edit to its own code is the IIFE wrapper, because this
console loads plain <script src> under script-src 'self' and its assets talk through globals.
Five invented numbers went out with it: a world scaled ×14 (and the 52° field, 900-unit home distance
and fog term invented to match it), 260×mass structural grains per chamber, a PointsMaterial that
cannot express the design's hard alpha < 0.5 sprite-mask discard, a raycaster picking against a
world-space threshold, and a SECOND halo per chamber — the reference builds a glow sprite and
deliberately never adds it to the group, which reads as an oversight and is not.
AND FOUR THINGS IN THE DESIGN WERE REFUSED, WHICH IS THE PART WORTH CARRYING. They are one defect
in four costumes: each is true of the reference's generated sample data and false of this colony.
Generated records (the handoff itself says to replace buildContext() with live sources, which is
why colony-topology.js is the one file NOT ported as-is). The 120 ms mission clock, which would
animate a per-task progress number this model does not have. The continuous conduit drift, which
claims work is passing through a permanent structural link. And the ant work timer, which would show
every ant busy in a colony doing nothing. Recolouring a chamber survives; renaming does not, because
the name is the registry's Colony value.
Verified by — 3,600 tests, zero failures, plus the headless harness: nine chambers built from the projection's own shape, focus → enter → cluster drill-in, orbit, the conduit grains measurably moving between two frames three seconds apart, a schema-1 layout refused and a schema-2 layout accepted, the Micromound panel opening from its own chamber with the stop reaching the host, and zero page errors.
The twenty new facts over .115 are almost all §18, which holds the design port. The one worth
naming is ThePortedConstants_StillAgreeWithTheVendoredReference: it reads
docs/design/colony-live-3d/reference/colony-renderer.js and the port side by side and compares 61
extracted constants. Every other guard in that file asserts a literal transcribed by hand, so it can
only catch a constant being DELETED; this one catches a constant being CHANGED to something
plausible, which is how a ported renderer actually drifts. It also fails loudly if a pattern stops
matching the REFERENCE, so it cannot pass by reading nothing — the failure mode four guards in this
file have already had.
Carried, and now overdue: the typed-row ratchet has not moved for three releases. The standing
rule is one slice per release; three misses means the rule is not being kept, and .117 should
either honour it or delete it rather than let it decay into a sentence nobody applies.
2e. What comes next — the shape of v0.3.8.136 and after
The universal-workflow program closed at .113 and R0 closed at .114. This section exists
because "what is next" was being reconstructed from three documents every release, and the
reconstruction kept losing the same items.
The FORAGER first-party program (A0–A6) — the successor program, opened at v0.3.8.142
The operator's directive: the complete Forager workflow (sources → pipelines → progress → review → conflicts → exports) usable without leaving ANTHILL, published knowledge and exports reaching retrieval automatically, missions proposed and executed under standing project policy, and scoped memory/learning that later missions actually consume — while Forager remains an independently installable, usable and sellable product that stays sole owner of parsing, canonical knowledge, evidence, entities, conflicts and search.
The governing agreement is docs/FORAGER_SHARED_CONTRACT.md (contract version 1-proposed,
received while the A0 audit was being recorded): the operator's contract verbatim, the wire-name
resolution against FORAGER 0.1.4, the Phase-0 decisions — the C# package provider is REJECTED in
favour of the recipient engine's import adapter (P9); interim delivery identities are consumer
constructions labeled as such — and the authoritative producer-requirement list P1–P10.
| Phase | What it delivers | State |
|---|---|---|
| A0 | Audit, compatibility decision, pinned producer version + artifact digest, feature/API matrix, reuse/gap map, baselines, producer-side requirements P1–P6 | ✅ CLOSED at v0.3.8.142 — docs/FORAGER_A0_COMPATIBILITY.md is the gate artifact |
| A1 | Integration modes: disabled / managed (ANTHILL starts, supervises and owns a bundled engine) / attached (existing instance, untouched); pairing, credential storage, project linking UI; module boundaries kept | open |
| A2 | The complete native Knowledge workflow: upload + ingestion UI, run/inspect/cancel/retry, review with a REAL typed proposal lifecycle (pending→approved→applied/failed) that actually calls the producer, conflict operations, exports; honest progress from producer state | open |
| A3 | Durable delivery: receipts/cursors with migrations, polling reconciliation (producer has no feed — requirement P2), package delivery validated by manifest sha256, revision-preserving invalidation, retrieval at real call sites (planning context, typed cited artifacts, reachable tools) | open |
| A4 | Knowledge-change analysis vs a persisted watermark → validated mission proposals (identity, causation, evidence revisions, budgets) through the EXISTING queue/policy gates/Director; policy modes knowledge-only / suggest / auto-execute-within-limits | DELIVERED at v0.3.9.3, with one deferral named rather than glossed. The analysis, the watermark, the findings and the three policy modes are shipped: knowledge_changes records every new or moved document keyed on the seeding action key (so a six-hourly pass files a finding once, not once a day); the watermark is DERIVED from MAX(detected_at) rather than stored beside the rows, so no second fact can disagree with them; knowledge_auto_study is now `off |
| A5 | Scoped memory + learning with real consumers: source-backed candidates with provenance and validation state, staleness on evidence change, tenant scope enforced at storage AND retrieval, no success credit from imported descriptions | open |
| A6 | Both-product qualification fixture (dated update, conflict, unverified claim, procedure, adversarial embedded instruction) across the 14 scenarios; docs; paired release notes with artifact digest + contract versions | open |
Standing constraints for every phase, from the directive: audit before replacing; refusal over substitution; document instructions are never authority to execute; producer capabilities that are genuinely absent get a precise producer-side requirement and a pending gate, never a fabricated success; the consumer lives in this repository and requests producer changes through handoffs rather than cloning engine logic into C#.
The external review of the mission workflow — verified at .129, and what of it is real
An outside review of the colony-wide mission workflow arrived against main at v0.3.8.125. Every
claim in it was re-checked against the tree at .128 before any of it entered this plan, because a
review is a reading and a reading ages — and one of its two headline findings had already been fixed
two releases before it was written down.
What was already closed, and must not be re-opened. Its lead finding was that the coder resolver
selected ui_coder on a bare text.Contains("ui"), so the word "req·ui·ring" routed backend work to
the UI worker. .126 replaced that with RoutingWords word boundaries and added
RoutingWordBoundaryTests.NoWorkerRoutingBranch_DecidesOnABareSubstring, a source guard that fails
if any routing branch decides on a bare substring again. The review's own failure logs predate it.
What is verified and queued. In order, and the order is the review's ranking corrected for what this tree actually does:
| # | Finding | Where | Pinned by a test? |
|---|---|---|---|
| 1 | ✅ CLOSED at v0.3.8.132. A mission that wants no worktree now enters a READ-ONLY scope over Project.Path, materialising nothing; MissionWorkspace.Writable (default true) is the discriminator, and the five consumers that read the ambient scope in order to WRITE — the agent-CLI working directory, the change harvester, the edit processor, the change summary and the acting-coder branch — read CurrentWritable instead, so a read-only scope is indistinguishable from no scope to every one of them. Was: A project mission's source root never becomes tool scope. Queen.RunMission gates workspace activation on EnableFileWriting || EnablePatchApplication || EnableActingCoder — all three default OFF. Project.Path is loaded and then used only to build a workspace that is never built, so under proposal-only defaults every file and check tool falls back to agent_workspace_dir. A mission about project X reads, patches and tests some other tree. |
Queen.cs ~775/812 |
No |
| 2 | ✅ CLOSED at v0.3.8.134. ToolEvidence.For now records NOTHING for a run_allowlisted_check that failed without an exit_code= line — the four not-started paths (id not in the catalog, check disabled, timed out, process could not start) leave the event log's tool_called/tool_completed and no evidence row at all, so a missing dotnet can no longer satisfy DiagnosisIntegrity's reproduction requirement. The discriminator is CheckRunner's own output, not a guess at the error text. Was: A check whose process never started is filed as a reproducible test failure. CheckRunner returns a typed failure with empty output; the tester greps for exit_code=, writes n/a, and then converts EVERY unsuccessful tool result into VerificationFailure + retryable + a medic handoff, discarding the tool's own class; ToolEvidence marks every run_allowlisted_check deterministic on the tool NAME alone. A missing dotnet and a failing test are the same row. |
CheckRunner.cs, SpecialistAnts.cs, ToolEvidence.cs |
Partly — the not-started path has no test at all |
| 9 | ✅ CLOSED at v0.3.8.135, and it was not in the review — the guard written for it found it. Two steps the colony inserts into its OWN plans named a task type the assigned role's contract does not declare, so the dispatch chokepoint blocked both every time they fired: the adaptive controller's delta-verification (verifier / verify, and Critical, so it failed the mission) and EnsureClassCoverage's research retrieval (web / research). HandoffGate had the same hole from the other side — it refused an unreconciled type, and a refused REQUIRED handoff is a deterministic block. All three now go through TaskTypeVocabulary.Reconcile; the authority and the refusal are unchanged. TaskTypeReachabilityTests reads every AssignedAnt/TaskType literal pair in src/ against the live catalog — the third population, after the seven pairs RosterContractTests names and the handoff routes HandoffTaskTypeTests reads. |
ExecutionService.cs, Planner.cs, HandoffGate.cs |
Yes — and the sweep found the second bug before it had run once |
| 3 | ✅ CLOSED at v0.3.8.137. ConversationRunner.Run fires an onMissionSettled callback from the one point that runs whether the pipeline returned or threw; ProjectScheduler keeps the run running until it fires and stamps the terminal status from the MISSION ROW (complete/partial/failed), so overlap detection finally has something to see — pinned by ARunStaysRunning_UntilItsMissionSettles_AndOverlapSeesTheFlight, with the second occurrence skipping mid-flight. ScheduleRun carries MissionId (schedule_runs.mission_id, surfaced through both run endpoints). The test fake now persists a real mission row and takes real time — the synchronous-fake blindness was fixed with the defect, not around it. Ask-mode runs still park as waiting_approval at the gate, where no callback comes. Was: A schedule run is complete before its mission has done anything. ConversationRunner returns as soon as the mission ROW exists; ProjectScheduler reads outcome.Started as "complete" and stamps FinishedAt. Overlap detection looks only for a running row, so occurrences can overlap. ScheduleRun carries no mission or job id. |
ProjectScheduler.cs, ConversationRunner.cs |
Yes — from both sides of the settle boundary |
| 4 | ✅ CLOSED at v0.3.8.134. ClassifyPatchJson grades a REASONED zero-proposal result a success, mirroring the acting path's NoChangesMarker; a zero-proposal result with no summary is still InternalDefect. Was: A justified no-change is an internal defect. ClassifyPatchJson grades a well-formed response with zero proposals as InternalDefect. The typed vocabulary already exists one path over — the acting coder's NO_CHANGES_NEEDED — and grades it a success. |
Ants.cs |
Yes, and the fixture is itself a justified refusal |
| 5 | ✅ CLOSED at v0.3.8.134. Both refusal paths now return through Queen.FinalizeRefusedMission, which persists the mission and its evaluation, compiles the operator report, logs mission_outcome carrying the refusal's own reason, invokes onMissionFinished, and returns what the normal path returns. Pinned by FinalizationOrderTests.BothRefusalPaths_GoThroughTheSharedExit. The try/finally around the mission body is NOT part of it and stays open. Was: Two early returns skip finalization. Dispatch-plan refusal and preflight refusal persist a status and return; neither compiles the operator report, emits the outcome event, or invokes onMissionFinished. So a BLOCKED mission — the state those very comments insist is not a failure — reaches the operator's job list as a bare failed with no reason. No try/finally wraps the mission body either. |
Queen.cs |
No |
| 6 | ✅ CLOSED at v0.3.8.137. The project travels submission → durable mission_jobs.project_id (migrated in place, legacy rows null) → requeue-after-crash → RunMission(projectId:). The Director passes objective.ProjectId; POST /missions accepts project_id and refuses an unknown project at the door; idempotent replay returns the ORIGINAL submission's project. DurableMissionRuntimeTests pin restart survival, replay, whitespace-is-null, and the legacy-database migration. Was: ProjectId is lost through the Director and the API job path (ApiMissionJob has no such field; POST /missions has no such field). The conversation/schedule path does carry it. Bites when the Director replaces a cron. |
ColonyDirector.cs, ApiJobRegistry.cs |
Yes — at the durable layer |
| 7 | ✅ CLOSED at v0.3.8.138 — and it was the review table's last open row. An operator-requested plan travels into planning on the MissionContext, and the executed graph IS the plan: same task ids (the persisted mission_dispatch_planned record and the graph join row-for-row), same types and roles, the operator's declared edges, nothing invented — no auto-wiring, no EnsureClassCoverage (a plan not positioned to deliver is REFUSED by preflight, not silently repaired). Materialized tasks go through the SAME admission pipeline as planner-authored ones (AdmitTask, extracted not copied), and a consequential plan still gets the policy verifier. Recorded as mission_plan_from_dispatch, the complement of mission_plan_substituted. Pinned by DispatchPlanMaterializationTests, with every plan obtained from DispatchPlanner.Plan itself. Was: The validated dispatch plan is logged and then discarded — PlanningService.CreatePlan builds the executed graph independently. Only bites when an operator actually requested a workflow, which is the one case the stage exists for. |
Queen.cs, PlanningService.cs, MissionContext.cs |
Yes — from the producer's side |
| 8 | ✅ CLOSED at v0.3.8.134. FallbackTasks and EnforceConstraints both split their keyword lists into a PREFIX set and an EXACT-WORD set, so repo, add, class, edit and create are matched whole and report, address, classification, editor and creature no longer route to the code lane. Was: FallbackTasks still decides the CODE lane on "add" and "class". .126 moved it to AnyPrefix, which anchors at a word START, so "req·ui·ring" and "un·change·d" were genuinely fixed and "address" and "classification" were not. The comment claims four; it earned two. |
Planner.cs |
No |
Where this plan disagrees with the review. Three of its findings are accurate readings of code that is the way it is on purpose, and adopting them as defects would undo argued decisions:
- Spec ingestion on length alone is real and is a standing design choice, defended in writing by
EvidenceGroundedPlanningTests— abandoning it trades a mission that inspects nothing for a mission that overflows its context. Changing it is a design change with its own case to make. - Structural status beside
verification_status: failedwas ATTEMPTED at.122, reverted, and the reason written down where the next attempt will find it. It is blocked on the per-task execution record, exactly as this section already said. - The tester judging the base tree is real, partially mitigated by the dispatcher re-entering a live revision scope, and its full answer is R6/R7 work rather than a repair.
Its P2 and P3 sections are a roadmap, and they substantially restate R6, R7 and R9. They do not reorder this one.
Items 1 and 2 are one slice. Together they are the whole of what the operator actually observed —
patches proposed against .anthill\workspace, a dotnet_test that "failed" in 0.06 seconds with no
exit code, and a medic loop that re-ran a coder against a failure nothing had changed.
The orchestration slice — .118 opened it, .122 closed two findings, and the rest waits on one row
WHAT .122 CLOSED, and why it could be closed without the execution record. Two of the three ▲
findings below turned out to be about a decision that was MADE and then not recorded anywhere a later
stage could read — which needs a join or an event, not a new table:
- The closure reconciliation was TRIED AND WITHDRAWN, and that is the more useful result.
mission.StatusandVerificationStatusstill do not meet. The join failed on a fact the map did not have:Verification.Faileddoes not mean "a check said no" —IsSatisfiedneeds the verifier's own verdict to be a PASS, andVerifierAntdowngrades a model-authored pass toUnknownwith no deterministic evidence behind it, sofailedspans "the check said no" and "nothing could satisfy the check". Demoting on it reclassified a legitimately complete mission. Do not begin the next attempt by making the status line readVerificationStatus. Splitting those two meanings IS closure enforcement, and it needs the execution record below. - The planner's substitutions. All five now reach the event log as
mission_plan_substitutedwith stable reason codes, so a fallback plan stops being indistinguishable from a requested one. Thegoal.Length > 6000gate is UNMOVED: it should route on whether aRequestedWorkflowwas supplied, and trading a known-bad heuristic for an unmeasurable one before the execution records exist would be a worse trade than leaving it.
THE ROW LANDED AT .139, AND IT WAS NOT A NEW TABLE. This paragraph said for eleven releases
that items 3–8 "all consume the same missing row, and .122 did not add it", and asked for "one
record per task attempt, written where the scheduler already writes the terminal state, carrying the
facts Domain.Task marks transient". Read against the tree, that describes task_attempts exactly —
live and load-bearing since v3.8.0, because the atomic claim runs on it. What was missing was never
the row; it was any fact about what the attempt DID. So .139 extended the row the claim already
writes rather than building a second one beside it, which would have been two records of one thing —
defect #5 on this document's own list, shipped deliberately.
task_attempts now carries assigned_ant, task_type, worker_basis, deliverable_ids,
required_capability, generation_degraded, produced_revision_id and ran_revision_id, written
at ExecutionService.CloseAttempt — the chokepoint whose own remark already explains why it is the
right one, since every path that ends a task passes through it with its final status set. They are
written there and not at claim time because they are properties of a FINISHED attempt: which tree a
check judged is not known when the claim is taken. TaskAttempt was already typed, so the ratchet is
satisfied without a new reader. ExecutionRecordTests asserts across a store close and reopen,
because "transient" is the thing being fixed and a test reading the facts back through live objects
would pass identically against the code this replaces.
NULL MEANS NOT RECORDED, NEVER "NO", and the gates built on top of this must honour that. Legacy
attempts migrate in place with no values. A closure gate reading those nulls as "this check ran in no
revision" or "this generation was fine" would refuse every mission the colony has already run —
inventing history to satisfy a guard, the direction evidence.revision_id and
patch_sets.base_fingerprint both refused.
CLOSURE ENFORCEMENT LANDED AT .140, AND IT DID NOT NEED THE RECORD. That is worth stating
plainly, because this section named the missing execution row as its prerequisite for eleven
releases and it was not one. .140 SPLIT THE WORD rather than the join, and needed nothing new to do
it: VerificationVerdict.Parse has separated Failed and NeedsImprovement from Unknown since
v2.19.0, and only the mission-level status flattened them. Now failed means a verdict-bearing
task said no, the new inconclusive means nothing could establish a verdict, and the evaluator
demotes a COMPLETE mission to Partial on the first only. inconclusive is exactly the set .122
demoted on and it still demotes nothing — which is why the plan's own instruction, "do not begin the
next attempt by making the status line read VerificationStatus", is followed literally. The truth
table in CharacterizationTests changes exactly one row, with the reason beside it as that table
requires. ClosureEnforcementTests owns it.
AND .140 TRIED A SECOND FIX AND REVERTED IT INSIDE THE RELEASE. VerifierAnt decides a verdict
(v3.8.27) and records it; the mission gate re-parses the model's prose instead, which reads exactly
like defect #5 with the authoritative implementation losing. Pointing the gate at the ruling failed
twenty integration tests across five mission classes, and they were right: the ant's verdict answers
a PROMOTION question — is there DETERMINISTIC evidence — and an audit, an external action and a
system action have none by design, so the ruling is Unknown for entire classes of legitimately
verified mission. "May this be promoted" and "did the verifier judge this acceptable" are different
questions. Recorded in MissionVerification.VerdictOf and pinned by ClosureEnforcementTests so it
is not re-tried.
.139's record is not thereby unused: it is what makes an attempt's execution survive the process,
which every gate below still needs.
WHAT STILL WAITS: the remainder of items 3–8 — artifact and evidence handoff, verification that reads execution rather than a narrative, and unsourced-claim rejection.
Read the colony before picking the next slice — v0.3.8.141
.141 was the first release chosen from the operator's LIVE DATABASE rather than from this document,
and the result argues for making that the habit. The plan's queue had closure enforcement, evidence
handoff and unsourced-claim rejection at the top; the colony's actual dominant failure was none of
them.
66 real missions: 37 escalated. Every one of the 39 required_handoff_refused events ended that
way — 18 mission task budget exhausted (12/12), 15 near-duplicate handoff suppressed, 5
unsupported task type, 1 depth limit. Thirty-three of thirty-nine were the runtime's OWN growth
bounds killing the mission they exist to protect, and v3.8.25 had predicted exactly that in a
comment while exempting only optional handoffs from it. Nothing in this plan named it, because
nothing in this plan was reading what the colony actually does.
The check .139 queued is also answered, and the answer is a deletion rather than a caller.
IEvidenceStore.HasDeterministicPass is "≥1 deterministic passing row exists"; EvidenceVerdict.For
returns Failed when ANY deterministic row failed. It is strictly weaker — it would pass a mission
holding a failed check — so it must never gate anything, and its doc comment claiming "every
promotion path asks it" is false: nothing asks it. It survives as a store-level test probe and should
be described as one.
One structural fact a session starting that work needs, and it is not in the brief:
IEvidenceStore.HasDeterministicPassis implemented and called only from tests. Before giving it a production call site, check whetherEvidenceVerdict.Foralready answers the same question — two implementations of one rule is this repository's defect #5, and adding a caller to close a "declared and reaching nobody" finding would be exactly the adjacent-question mistake if so.Closed atDispatchPlan.Tasksis validated, logged, and then discarded..138— the executed graph IS the plan, same task ids, so an attempt row'stask_idalready joins to the dispatch plan row..139therefore did NOT add aplan_task_idcolumn: a second identity for one thing is the same defect this section keeps naming.
docs/ORCHESTRATION-FINDINGS.md remains the evidence, measured against fcf12a7; the ▲ items below
are kept because their reasoning is what redirects the work, not because they are all still open.
.118 shipped the input contract (RequestedWorkflow) and the pre-dispatch stage
(DispatchPlanner / DispatchPlan) and claimed nothing beyond them. The remainder of the brief —
authoritative execution records, artifact and evidence handoff, verification that reads execution,
closure enforcement, unsourced-claim rejection — all depend on ONE missing thing, and naming it is
what .118's investigation bought: there is no per-task authoritative execution record. Every
downstream item is a consumer of a row that does not exist yet.
docs/ORCHESTRATION-FINDINGS.md is the map, gathered by reading the eight code locations the brief
named rather than reasoning from the symptom. Three of its findings change what the remaining work
should be, and a session that skips them will build the wrong fix:
▲
checks: 0counts the wrong thing and gates nothing. Two unrelated evidence concepts exist.IEvidenceStoreholds the durable rows verification actually reads.MissionReport.Checkscounts per-taskAntEvidencerows of kind"check", written in exactly one place —TesterAnt. Sochecks: 0means "no tester dispatchedrun_allowlisted_check", not "no evidence", and its only consumer isRenderitself.IEvidenceStore.HasDeterministicPassis implemented and called ONLY from tests. Closure enforcement must gate on the store, not on the display counter.▲ Verification is stronger than the brief assumed; the leak is upstream of it.
VerifierAntasks the evidence store first and downgrades a model-only PASS toUnknown;EvidenceVerdict.Forcounts onlydeterministic == truerows;MissionEvaluator.EvaluatedistinguishesNotRunfromFailed. Prose alone cannot producecompleted_verifiedin the wired configuration. The defect is thatmission.Statusnever consults any of it —Queen.cs:1299-1302computes structural status purely from task terminal states. The fix belongs atQueen.cs:1299-1302, not inVerifierAnt.▲ The planner's provider fallback is invisible to the mission record.
Planner.CreateTaskssubstitutes a static plan on four conditions, every one of them recorded only byConsole.Error.WriteLine— no event, no artifact, nothing an operator or a later stage can read.ResearcherAntandBuilderAntoffline fallbacks return plain"succeeded"with no warning, andWebResearchAnt.SummarizeSourcesilently substitutes a truncated snippet. This is the cheapest item on the list and the one with the worst failure mode: a colony that ignored the goal, with a green run behind it.
Two more from the same inspection, lower cost and worth carrying:
Strategist.GenerateGoalhas a model rewrite the charter for standing objectives with no check that the rewrite preserved meaning. Every other human path reachesMission.Goalverbatim.- Cross-mission recall is prose with artifact ids discarded (
SqliteMemory.Operations.cs:1299-1301) while within-mission recall is typed and keeps them (ArtifactContext.cs:192-195). Unresolved artifact ids are reported ONLY inline in the prompt text —ArtifactContext.cscallsLogEventnowhere, so an unresolved handoff is invisible to the record.
And the gate .118 deliberately did not move. Planner.cs:128 still routes on
goal.Length > 6000. It should route on whether a RequestedWorkflow was supplied, but changing it
before execution records exist would trade a known-bad heuristic for an unmeasurable one.
The named findings from .114–.115, in the order they cost the most.
A MODULE CANNOT PERSIST AN EVENT, so its events cannot be replayed or reconstructed. Measured after
.115shipped, correcting a figure that release stated wrongly: there are 12 bus-onlyEvents.Publishcall sites across 11 files, 8 of them Micromound — not the "31" the.115entry claims. The other 198 event sites go throughSqliteMemory.LogEvent, which already writes the row and THEN publishes it, so the ordinary path was never the problem.The 12 are not carelessness, and that is the finding.
Anthill.Coreis off-limits to a module by the boundary rule; a module receives anIModuleContextand an event bus and has no persistence path at all. Micromound publishes to the bus because publishing is the only thing it can do.Three consequences, all live: a reconnecting console silently loses those events, Colony Growth Playback cannot show them, and no audit after the fact can prove a charter was issued.
.115states the limit honestly in the timeline's coverage line, which is the right thing to do about a gap and is not a fix.The fix is a SEAM, not a sweep.
EventTypesalready requires every event type to be declared andEventVocabularyTestsenforces it, so the declaration is where durability belongs: each type says durable or transient, and the composition root — which may name a module — persists the durable ones. The guard is available at the TOP tier for once, which is rare for this kind of work: publish a declared-durable event, then assert the row is in the table. No source scan.The typed-row ratchet, still at 45, skipped twice.
.114and.115both spent themselves elsewhere.TheUntypedStoreSurface_OnlyShrinksenforces ≤ 45 so it cannot reverse, and nothing makes it fall. One slice per release, each lowering the ratchet in the same commit, is the standing rule — a third skip means the rule is not being kept, and it should then be rewritten or dropped rather than quietly missed.Three widget payloads read by nobody.
mound_fleet,mission_statusandevidence_feedare built on every Micromound mutation for a widget runtime that was going to render them..115's console reads the routes directly instead. That is either a second surface to retire or a first surface to finish, and it should stop being both — the same shape asollama_model_presentand/config/healthbefore it.The console has no runtime test. Every Colony Live guard is a source scan or a C# fact; no browser loads the renderer in CI, which is why
.115shipped a view with no orbit control and no record particles and went green. A headless Chromium harness that mounts the renderer against a synthetic scene and asserts on the rendered frame — it built, it drew N chambers, no page errors — would have failed on every one of those..116built that harness ad hoc; it belongs in the repo.Anthill.Tests.Micromoundis outsideAnthill.sln. It builds under theMICROMOUNDdefine against a sibling checkout, so a solution-wide run is complete for the solution and silently excludes 166 tests — plus, as of.115, the console vocabulary guards. Every release has to remember a second command. Folding it in behind a property, or making the solution run fail loudly when it was skipped, removes a step that currently depends on somebody remembering.Colony Live's remaining gaps. No pheromone overlay bound to real scores (nothing in the model carries a per-record score); the canvas fallback plays no transition flights; and sector membership has no history, so a reconstructed frame is drawn in today's chambers and says so.
The R-numbered order is unchanged. R1 is the live one. R4 needs the live pack, which is the operator's step. R6 gates R9, and nothing below R6 has started.
R1 — Contract and adapter truth · 1–3 releases · in progress
Nothing downstream can be attributed until the colony's declarations match its runtime. A live failure today cannot be pinned on the model, the adapter, or a contract that was never true.
- ✅ Five deterministic roles declared
AllowsModelCalls: true(v0.3.8.76 — contracts state what their ants can do;ContractDeclarationTestsasserts a declaration belongs to a role that can act on it, and that a contract agrees with itself aboutmodel.invoke). - ✅ Three routes called models with no declaration at all (v0.3.8.76 — found while fixing the
above, and the more expensive half.
planner,strategistand the answer-synthesisscriberoute were never graded, because the fitness report enumerated contracts and they have none.ModelRouteRequirementsnow declares all eight routes that reach the router, in both directions). - ✅
ResponseSchemaJsonwas declared and reaching nobody (v0.3.8.76 —GenerateTypedtakes a schema;coder,plannerandstrategistsend one. Each schema written against the parser rather than the prompt, which caught three disagreements that would each have been an outage). - ✅ Verifier's structured-output requirement (v0.3.8.76 — checked, and it described a use of the model that v3.8.22 ended. The verdict is deterministic and the model's reading never promotes; removed, and the context requirement that is real kept).
- ✅ A guard for tick-versus-body disagreement (v0.3.8.76 —
ChecklistIntegrityTests). - ✅ Provider-adapter conformance suite (v0.3.8.77 —
AdapterConformanceTests: four adapters × eight capabilities, thirty-two cells, each either citing the test that proves it or naming why the transport cannot. Citations are checked to resolve, and a fifth provider added to the module fails the matrix rather than passing by being unknown to it.) It found a live defect on its first pass:ModelCapabilityCatalogdeclaresanthropicasStandard— structured output included — soNegotiatekept a response schema for it, andAnthropicBodynever read the field. Fixed by binding the schema the only way Anthropic serves one, as a forced tool call, withReadAnthropicunwrapping the reply back into content. - ◻ Ollama capability discovery. Reads
/api/tags, whichApiHost.Providers.cs:112documents as a deliberate choice — Ollama publishes a per-modelcapabilitiesarray there, and against three real local models the hand-written table was wrong twice./api/showmay still be richer. Treat as a contested design decision to extend, not an oversight to fix. Carried to R4, where a live multi-provider run is what would settle it; it is not a gate on anything before that.
Exit gate — ✅ CLOSED at v0.3.8.77. Every role's declared capabilities match what it can do, asserted by a test (v0.3.8.76). Every adapter passes the conformance suite or is explicitly marked unsupported for a named capability (v0.3.8.77).
R2 — Finish deterministic qualification · ✅ CLOSED at v0.3.8.79
The two scenarios the ledger overstated, and the two defects closing them uncovered.
✅ Scenario 5 — the composed UI-patch lifecycle (v0.3.8.78 —
ComposedUiPatchLifecycleTests. The gate and the producer were each proved and the JOIN was not: a map with the right shape and the wrong mission, or a gate reading a store the cartographer never wrote to, satisfies both ends while the middle is broken. The composed run drives goal → cartographer reading REAL UI files → gate admits → coder → tester and soldier → verification → applied bytes. The map is not scripted, becauseUiCartographerAntholds no router to script — if it reads nothing, the gate refuses and the test fails, which is the coupling the scenario claims.)✅ The build verifier reads
CheckSource(v0.3.8.78 — surfaced by the above and blocking it.RunAllowlistedCheckToolhas resolved ids throughCheckSourcesince v0.3.8.73;BuildVerifierstill asked for the literaldotnet_build, so a code patch in any non-.NET workspace randotnet buildagainst a directory with no project and could never be verified. Where an operator declares checks, those checks are now the build — all of them, any failure fails, the result stays deterministic, an empty selection fails closed, and the no-declaration fallback is unchanged.)✅ Scenario 17 — process death mid-apply (v0.3.8.79 —
ProcessDeathMidApplyTestsand theAnthill.CrashHelperexecutable. Recovery was proved by ABANDONING a transaction object, which is still a healthy process with flushed buffers and run finally blocks; a killed one has neither. The test starts a real process, waits for it to signal that the journal and patched bytes are durable,Kill()s it, and recovers in the parent. The sentinel is what makes it deterministic — killing on start would sometimes find no journal, and "recovered cleanly from nothing" is a pass that means nothing happened.)✅
RunShellmangled double quotes on Windows (found at v0.3.8.78, fixed at v0.3.8.79.)AutoApplyRunner.RunShellpasses the whole command throughProcessStartInfo.ArgumentList.Add, which escapes by C-RUNTIME rules — an inner"becomes\"— andcmd.exedoes not follow those rules. A verify command written asfindstr /C:"aria-label" filereaches findstr as/C:\"aria-label\", matches nothing, exits 1, and a correctly applied patch is rolled back with "Verify FAILED" against a tree where the change is present. Two instances:autonomy_autoapply_verify_cmd, and the auto-commitgit -c user.name="ANTHILL Auto-Apply" … -m "{msg}"atAutoApplyRunner.cs:639, which no test exercised. Every quoted verify command configured in the field was affected. Fixed by handing Windows the raw string (psi.Arguments) so cmd applies its OWN rules, while Unix keepsArgumentList— there is no command-line re-parsing on that side, so a string would introduce the very re-quoting this removes.ShellQuotingTestsproves a quoted phrase survives, that an absent phrase still fails, and that&&still composes. The sweep for the same class found a THIRD instance, and the worst-placed of the three:OperatorShell.Execute, the dashboard's shell box, where an admin typinggit commit -m "fix the thing"had it delivered as-m \"fixplus two stray arguments — a human typing a correct command and watching it come back wrong, with nothing in the output explaining why.ShellSpawnTestsnow pins the rule repo-wide in both directions: no/cthroughArgumentList, and sites invoking a real program (git,docker, an agent binary, a declared check) keep the list, because over-applying the fix is the same defect from the other side.
And the guard that predicted its own expiry.
PartialCoverage_IsDeclaredRatherThanImpliedassertedNotEmpty(partial), which would have failed for the single outcome the ledger exists to reach. v0.3.8.78 recorded that here and left it standing, because it was still true then; v0.3.8.79 removed the assertion in the release that made it false. Same correction v0.3.8.74 made to its sibling — a guard that cannot express success is not a guard, it is a deadline — and the second time this file has needed it.
Exit gate — ✅ CLOSED at v0.3.8.79. 20 of 20 closed by substance, with no note admitting an unproved claim.
R3 — Per-role cancellation and timeout · ✅ CLOSED at v0.3.8.88
The largest single item remaining, and the other large area where nothing had ever asserted the outcome.
✅ The cancellation matrix and its harness (v0.3.8.80 —
RoleCancellationTests. Twelve roles × four points = 48 cells, every one decided: 24 driven live by the harness, 11 cited to the tests that already prove them, 13 marked not-applicable — and the not-applicable claims are CHECKED against the contracts, so a role that acquires a tool or a router stops being exempt and the suite says so.) What it closed that nothing had: every prior cancellation test was about a MECHANISM — the ambient model-call scope, the process tree, a hanging subprocess. None was about a ROLE. "Does cancelling stop the archivist without writing a lesson to durable memory" had no answer anywhere, and the damage differs per role: a cancelled tester leaves a process, a cancelled archivist leaves a memory, a cancelled coder leaves a patch set.✅ Acceptance gates 1 and 2 (v0.3.8.80 —
AcceptanceGatesOneAndTwoTests. Gate 1 asks for READY, which is stronger than the gate-only check that existed:Readyis a conjunction of five conditions and "no gate blocks it" reads exactly like "it is ready" in a summary. Gate 2 reads its four clauses from four different sources on purpose, because the gate is about them agreeing.)✅
during_generationandduring_tool_calldriven live for six of the eleven applicable cells (v0.3.8.81 — researcher, web and coder mid-generation; file, researcher and web mid-tool-call). It found the defect this item predicted, and a second one behind it. Every model-calling role reads a non-Ok call as "the routed model is unavailable" and DEGRADES — which is right for the case it was written for. A cancelled call is non-Ok, so cancellation came through that same door: the researcher and the builder returnedSucceededWithWarnings, the task COMPLETED, and a completed task ingests handoffs, inserts a verification task after a deliverable, hands the archivist something to remember and processes the coder's proposals. The operator pressed stop and the colony answered with a fabricated fallback deliverable and more scheduled work.DrainRunningTaskshas recorded this state since v2.26.0 — for tasks still RUNNING at the grace deadline; one that finished INSIDE it by degrading was never its business, so the faster a role gave up, the more likely its cancelled work was recorded as a completion. Fixed once inExecutionService, where what a stopped mission may record is decided, rather than at the eight ant call sites, where how to handle a bad model call is decided.✅ A cancelled call is no longer reputation (v0.3.8.81 — the second defect, and the one that outlived its mission).
ModelRouter.SendCoreheld two implementations of one rule four lines apart: the breaker readCancelledas Neutral — "we stopped the call ourselves" — while the pheromone trail derived everything fromOkand wrote a FAILURE againstmodel:{provider}:{model}:{role}. The transient copy was right and the durable copy was wrong, so every operator stop taught the colony a little more firmly that the model its cancelled role was using is unsuited to that role. This is a wrong memory that R8's exit gate could not have traced, because the mission that wrote it looked fine. One authority now,IsColonyStopped, asked by both readers.✅ The graduation record's cancellation column, and the
ui_cartographerfault cell (v0.3.8.81 — the last nulls in the record.UiCartographerFaultTestscovers the branch nobody exercised: a failed listing TOOL, as against the empty WORKSPACE the unit tests already drive. A broken tool producing an empty map would be admitted byUiChangeGate, which asks whether a usable map exists rather than whether the task succeeded.) Both gap-asserting tests were rewritten rather than relaxed, on their own instructions — the third time this file has recorded that correction after v0.3.8.74 and v0.3.8.79. A guard that cannot express success is not a guard, it is a deadline.✅ The harness ran the plans it wrote (v0.3.8.82 — and this is the release, because until it none of the above was about the roles it named).
Planner.TasksFromJsonrejects any plan with fewer thanMinDynamicTasks(3) usable tasks and the mission silently runsFallbackTasksinstead — a static researcher/file/coder/builder/verifier graph. Every plan this harness ever scripted was below that minimum: one task in the live fixture, two in the pre-dispatch one. So each cell passed because a fallback branch happened to contain the role the assertion was looking for, which is this repository's oldest defect shape pointed at its own test fixtures.AssertTheMissionRanTheScriptedPlannow compares the planned roles against the scripted ones on every cell, andScriptedPlanbuilds a three-task graph with the role under test first.✅
during_generationandduring_tool_calldriven live, for real this time (v0.3.8.82 — researcher, web, coder and builder mid-generation; file, researcher, web, ui_cartographer and scribe mid-tool-call). The three cells v0.3.8.81 recorded as "attempted and did not reach the point" all had the same cause, and none of them was what that release guessed at: the builder was never planned, the scribe was never planned, and the cartographer's gate was tripped by the FALLBACK plan's researcher dispatching a tool inside the cartographer's grant. The v0.3.8.81 note worrying about "a dispatch outside per-role authorization attribution" described a fixture artefact, and is withdrawn.✅ The medic and the archivist cannot be driven by planning a task for them (reasoned at v0.3.8.82, LANDED at v0.3.8.83 — that release documented the reclassification and shipped a matrix that still drove all twelve, because the edit was lost and nothing compares a stated count against the matrix that produces it. A contract fact, found only once the plans became real).
AntRegistry.ValidateTaskrefuses a planner-produced task for aFailureTriggeredorPostFinalizationrole, so theirbefore_dispatchandawaiting_dependencycells are now NOT-APPLICABLE with that reason, checked against the contract rather than asserted. They looked driven for two releases for the same reason everything else did.✅ The sweep for the same defect elsewhere in the suite (v0.3.8.83 — and it found a second instance).
ScriptedPlanConformanceTestsreads every.Role("planner", …)in the test suite and requires each scripted plan to be one the Planner would ACCEPT: at leastMinDynamicTaskstasks, and only planner-eligible roles.EarnedRepairLifecycleTestsscripted two — researcher and coder — for a goal containing both "document" and "add", which selects the fallback's CODE branch. That branch has a coder, so the patch → failing check → medic → repair → passing check loop the whole scenario asserts still happened, and qualification scenario 15's last edge was being proved about a plan nobody wrote. A plan may satisfy the guard statically, or its fixture may verify at runtime the wayRoleCancellationTestsnow does; the guard also asserts it found at least five plans, because a sweep that silently stops sweeping is the failure it exists to prevent.✅ The last two cited cells, driven (v0.3.8.84 — and the citation was wrong, not merely unfinished). Both
verifier/during_generationandtester/during_tool_callwere recorded as unreachable becauseSchedulingMode.PolicyInsertedmeant "no plan may assign it".AntRegistry.ValidateTaskrefuses onlyFailureTriggeredandPostFinalizationfrom planner output, and v0.3.8.51 narrowed it to those two on a field report: a planned tester or soldier step is a plan asking for MORE safety, not less; PolicyInserted is a floor, not a ceiling. The soldier — also PolicyInserted — was being driven by this same harness at both universal points the whole time, which is the contradiction that should have been visible from inside the file. A declaration disagreeing with the runtime, written into the matrix whose job is catching those. The tester's cell is SPLIT rather than claimed whole: the harness proves what the role leaves behind, and the orphan-process half stays cited to the two tests that prove it, in the same cell, where all three citations are checked to resolve.✅
archivist/before_dispatch, and the defect it was hiding (v0.3.8.85). The cell was recorded not-applicable becauseAntRegistry.ValidateTaskrefuses a planned archivist — true of the PLANNER and false of the role.Queen.RunArchivistAfterFinalizationis its real dispatch site, invoked directly once the canonical evaluation is persisted, and "cancel before this role acts" is answerable there. The answer was that nothing stopped it: that path does not go throughExecutionService.RunSingleTask, so it could not inherit v0.3.8.81's stop check and had none of its own — a cancelled mission still ran the archivist over its partial work and ingested the candidates it proposed. Fixed by reading the persisted outcome, the authority this method's own documentation already insists on, and skipping with the existingarchivist_skippedevent under a distinct reason. And it shows how this harness can pass for the wrong reason even now. The no-positive-memory property watchesmemory_candidate_archived, and a stopped mission usually gives the archivist nothing worth proposing — so the assertion passed because the archivist found nothing, not because it was prevented from looking.✅
medic/before_dispatch— DRIVEN LIVE (v0.3.8.88). The last gap, closed from the medic's real trigger rather than excused. v0.3.8.83 had already written down what it would take — "a critical task that fails under adaptive mission control" — and the fixture existed one file over:CodePatchLifecycleTestsdrives a patch mission whose policy-inserted tester runs a check against the materialized revision, fails, and hands off to the medic.The window is exact rather than approximate. Both admission paths —
IngestHandoffsandApplyAdaptiveDecision's repair arm — admit the task FIRST and log afterwards with the destination role as the event's ant name, so that event means scheduled, persisted, not yet dispatched. The fixture stops the colony on it, through a synchronous test bus: the productionInProcessEventBusdispatches off the publisher's thread by contract, which would make the stopping instant a race and the cell a coin toss.◻ The two cells that remain are FACTS, not gaps:
medic/awaiting_dependencyandarchivist/awaiting_dependency. Named individually becauseRoleQualificationRecordTestsasserts this document keeps naming whatever is not driven. Neither can be produced by the runtime, and each says why in the matrix.medic/awaiting_dependency— not-applicable on a runtime fact, since v0.3.8.87. Its old reason was "a role the planner may not assign has no planned task that can sit waiting", which is true, and answers a question adjacent to the one the cell asks: the medic DOES get a task, from two runtime paths. Neither gives it a dependency, and neither can. Its parent is a task that has already FAILED, so an edge onto that parent would never be satisfiable and the role would deadlock rather than wait — which is whyApplyAdaptiveDecision's repair arm setsParentTaskIdsand leavesDependsOnempty while the delta-plan arm four lines below sets both.HandoffGatelikewise constructs its task with no dependency.RoleCancellationTests.ANoFailureTriggeredRole_IsEverGivenADependencyholds the creation sites to that, so the claim now fails when the code changes rather than when someone rereads it.archivist/awaiting_dependency— the strongest not-applicable in the matrix. The role is never SCHEDULED at all; the Queen invokes it directly after finalization, so there is no queue entry for a dependency to hold up, now or ever. It does not depend on how a task happens to be constructed.
Exit gate — ✅ MET at v0.3.8.88. All forty-eight cancellation cells decided; every driven cell proved to have run the plan its fixture wrote; 33 driven live, 0 cited, 15 not-applicable. The medic's and archivist's universal cells were the gate's real substance and none of them was excused:
archivist/before_dispatchwas driven at v0.3.8.85 by looking at the role's own dispatch site instead of the planner's,medic/before_dispatchat v0.3.8.88 from the tester→medic handoff, and the twoawaiting_dependencycells are recorded as facts the runtime cannot produce — with the medic's reason rewritten at v0.3.8.87 from a claim about the planner to a proof about the scheduler, pinned by a source guard. The graduation record — ✅ complete at v0.3.8.81. Acceptance gates 1 and 2 — ✅ closed at v0.3.8.80.
R0 — Correctness before capability · 4–6 releases · in progress, gates everything
Inserted at v0.3.8.91 after an external repository review, and placed FIRST because its findings are not stylistic: several documents promise guarantees the runtime does not enforce, and two of them were reachable. Every claim was verified against the code before being accepted; two were revised (one worse than reported, one narrower). The reviewer's frame is the right one for this whole group: the foundations are sound, and the work is deleting the alternate paths around them.
Nothing below R0 proceeds until it closes. No new ant, no memory feature, no broader autonomy, no connectors. R4's live runs wait too — a qualification report is only as true as the instrumentation under it.
- ✅ The window before the first administrator (v0.3.8.91).
/auth/setupwas unauthenticated whileCountUsers() == 0, on a listener every profile forces to0.0.0.0, withoperator_shell_enabledshipping true — reach the port, win the race, get a host shell. Closed bySetupAuthority(single-use bootstrap token, rule written on the BIND not the caller's address), a transactionalCreateInitialAdministrator, and the operator terminal shipping off. TheDEPLOYMENT.mdparagraph that argued the bind was safe is corrected rather than deleted. - ✅ A verification fault fails closed (v0.3.8.91). The catch left
DeterministicBlocknull andApplyUnderBypassgates on exactly that, so a crashed verifier let an unverified patch reach the operator's tree under a Bypass conversation. - ✅ No control decision is read out of prose (v0.3.8.91). A REJECTED patch satisfied
Contains("applied") && !Contains("not applied"), returned HTTP 200 and fired a real git commit. - ✅ One promotion gate (v0.3.8.91).
PatchPromotionGateis the single authority and the Apply button and bypass lane consult it. The actor changes exactly ONE condition — who satisfies the human; a test pins every other condition above the actor switch so no lane can be silently exempted. Found while building it:Task.DeterministicBlockhad no database column, so the block that gates bypass application has never survived a restart. Auto-apply still runs its own nine checks — stricter than the gate on every axis, and folding it in is left as a named follow-up rather than done blind. - ✅ Patch sets apply as a unit on every path (v0.3.8.91).
PatchSetApplypreflights every target, journals before the first mutation, stages each file's pre-state, and rolls the whole set back on any failure. The bypass lane'sforeach (proposal) => apply— which continued past a failure — is gone.AutoApplyRunner's duplicate preflight now delegates to the one in Core. - ✅ The live tree must match the verified tree (v0.3.8.91).
WorkspaceFingerprint— HEAD plus the fullgit status --porcelain -ualllisting — captured when the sandbox is built and compared by the gate before any lane writes. HEAD alone was the first design and would have been wrong: it does not move on an uncommitted edit, which is the case the check exists for. Three states, andNotCaptured(non-git, or a set predating this) is deliberately not a refusal. - ✅ File mutation and database state, one recoverable transaction (v0.3.8.91). An apply intent
journal — Prepared → Mutating → Applied → Recorded, on both lanes, written before the write — plus
startup reconciliation that decides from hashes: prepared discards, applied completes the records,
and a mid-write state matching neither hash is left for an operator rather than resolved in favour
of success. It never re-applies and never rolls back.
✅ The crash-injection matrix (v0.3.8.94) —
Anthill.CrashHelpergained an intent-journal mode that drives the live apply sequence to a chosen phase and blocks;PatchApplyCrashMatrixTestsreally kills it at each window (prepared / mutating-unwritten / mutating-written / applied) and reconciles in the parent process: discard, discard-by-hash, needs-operator-with-intent-open, and the database catching up with the disk — one deterministic recovered state per row. - ✅ A refused lease prevents execution (v0.3.8.91). The claim is taken before anything is
committed and a refusal returns. Committing first was what forced the old code to ignore the
refusal; there is now nothing to strand. A claim the scheduler then declines is released as
Abandonedrather than held. - ✅ The configuration surface agrees with the runtime (v0.3.8.91). The
api_token_envself-referential fallback (sticky, and kept using the variable the operator had abandoned); four??env overrides an empty string won;ANTHILL_PORTunclamped and silently falling back; an unreadable config running on defaults that bind0.0.0.0;config.example.jsonshowing seven roster flags asfalsethat the migration forcestrue, with the two real controls undocumented; andLastConfigMigration, which its own doc comment claimed two endpoints surfaced and neither did.ConfigurationSurfaceTestspins both directions against an explicit undocumented-on-purpose ledger. ✅ The GENERATED schema (v0.3.8.114) —ConfigCatalogis one declaration per key, carrying exposure, security class, env override, range, aliases and section, with the CLR type and the DEFAULT VALUE derived by reflection fromAnthillConfigrather than restated (a declared default that disagrees with the property's initializer is the same defect one layer up).--emit-configregeneratesconfig.example.jsonanddocs/CONFIGURATION.md, andConfigCatalogTestscompares the committed files byte for byte against a fresh render, so drift is a failing build rather than a discovery.AnthillRuntime.EditableConfigKeys— a seventy-line hand-maintained HashSet — is now a projection of the catalog. What the example file taught: it is NOT a dump of the defaults; it is a curated example whose illustrative values differ from them on purpose, which is why the first render deleted 148 lines.ExampleJsoncarries those, and aSecretblanks BEFORE the example is consulted, guarded byNoSafetyGate_IsIllustratedAsEnabledafter four safety gates were found shown enabled against shipped defaults offalse. - ✅ The evidence and artifact vocabulary mismatches (v0.3.8.94).
AntEvidenceKindsis the closed vocabulary for what an ant reports (disjoint from the store'sEvidenceKinds, on purpose); the kind-"tool" filter bothFailureContext.ToolandTaskResult.Toolwaited on for six releases finally has a producer (the registry records dispatched tool names; the measurement boundary turns them into evidence);deterministic_work_completedstopped consulting the wrong witness's vocabulary;patch_jsonandtextare DECLARED transport-only at the artifact bridge instead of falling into the null arm beside typos; andVerificationPolicyReachabilityTestsis the dormant- keys ledger (code_patch_full,config_change,artifact_production— real policies, recorded as dormant, failing the ledger when either fact changes). - ✅ Auto-apply consults the one gate (v0.3.8.94). The Director evaluates every eligible
proposal as
PromotionActor.Automation— the actor v0.3.8.91 declared and nothing used — and its private copies of the evaluation / write-gates / rollback-marker checks are deleted. The fold STRENGTHENS the lane (deterministic block, incomplete or blocking policy review, moved workspace — conditions the runner never checked); what stays runner-side is exactly the set-level part a per-proposal gate cannot own: the evidence content hash, mixed deterministic rows, patch-set identity, whole-set preflight, the durable transaction. - ✅ Enforcement — v0.3.8.112. Warnings are errors across the solution (flipped on a measured
zero-warning tree, not declared over a backlog); the repository scans itself for committed
credentials through
PolicyScan's own rule table; a size ratchet set above the tree as it stood; one version per package across every project; module discovery replacing a hand-maintained list; and the guard hierarchy written down indocs/GUARDS.mdand enforced byGuardHierarchyTests. Analyzers beyond the SDK default and typed database rows are.113— each named in §2c with why. Including the guard hierarchy, which v0.3.8.92 paid for: a runtime black-box test first, then a typed registry, then compiled/Roslyn inspection, and a source scan last. When a source scan IS the right tool it may never depend on a character count — v0.3.8.91 shipped a guard reading a 4,000-character window whose marker sat 27 characters inside it on Linux and outside it on a CRLF checkout, and main went red on a property that had not changed.
Exit gate — ✅ CLOSED at v0.3.8.114. Every write to the operator's tree passes one gate; every externally visible mutation is recoverable after a crash; no security decision reads prose or falls back to a broader source when its authoritative input is missing; and the configuration surface has one authority.
The fourth clause is what
.114closed, and.112should not have implied otherwise. That release's changelog opens "R0's LAST ITEM" — it was not; the generated schema was still open, and.112and.113both shipped without it. A tagged entry is frozen, so the correction is recorded here and in.114's own entry rather than by editing.112's.Standing R0 hygiene continues past the gate and is not folded into it: the typed-row ratchet, central package management,
AnalysisMode, and the four remaining literal-only guards. A closed gate means the four conditions hold, not that there is nothing left to tidy.
R0-B — One real task, end-to-end, through the production path · v0.3.8.93
The corrections brief: make operator request → plan → routed worker → mission worktree → patch/evidence pipeline → real answer true at every joint, not just at most of them. Closed by
substance at v0.3.8.93, with one honest residual:
- ✅ The role's contract clamps the agent CLI —
AgentAccessScopecarriesRoleMayWrite, derived from the ant registry at dispatch; a read-only role gets no Edit/Write/Bash flags or settings under ANY policy. Skip All Approvals skips the operator's prompts, not the role's contract — the promotion gate's sentence, now true one layer earlier. Enforced in process argv and the materialized settings file, never in prompt prose. - ✅ Both coder modes end in one pipeline —
ProcessPatchSetis the single consumer; the harvested worktree diff now gets verification, a patch artifact, approval cards and the bypass gate exactly as a structured-JSON proposal does. The one divergence (review tasks cannot be inserted after the task graph closes) is evented, not silent. - ✅ The prompts tell the truth — the operator request travels in an OPERATOR REQUEST fence and
is the instruction; fetched/prior-model content keeps the UNTRUSTED fence; embedded fence markers
are defanged in both, so a hostile string in a fetched document cannot close its fence or forge
the operator's. Worker prompts name REAL dispatchable tools from the authorization table, or say
"none" —
read_workspace_docsand its phantom siblings no longer reach a prompt. - ✅ Mission size is proportional — a one-task informational plan is a declared, accepted outcome; the three-task minimum and the guaranteed verifier now bind exactly the plans with consequential (patch-producing) work. The guard was split, both halves pinned.
- ✅ One consequential pheromone decision — verified worker trails break the
declaration-order tie in worker selection; capability keywords outrank any trail; evented as
worker_selected_by_trail; A/B-replayed inPheromoneDecisionTests. - ✅ The real result reaches the operator — already true since v0.3.8.73 (
operator_summarycompiled from persisted rows, no model in the path); re-verified rather than rebuilt. - ◻ The live CLI gate has still not run.
CliBoundaryCharacterizationTestsrecords the exact argv/settings per role×policy cell at the pure-function layer, and says in its own header that it is NOT the live gate: no vendor process starts in the suite. The live run remains R4's item, and nothing in this release claims it.
R4 — Protocol-compliant live qualification · 2–5 releases
Only meaningful after R1: without adapter conformance a live failure cannot be attributed.
Ollama, an OpenAI-compatible provider, and Claude Code or another agent CLI; ideally one small local
and one strong cloud model. Record provider and model version, tokens, cost, durations, failure
classes, which trigger reached each role, artifacts produced and consumed, and whether
MissionReconstruction replays the result.
Budget the run at one release and its findings at one to four. The single unstructured live mission at v0.3.8.73 produced three real defects, one of which — the operator report having no compiler — was architectural.
- ✅ The recorder, built and proved before any live run (v0.3.8.89).
LiveQualificationRecordassembles the exit gate's telemetry table out of records the colony already keeps — provenance for the model that actually served each call,model_callevents for tokens and durations, typed failure classes, the consumption ledger for what each role really read, andMissionReconstructionfor whether the run replays.LiveQualificationRecordTestsholds it toQUALIFICATION.md's table one-to-one and drives it against a scripted mission, so the live run is an operator pressing go rather than a live run plus an argument about whether its telemetry was complete. - ✅ Cost has a producer: the operator (v0.3.8.90).
ModelPricingconverts the tokens the runtime already measures against amodel_pricingtable inconfig.json—provider/model, orprovider/*for a whole provider, which is how a local run reports a MEASURED zero instead of an unknown. A pure function over a table passed in, not a reader of statics, and deliberately not inside the recorder: this plan's own wording for the change. The refusals are the safety property, and there are three, distinguishable because they are three different things to do: no table configured, a provider that reported no usage, and a served model the table does not cover. A run is priced only when EVERY model has an entry and EVERY call reported usage — a partially priced run is not priced, because a total lower than the run's real cost wearing a currency symbol is worse than an absent figure. Nine tests inModelPricingTests, most of them about the refusals. - ◻ The runs themselves. Ollama, an OpenAI-compatible provider, and an agent CLI.
Exit gate. A recorded run per provider with complete telemetry, and
QUALIFICATION.mdupdated from "never happened" to the run's own evidence. The recorder is ✅ at v0.3.8.89 and cost ✅ at v0.3.8.90; the gate is now open on nothing but the runs themselves.
R5 — Public-QA readiness · 2–4 releases
Pulled forward from the old plan's tail, because outside testers are a stated goal and a stale tutorial fails a new user the same way a stale handoff failed a new session.
Windows, Docker and LXC/Linux; fresh install; upgrade and migration; restart recovery; diagnostics
and redaction; tutorial accuracy; the complete QA-CHECKLIST.md walked end to end by someone who did
not write it.
Exit gate. A person who has never seen the repository can install, run a mission, and file a report using only the shipped documents.
R6 — Execution sandbox · 3–6 releases · gate for R9
The plan still admits a filesystem TOCTOU window, dotnet on the shell allowlist as arbitrary
workspace code execution, and incomplete Windows junction testing. Containers or seccomp on Linux and
a job-object or equivalent on Windows.
This is a gate before unattended coding, not a follow-up to it. Autonomy on top of an escapable sandbox enlarges the blast radius without improving the system — §5's own argument, applied to the one boundary it left open.
Exit gate. A hostile patch cannot escape the workspace on either platform, proved by test.
R7 — Finish the typed channel · 4–7 releases
- Authoritative task inputs everywhere. The scheduler derives inputs from dependencies, required schema types, the current revision, artifact generation and policy — never "an artifact exists somewhere in this mission".
- The last of prose as the primary channel.
Task.Resultis still central and the builder still produces free prose.MissionReport(v0.3.8.73) was the first half. - Complete artifact provenance. Mostly filling fields that exist.
- Research citation quality. Scenario 1 has admitted since v0.3.8.57 that a brief citing its sources badly still parses.
Exit gate. No downstream decision reads
Task.Resultprose where a typed artifact exists.
R8 — Graduation, reputation routing, memory · 6–12 releases
- The per-role graduation record, complete — only fillable after R3.
- Reputation-aware routing. The first item where behaviour changes based on the colony's own history: new failure modes rather than newly-discovered ones.
- Semantic and procedural memory. The widest range here. Memory that influences future missions is where a wrong answer compounds instead of failing.
Exit gate. A role's history changes what it is asked to do, and a wrong memory can be traced to the mission that produced it.
R9 — Sandboxed autonomy and the PR lifecycle · 4–8 releases
Everything above, pointed at a real repository with a real pull request at the end. Requires R6.
Exit gate. An unattended mission opens a PR a maintainer would accept, and every refusal on the way is attributable.
R10 — Connectors, self-improvement, production hardening · 8–15 releases
Five programs, not one item. The old plan's item 14 concealed this:
- a generic connector framework;
- safe self-improvement;
- production qualification — SLOs, alerting, runbooks;
- a thirty-day soak;
- external security review, SBOM and signing, backup/restore and migration drills.
Exit gate. Thirty days unattended with no unexplained incident, and an external review with no open P0.
Ongoing — technical cleanup
✅ The event vocabulary is complete and consumed (v0.3.8.86 — 67 emitted names added, 2 phantoms removed, both directions enforced by
EventVocabularyTests). Publishers still pass literals rather than the constants; that conversion is a separate, larger change and the guard makes the drift it would prevent impossible to widen in the meantime — a new literal must be declared to pass.✅ One catalog declares what a role may do (v0.3.8.87 —
ToolCatalogremoved,CapabilityDeclarationTestsadded).AntExecutionCatalogandToolCatalogboth declared each role's capabilities and side effects; only the first was enforced, and the second was read by the gate that decides whether a planned task may enter the queue. They disagreed about capabilities for four roles, about side effects for two, and about six roles the second did not list at all — which is where v0.3.8.76's deletedmodel.invokelie survived. The pre-execution permission check that lived beside it,ToolCatalog.CanRun, had no production caller in its whole life; its one caller was a test that built both sides itself.◻ The
Capabilityvocabulary is half-unwired, and now says so. Seven of the fourteen names are granted by nothing and required by nobody.repo.patch.applyis withheld on purpose and always was; the Proxmox, infrastructure and credential names belong to a module surface that authorizes throughActionExecutorinstead. Each is recorded inCapabilityGrant.DeliberatelyUngrantedwith its reason, and a new capability can no longer join the vocabulary and quietly reach nobody. Wiring or removing them is R6/R10 work, not cleanup.◻ The rest of the v0.3.8.90 sweep, with its sites. Sweeping "a filter that could not match" across every vocabulary in the tree found ten instances. Six are closed in v0.3.8.90 (the four builder handoffs, the failure-event set, the console's mission notifications, the autonomy dedupe, the category switch, and the diagnostics counters). Four are recorded here rather than half-fixed, because each needs a decision and not an edit:
AntEvidence.Kind == "tool"matches nothing —SqliteMemory.TaskResults.cs:127andExecutionService.cs:1710. No ant emits atoolcitation (the kinds arecheck,file_path,policy_rule,mission_id,failure_id,revision,workspace,failure_signature), soArtifactProvenance.ToolandFailureContext.Toolare null in every row ever written — reading as "no tool was involved" rather than "unknown". The decision is which side is wrong: producers should cite the tool, or the field should go.ArtifactSchemas.ForAntKindhas no arm forpatch_json—Evidence.cs:248. The coder emitspatch_json; the switch has an arm forpatch_set, which nothing produces, and six other dead arms. The coder's artifact therefore maps to null andBridgeArtifactsToStoreskips it. It is NOT simply a rename:ExecutionService.cs:951alreadyPuts the patch set directly, so adding the arm may double-store. That has to be understood before it is touched. Related:AntExecution.cs:358declares the coder'sProducedArtifactTypesaspatch_set— a name the coder does not produce.EvidenceKinds.Reproducible.Contains(e.Kind)—ExecutionService.cs:1479— tests ADR-004 verdict kinds against citation kinds. The disjunct is dead; onlye.Kind == "check"ever fires. The file one directory over states the distinction being violated: anAntEvidenceis a CITATION, ADR-004 evidence is a VERDICT.VerificationPolicykeys no task type can reach —Verification.cs:79-82, 110-111:code_patch_full,config_change,artifact_production, and the aliasesdocs_updateanddocumentation. Lower confidence than the others (they are table keys, and the type is public), butDeterministicBlockTestsasserts against one of them directly — the same shape the file's own v3.8.21 note says caused a defect.
◻ The emitter detector still has a blind spot, and it is wider than its comment admits. v0.3.8.86's sweep reads event names handed to
LogEventas a literal first argument. A name passed through a wrapper is invisible — v0.3.8.89 named that — and so is one built by a ternary (Queen.cs:1074emits all three mission terminals that way) or one whose first argument contains a call, which the[^,()]+in the pattern rejects. v0.3.8.90 declared six of the affected names by hand because two consumers needed to reference them; roughly a dozen remain emitted and undeclared. Widening the detector is the fix, and it is a sweep of its own rather than a line in another release.◻ The operator's configuration surface disagrees with itself. A second sweep, over config rather than vocabularies, found: 25 parsed keys
config.example.jsonnever documents — includingroster_profileanddisabled_roles, which are the ONLY working off-switches for the seven specialist ants the example file shows asfalseand the roster profile then forces totrue;config_version,logs_dirandexports_dir, which are documented and read by nothing;AnthillRuntime.LastConfigMigration, whose own doc comment says two endpoints surface it and neither does; nineRuntimeOptionsfields nobody reads, against that file's own stated rule; anapi_token_envwhose fallback is the static's prior value rather than a config value, so redirecting it to an unset variable silently keeps authenticating againstANTHILL_API_TOKEN; and three env overrides that use??, so an empty string set in a compose file wins over the file. This is a release, not a cleanup line — the token precedence is security-adjacent and the roster one is a safety claim the file gets wrong. v0.3.8.90 fixed only the piece it touched (ResetConfigsilently discarding the operator's priority route, now also the price table).
Not a tail. Fully async execution, ~166 runtime statics, VRAM scheduling, multi-platform QA,
event-loss accounting and deployment verification are independent workstreams. The statics in
particular have caused three defects in this session alone (a leaked UseOllama, two roster-gate
leaks), and they get harder to remove as more code depends on them. Take slices opportunistically;
some belong before R9 rather than after.
3. Acceptance gates
Non-negotiable. The colony is not a twelve-role colony until all of these pass. 12 of 12 (v0.3.8.80).
- ✅ All twelve roles report Ready under the full profile (v0.3.8.80)
- ✅ Every enabled role has a handler, contract, real production trigger and typed output (v0.3.8.80)
- ✅ A compile-breaking proposed change fails when built in the patched mission workspace (v3.8.23)
- ✅ That failed patch cannot become
completed_verified(v3.8.22) - ✅ A Soldier block cannot be overridden by model text (v3.8.22)
- ✅ Tester failure triggers exactly one bounded Medic repair and a mandatory retest (v0.3.8.57)
- ✅ A UI change cannot reach Coder without a valid
ui_map(v0.3.8.57) - ✅ Scribe and Archivist cannot act positively on unverified work (v0.3.8.57)
- ✅ Archivist runs only after the persisted canonical evaluation exists (v3.8.26 / v0.3.8.41, pinned at v0.3.8.57)
- ✅ Replaying artifact IDs reconstructs every role's inputs and evidence (v0.3.8.57)
- ✅ No mission ant can dispatch shell, direct file-write, or primary-workspace patch tools
- ✅ Disabled or unavailable roles never receive negative reputation for not running (v3.8.26)
4. What "done" would mean
A colony that plans, changes code, verifies with evidence, refuses on reproducible grounds, applies what it has proved, remembers what it learned, and can be handed to someone who did not build it — running unattended without an unexplained incident, on a sandbox a hostile patch cannot leave.
Every clause above is a gate in §2. None is a matter of opinion.
5. The security review — CLOSED, kept as the record
S1–S7 all closed (v0.3.8.59–.65). This section is history now, not a queue, and it sits below the forward plan for that reason. It is kept in full because the findings and the sweeps they prompted are the clearest worked examples of this repository's defect classes — the review named two confinement sites and the sweep found six — and because the RESIDUALS recorded in each subsection are real and feed R1 and R6 above.
An external source-level review found four P0 and two P1 defects. They took priority over everything else, because they were about existing autonomy being trustworthy rather than about the colony doing more. Shipping more autonomy on top of a broken confinement boundary enlarges the blast radius; it does not improve the system.
The biggest benefit is not additional autonomy — it is making existing autonomy trustworthy: no workspace escapes, no silent secret disclosure, no partial trees described as rolled back, and no database failure turning into permission to write.
What was reviewed
| Release reviewed | v0.3.8.57, commit c62a27a |
| Main at review time | 527b4a7 — its only post-release changes are the PR #11 documentation reconciliation, so the runtime code is identical |
| CI | green on both the release and current main |
| Issue tracking | none — no open GitHub issues track any of these findings |
| Method | source-level review; the reviewing environment had no .NET SDK, so nothing was executed locally |
The last two rows matter. Green CI is not evidence against these findings — every one of them is a path the suite does not exercise, and two of them (S2, S5) are places where a test asserts something adjacent to the claim and passes. And with no issues open, this section is the only record.
Immediate containment, until the P0s close
{
"autonomy_autoapply_enabled": false,
"patch_application_enabled": false,
"file_writing_enabled": false,
"file_tools_enabled": false,
"shell_tool_enabled": false
}
Also deny or restrict /projects/{id}/file and /projects/{id}/files at the proxy or API layer.
The Files-pane endpoints do not consult the runtime write flags, and their READ route is
escapable on its own — so the flags above do not contain S1 by themselves.
The findings
| Priority | Finding | Worst consequence |
|---|---|---|
| P0 | Files-pane and workspace confinement can be escaped | Read/write files outside the selected workspace |
| P0 | Auto-apply is not actually atomic | Partial/truncated tree, with logs claiming rollback succeeded |
| P0 | Verification and evidence fail open | Unverified patches reach live auto-apply |
| P0 | Secret artifacts can be sent to models | Credential / private-data disclosure |
| P1 | UI-map enforcement fails open | UI code dispatched without a trustworthy map |
| P1 | Some subprocess timeouts are ineffective | Hung worker or director; unbounded output/memory |
Repair order
This is the reviewer's order and it is not the severity order. Confinement comes first because every later fix is verified by reading and writing files, and evidence comes before transactional apply because a correct transaction around an unverified patch is a reliable way to ship the wrong bytes.
- ✅ S1 — Files-pane traversal and symlink-safe confinement (v0.3.8.59 — one resolver,
PathContainment; TOCTOU remains, see below) - ✅ S2 — shell tool confinement: disable, or fix (v0.3.8.59 — arguments contained;
dotnetresidual recorded) - ✅ S3 — evidence fail-CLOSED (v0.3.8.61 — verifier:
verification_unavailable; auto-apply: five refusal arms; see below) - ✅ S4 — transactional patch application and durable recovery (v0.3.8.62 —
ApplyTransaction; see below) - ✅ S5 — Secret-artifact filtering (v0.3.8.63 —
IsModelReadableallowlist; WITHHELD reporting; see below) - ✅ S6 — UI gate (v0.3.8.64 — throwing store refuses;
{}no longer conforms); remaining subprocess handling ✅ v0.3.8.59 - ✅ S7 — runtime and fault-injection tests before auto-apply is re-enabled (v0.3.8.65 —
ApplyTransactionTests,EvidenceFailsClosedTests,SubprocessHangTests; see S8 below. This line sat unticked for ten releases while its own body recorded the work as done — documentation drift thatDocumentCurrencyTestscannot see, because it names no version.)
S1 — Filesystem confinement (P0) ✅ v0.3.8.59
Broken in two independent ways, either of which is sufficient.
- The Files pane checks
full.StartsWith(root, StringComparison.Ordinal)with no separator requirement. A project at/srv/projecttherefore serves../project-secret/key.txt, which resolves to/srv/project-secret/key.txt— a SIBLING whose name merely starts with the root string. The vulnerable helper feeds read, create and edit alike:ApiHost.Providers.csL855–980. WorkspacePathGuardusesPath.GetFullPath, which removes..but does not resolve symlinks or Windows junctions. A link inside the workspace pointing outside it passes containment and is then followed by the file tools and the patch applier:WorkspacePathGuard.csL63–78. Worse,RepositoryIndex.csL232–249 CLAIMS symlinks are resolved while the guard does not — a declaration that disagrees with the runtime, in the security boundary.- With
shell_tool_enabled, confinement is weaker still:cat,findandgrepaccept unrestricted absolute paths, and setting a working directory does not sandbox a process:ShellAndWebTools.csL31–56.
Fixed in v0.3.8.59. Anthill.Core.Security.PathContainment is the one resolver. It requires
exact-root equality or root-plus-separator, and it walks the path from the volume root resolving
EVERY component through its own chain of links, bounded at 40 hops so a cycle is a refusal rather
than a hang. Components that do not exist cannot be links and are appended literally, which is what
lets a file be created at a path whose parent is real. WorkspacePathGuard.ResolveSafePath delegates
to it, so all twenty call sites behind it are covered at once.
The review named two sites; there were six. The sweep the fix prompted found the identical
missing-separator comparison in PatchVerifyRunner, SandboxWorkspace.Harvest and
Verification.Verify, and a separator-correct but link-blind one in PatchSetMaterializer. The
Verification copy was the worst: it hashes a required artifact as EVIDENCE, so a link or a
sibling-prefixed path meant a hash recorded as proof of a file inside the workspace could be of a
file outside it. PathContainmentTests now carries a detector keyed on the ROOT side of the
comparison rather than the variable name — the first draft keyed on the variable and found only the
two already known, which is what a detector written around the examples in hand always does.
RepositoryIndex's comment is now true. It claimed a symlink out of the workspace "resolves
outside the root and is refused here". It never was. Deferring to the guard was the right call for
the right reason; the guard just did not do the thing the comment credited it with.
STILL OPEN — the TOCTOU race. This is resolution-time containment. A component swapped for a
link BETWEEN the check and the caller's open is not closed, and cannot be with the APIs .NET exposes
portably — it needs handle-relative, no-follow syscalls (openat with O_NOFOLLOW). The window is
narrow and requires an attacker already able to write inside the workspace. Recorded here rather
than described in the code as handled.
Tests. PathContainmentTests: sibling-prefix traversal, absolute and relative link targets,
intermediate-component links, links pointing back inside, a root that is itself a link, link cycles,
non-existent leaves inside and outside, and the guard enforcing the same boundary. The link tests
probe for the privilege to create a link and skip without it — Windows needs Developer Mode or
elevation — so on such a machine the link half is unverified and the sibling half still runs. Linux
CI covers both. Windows junctions are covered by the same LinkTarget path .NET uses for symlinks
but are not separately exercised; that gap is real and small.
S2 — Shell tool confinement (P0) ✅ v0.3.8.59
Called out separately by the reviewer because it had its own remedy — the tool can be disabled
outright — and because the defect is a category error rather than a bug. WorkingDirectory decides
where RELATIVE paths resolve and confines nothing; cat /etc/passwd, grep -r secret / and
find / -name '*.key' ran exactly as written. The nine-command allowlist says WHICH PROGRAM may
run and nothing about what it is pointed at, and it was being asked to do a sandbox's job.
Fixed in v0.3.8.59. Every path-like argument resolves through PathContainment before a process
starts. An argument counts as a path if it is rooted, contains a separator, or contains ..;
--flag=value is split so a path on the right of the equals is checked rather than skipped. Bare
tokens are left alone, so grep -r secret . still searches for the word. The command now runs in
EffectiveRoot rather than Root — inside a mission the workspace is a disposable tree, and the
old value pointed every shell command at the live checkout the mission exists to stay out of.
Beyond the review: find -exec, -execdir, -ok, -okdir, -delete and the -fprintf
family are refused. find . -exec rm {} ; passes every containment check because the path IS the
workspace — the flag is what runs the other program, which is the same question the review asked
about paths applied to arguments.
STILL OPEN — dotnet is arbitrary code execution. It is on the allowlist deliberately, for
build and test, and dotnet run executes whatever the workspace contains. Argument containment
cannot address that; it is governed by shell_tool_enabled, which is off by default. A real sandbox
(container, seccomp profile, job object) is the correct answer and is not reachable in-process across
the three platforms the colony supports.
Tests. ShellConfinementTests — absolute paths outside the root for each command the review
named, relative traversal, paths hidden in flag values, execute/delete flags, bare tokens that must
NOT be treated as paths, and paths inside the workspace that must still work. The gates fixture
opens the shell deliberately: with it closed every one of these would pass, refused by the enable
flag before containment was consulted, which is the adjacent-question defect in its purest form.
S3 — Verification and evidence fail OPEN (P0)
Two consecutive fail-open boundaries, and the direction is the wrong one — a store failure WIDENS authority instead of stopping a live write.
- A verifier that cannot read evidence returns null and falls through to model or static prose:
Ants.csL1001–1051. The static fallback can emitVerification Passedfrom completed task counts alone: L1157–1165. - Auto-apply that cannot read evidence returns zero refusals and continues. It also
deliberately accepts missions with no revision-identified evidence, and skips proposals with a
null patch-set id:
AutoApplyRunner.csL286–329.
Fix. An unavailable evidence store must produce verification_unavailable and never prose
fallback for production verification. Live auto-apply must require a non-empty patch-set id and
complete revision identity; the complete verification bundle for the exact revision and tree; at
least one deterministic pass; no deterministic failure; every policy-required check. Compare the
patch-set CONTENT hash as well as revision id and tree hash. Legacy unidentified evidence stays
readable for history and is manual-apply only.
Tests. Evidence-query exceptions, no evidence, legacy evidence, null patch-set id, mixed revisions, mixed pass/fail rows, wrong tree hashes.
This subsumes §2 item 2, which asked only that the canonical evaluator consume Evidence.Judges().
Closed v0.3.8.61. The verifier distinguishes a store that FAILED from a store that never
existed: failure produces verification_unavailable — a verdict Parse cannot emit and IsPass
never accepts — while the no-store CLI/test configuration keeps its static contract. Auto-apply's
gate (RefuseEvidenceAboutAnotherRevision, now taking IEvidenceStore so tests drive the real
function) refuses on: store read failure; no revision-identified evidence (legacy rows are
manual-apply only); missing patch-set id; evidence not deterministic-and-passing for the exact
revision and tree; any deterministic FAILURE for the revision (a pass cannot outvote it); and a
patch-set CONTENT hash mismatch — the evidence judged bytes, so the gate compares bytes, which
also makes a policy-filtered subset self-refusing. EvidenceFailsClosedTests covers the full
test list above behaviourally. Residual: "every policy-required check" is enforced upstream by
the canonical completed_verified evaluation auto-apply already requires; it is not re-derived
inside the gate, deliberately — one authority (MissionVerification) owns that rule.
S4 — Transactional patch application (P0)
The v0.3.8.57 "a patch set applies as a unit or not at all" guarantee does not survive a mid-write failure.
ApplyPatchToolbacks up and then performs destructive I/O, but its outer handler returns only an error — the backup and path metadata are lost. AWriteAllTextthat truncates or partially creates a file before throwing is therefore unrecoverable: L75–95, L156–198.AutoApplyRunnerrolls back only the EARLIER successful patches, never the operation that failed mid-write. It ignores every rollback return value and logs the whole batch as rolled back regardless: L176–204, L243–258.- Rollback itself can destroy newer work: it deletes added files and overwrites modified or renamed
ones without checking whether they changed after apply. Manual revert has the same behaviour:
Queen.Views.csL173–226, L268–318.
Fix. Stage writes into temporaries and atomically replace or move. Write a durable transaction
journal before the first mutation, and recover incomplete journals at startup. Return recovery
metadata even when the current operation fails. Record pre- and post-apply hashes, and roll back
only where the current bytes still match what was applied. Treat incomplete rollback as a critical,
durable rollback_failed state that halts auto-apply.
Tests. Injected disk-full, permission change, partial write, rename, rollback failure,
concurrent edit, process crash — asserting a byte-identical restored tree.
AutoApplyAtomicityTests.cs L141–155 currently asserts only that the SOURCE contains a rollback
call. That is a check answering a question adjacent to the one asked, in the file whose whole
purpose is to prove atomicity.
Closed v0.3.8.62. ApplyTransaction (SDK): journal durable before the first mutation; staged
atomic writes (a target is never half-written); hash-checked rollback that preserves and reports
newer work as conflicts; a durable ROLLBACK_FAILED marker that halts auto-apply until an
operator clears it; startup recovery replaying interrupted journals under the same rule. The tool
reports recovery metadata on failure and applied_hash on success; the runner journals the batch
and believes the rollback report; manual revert applies the same hash gate; the un-journaled
RollbackAutoApplied is deleted rather than left to drift. ApplyTransactionTests covers the
fault list behaviourally, byte-identical assertions included; the adjacent-question source scan is
replaced. Residual: disk-full and permission faults are injected at the transaction's write seam
rather than by filling a real volume — the seam fires between the temp write and the atomic swap,
which is the worst real moment; and legacy patches applied before applied_hash existed keep the
old unchecked revert behaviour, stated in the revert reply.
S5 — ArtifactVisibility.Secret does not prevent model disclosure (P0)
Artifact.cs L63–71 states that a Secret artifact is "never rendered, never sent to a model".
Nothing enforces it:
- mission queries return every visibility:
SqliteMemory.Artifacts.csL146–158; ArtifactContext.Compiledoes not filter Secret and emits their payloads, including declared inputs:ArtifactContext.csL98–143, L168–205;- those blocks are appended directly to model prompts:
DomainHelpers.csL153–162; - the soldier reads payloads directly with no visibility check:
SpecialistAnts.csL325–354.
Built-in producers currently write mostly Colony or Operator, so exploitation needs a module, a custom producer, or a corrupted/imported row. That is not much comfort: the public SDK exists to support modules, and malformed visibility is deliberately coerced TO Secret — so the unsafe value is precisely the one a malformed import lands on.
Fix. An audience-aware retrieval and render policy. Secret artifacts never enter any model context or narrative renderer. A declared Secret INPUT is reported as WITHHELD rather than silently omitted — a silent drop is how a role reasons confidently about a premise it never received. Apply the check again at every direct consumer.
Tests. Prioritized, declared, and corrupt-visibility Secret artifacts.
Closed v0.3.8.63. Artifact.IsModelReadable is the one definition, and it is an ALLOWLIST
(Colony or Operator) so an out-of-range enum value fails closed — the coercion of malformed
visibility TO Secret finally means something. The context compiler removes Secret payloads from
mission-wide blocks (unadvertised) and reports declared Secret inputs as WITHHELD by id and
schema, never by content; the soldier's direct read applies the same check and names what it
withheld. SecretArtifactTests covers the review's three cases. Residual: the enum doc also says
"never in an API response" — the API surfaces return artifact METADATA through their own routes
and were not part of the four defect sites; a sweep of those routes belongs with S6's UI work.
S6 — UI-map gate fails open, and {} is a valid map (P1)
UiChangeGate.Check allows when the artifact store is absent or throws: L88–107. That was a
deliberate choice — a missing store is evidence about the WIRING rather than the mission, and
failing closed would block every CLI and test caller — but production dispatch always has a store,
and the two cases are distinguishable. Separately, the ui_map schema requires no keys, so {}
conforms: ArtifactSchemaCheck.cs L121–124. UiChangeGateTests proves a truncated map is refused
while an empty one passes.
Fix. Fail closed on production dispatch when the store is unavailable. Require an intact map
from ui_cartographer carrying files_examined, routes and api_calls.
Closed v0.3.8.64. The gate distinguishes the absent store (CLI/tests — permissive, evidence
about the wiring) from the throwing store (an incident — refuses, naming the outage), the same
two-state repair the verifier received in S3. The ui_map schema requires the three keys the
cartographer has always emitted unconditionally, so {} no longer conforms while an honest empty
map (routes: []) still does. S5's API residual was swept and closed by observation: no
Anthill.Api route serves artifact payloads at all, so "never in an API response" holds vacuously;
any future artifact-serving route must filter on IsModelReadable and this note is its warning.
S7 — Subprocess timeouts that cannot fire (P1) ✅ v0.3.8.59
ShellCommandTool and RepoOps.Git call synchronous ReadToEnd() before
WaitForExit(timeout). A process that never exits therefore never reaches the timeout, and
sequential stdout-then-stderr reads deadlock when the other pipe fills: ShellAndWebTools.cs
L44–56, RepoOps.cs L27–51.
v0.3.8.57 fixed five git sites to kill their process trees on timeout and did not fix the read that prevents the timeout being reached — the guard was added upstream of the thing that blocks it.
Fix. Concurrent asynchronous draining, bounded output, cancellation, process-tree termination.
Half done in v0.3.8.59, unavoidably. ShellCommandTool's reads and its timeout are the same
method S2 had to change, so leaving the ordering broken there would have meant shipping a security
fix into a method that still hangs. Both pipes now drain concurrently, the wait bounds the whole
thing, the kill takes the process tree, and output is capped at 20,000 characters — find over a
large tree previously returned everything, into a ToolResult, into an artifact, into a prompt.
RepoOps.Git closed too. Same fix: both pipes drained concurrently, the wait bounds the whole
call, the process tree is killed. Worth naming what it looked like — v0.3.8.57 added
Kill(entireProcessTree: true) to the very line below the reads and did not touch them, so the
colony spent two releases with a correct kill on an unreachable path. A guard placed downstream of
the thing that hangs.
git clone on a large repository is the concrete deadlock: it writes progress to stderr
continuously while this side drains stdout, each waits for the other, and neither is timed out.
STILL OPEN — the behavioural tests. A child that writes heavily to BOTH streams and one that
never exits are not written. ShellConfinementTests pins the ORDER (no synchronous read between
start and wait), which is the defect's shape rather than a proof that the fix survives a real hang.
Closed v0.3.8.65. SubprocessHangTests runs the real things: a git that genuinely never exits
(a pre-commit hook that sleeps, against a shortened test-seam timeout) proves the timeout FIRES
and the call returns bounded; a hook that writes ~130KB to BOTH streams proves the sequential-read
deadlock is gone; and a find over four thousand files, through the production ShellCommandTool,
proves the flood drains concurrently and the output cap holds. POSIX-only by early return — the
children are shell scripts, and CI and both operator gates run them.
S9 — The colony asserts its roles through a channel that carries no authority (P0-adjacent) ✅ v0.3.8.59
Found in the field, v0.3.8.59, and it is the release's own defect class one layer out. With every message now a mission, the colony's role prompts reach an agent CLI and the agent REFUSES them as a prompt-injection attempt. It is right to.
AgentCliProvider.Flatten collapses every ModelMessage into one string, prefixing non-user roles
with a literal [system] text header, and hands the result to -p "{prompt}". -p is a USER
TURN. AgentCli has PromptArgs, StreamArgs, AcceptEditsArgs, AutoApproveToolArgs,
BypassArgs, AddDirArgs and LocalSettingsRelativePath — and no system-prompt flag at all.
So what the agent receives is a user message that assigns it a persona, cites mission IDs, asserts
tool permissions the session does not have, demands a fixed output format, and carries a line of
prose that says [system]. That is the signature of an injection, not a resemblance to one. A model
that complied would be a model that does whatever any user turn claiming to be a system tells it.
The direct agent lane never hit this because the operator's words arrived as what they were: a person asking a question. Nothing was impersonating anything. Deleting that lane was still right — but it exposed that the colony's authority over its workers was never carried by anything except prose.
Fix. Use the real channel. Claude Code has --append-system-prompt and --system-prompt (and
--system-prompt-file), all valid alongside -p. Add SystemPromptArgs to AgentCli; route the
role contract there and ONLY the operator's actual task through -p; delete the [system] text
header, which exists solely because the system channel was missing. An agent with no such flag either
keeps the flattened form and its refusals, or is not routed roles that need a persona — recorded
per agent rather than assumed uniform.
And stop laundering operator text as colony framing. The field report showed an operator's off-topic question surfacing as a project's "stated purpose". Operator text must be labelled as operator text; presenting it under a heading that claims it is something else is the same defect pointed inward, and it is what makes a legitimate prompt read as a fabricated one.
Fixed in v0.3.8.59. AnthillRuntime.PromptInjectionPrefix is deleted. RoleSystemPrompt(role, mission) replaces it on the SYSTEM channel and says where it comes from — "this message is your
operating contract and comes from the harness itself, not from the person who wrote the request" —
which is a claim the transport now makes true. UntrustedBlock(label, text) fences the spans that
genuinely are untrusted, with paired delimiters and a subject.
GenerateTyped takes system:, composing a System + User pair; null keeps the old single-message
shape so an unconverted caller loses nothing. AgentCli.SystemPromptArgs carries the contract to an
agent's own flag — Claude Code's --append-system-prompt, appended rather than replacing so the
agent keeps its own tool guidance and safety instructions. AgentCliProvider.Flatten became Split;
the [system] literal is gone, and an agent with no such channel folds the contract into the prompt
plainly rather than impersonating a system header.
All eight model-calling roles now send a contract. The scribe is worth noting: it never carried the old prefix, so it was the one role NOT sending an injection-shaped prompt — and also the one sending no operating rules at all. Same gap, opposite symptom.
Tests. RoleContractChannelTests — no source anywhere asserts a system boundary from inside a
prompt; the contract names its own origin; the untrusted block fences only what it labels; every
GenerateTyped call passes system:; the catalog appends rather than replaces; an empty contract
sends no flag (a blank --append-system-prompt "" reads as an instruction to have no contract); and
the contract travels as discrete argv, never shell text.
THE TALKING POINTS WERE CONDITIONAL FACTS ASSERTED UNCONDITIONALLY, which is worse than either
"true" or "false" and is why a worker refusing to vouch for them was right. Both features exist:
missions_fts USING fts5 is created in SqliteMemory.Schema, and EnableParallelExecution /
MaxParallelWorkers drive a TaskScheduler that reads DependsOn. But both are RUNTIME-CONDITIONAL.
FtsAvailable is a mutable flag the memory layer sets to false when SQLite throws
(catch (SqliteException) { FtsAvailable = false; }), so on an install without FTS5 the claim is
false — and SelfTest already reports exactly that: "FTS5 not available; keyword fallback in use".
EnableParallelExecution is an operator toggle, surfaced in Queen.Views as Parallel Execution: {flag}.
So the colony KNEW the answer at runtime and asked the model to assert it from prose instead. That is this repository's own recurring shape — a claim derived from a sentence rather than from the state that knows — pointed at the operator reading the answer.
The strongest evidence it was a known hedge that got lost: the non-LLM FallbackResponse in the same
class says "uses FTS5 WHEN AVAILABLE". The deterministic path was more truthful than the instruction
given to the model.
Related and NOT fixed: that same FallbackResponse still opens with "1. Review patch proposals
using /patches…" unconditionally, on missions that produced no patches. Same untruth, deterministic
rather than generated, so no model will ever flag it.
COMPLETED in the same release. All six persona-bearing prompts converted — builder, coder,
verifier, planner and strategist by name; researcher and web carried only the banner. Operator text
(mission goal, prior task output, standing objective) is fenced with UntrustedBlock. The
strategist's objective matters most: it is text an operator wrote that the colony re-reads unattended
on every run, which makes it the highest-value place in the colony to plant an instruction — authored
once, obeyed forever, with nobody watching that turn.
The builder's FallbackResponse no longer opens with "Review patch proposals using /patches" on
missions that produced none. Same untruth as the deleted talking points, deterministic rather than
generated, so nothing downstream would ever have flagged it.
STILL OPEN. Only Claude Code has a verified system-prompt flag; Codex, Gemini, Aider and OpenCode are declared as having none and fall back to folding. That is recorded per agent rather than assumed uniform, but it means the fix is partial for four of five agents until each flag is confirmed.
S8 — Re-enable ✅ decision recorded v0.3.8.65
Fault-injection tests land before auto-apply is switched back on. §2 resumes after that.
The precondition is met. Every rung of this ladder is closed: confinement (S1/S2, .59),
evidence failing closed (S3, .61), transactional apply with durable recovery (S4, .62), secret
filtering (S5, .63), the UI gate (S6, .64), subprocess hangs proven survivable behaviourally
(S7, .59 + .65), and the authority channel (S9, .59). The fault-injection suites the reviewer
required exist and run on every gate: ApplyTransactionTests (mid-write faults, crash recovery,
concurrent edits, vanished backups — byte-identical restores or durable halts),
EvidenceFailsClosedTests (throwing stores), SubprocessHangTests (real hangs and floods).
Re-enabling is an OPERATOR act, and this is its checklist. autonomy_autoapply_enabled stays
off in the shipped defaults; an operator turning it on should verify, in order: (1) this section's
ladder shows every rung ✅ in the running version; (2) no ROLLBACK_FAILED marker exists under the
workspace's .anthill/apply-journal/ — the runner refuses while one does; (3) both write gates
(patch_application_enabled, file_writing_enabled) are deliberate choices, not leftovers;
(4) autonomy_autoapply_verify_cmd names a check the deployment can actually run, or the
break-glass keep-without-verify is a knowing, logged choice; (5) the auto-apply path allowlist
names only trees whose partial states the operator could tolerate diagnosing. The colony enforces
(2) itself and logs the rest; the list exists so the decision is made once, with eyes open, rather
than discovered in fragments during an incident.
6. The record — the shape of the mistakes
Kept because the shape recurs and recognising it is worth more than any individual fix.
A check that answers a question ADJACENT to the one asked, and passes. Found fifteen times. The newest: a graduation record cited two real cancellation test files that prove real things and name no role, and a qualification index lived in a doc comment where a citation could rot into a deleted file without anything noticing.
Declared, and reaching nobody. RequiredInputArtifactTypes, EvidenceKinds.SchemaValid and
Task.InputArtifactIds were each declared before anything populated them, and each looked exactly
like a working feature for releases. Evidence.Judges joined them inside a single release — added
and read by nothing until the same release closed it. v0.3.8.86 found the vocabulary case: two event
constants nothing published, both NEAR-MISSES of real event names, so a subscriber filtering on
either compiled, ran, and matched nothing forever. v0.3.8.87 found the version with a passing test
attached — ToolCatalog.CanRun, a pre-execution permission check with no production caller in its
whole life, whose only caller was a test that built the descriptor AND the grant set itself. The
sharper reading is the one FailureClassNames had already written down about a different bug: no
test anywhere ran a value from a real producer into a real consumer.
A filter that could not match, guarding the property that mattered most. v0.3.8.89. Five
assertions queried memory_candidate_archived; the ingest emits memory_candidate. One of the five
was the cancellation harness's "no memory survives a stopped mission" — the property R3's header
singles out as the one that outlives the mission — and it was checking for an event no producer
writes. v0.3.8.85 wrote the sentence that names it ("held by luck rather than by design") and
attributed it to the archivist usually having nothing to propose; the real reason is that the filter
matched nothing. Found from the CONSUMER side: an event type queried by name is a position that can
only be an event type, and four of eighteen such names were declared nowhere — a blind spot in
v0.3.8.86's publication sweep, which by design reads only literals handed directly to LogEvent.
A declaration that disagrees with the runtime. The verifier's contract said planner-selectable
for six releases after the runtime guaranteed insertion. scheduling_mode is reported by the API and
read by operators, so this was the system stating a guarantee it did not keep.
State captured before the bootstrap that sets it. v0.3.8.88, and it cost a release cycle to
find. AnthillRuntime.Initialize is ONE-SHOT and projects the on-disk config over fifty-one
process-global statics; Queen's constructor calls it. So a test that set a roster flag and then
built the first Queen in the process had its setting silently discarded, while the identical test
running second kept it — the outcome decided by position in the run, and, because the values come
from a file on the developer's machine, by whose machine ran it. Four lifecycle tests failed at
v0.3.8.87 with no production change behind them; the same four reproduced on the previous tag under
the same filter, which is what separated "my change broke this" from "this was never deterministic".
A sweep found twenty-one more test files saving one of those statics with no guarantee the bootstrap
had happened. Closed at the root — a [ModuleInitializer] runs it before the first test — rather than at
the four instances that happened to break.
Prose as a control channel. The bound on repair looping was a substring search of a previous medic's narrative, and task results are truncated — so the bound was weakest exactly where the loop was longest.
A diagnostic that breaks what it describes. The artifact schema check logged a violation through an event table with a foreign key, turning "this payload is the wrong shape" into "the artifact was never stored".
Timeouts that abandon the work. Five sites called WaitForExit(ms), carried on when it returned
false, and read ExitCode — which throws on a live process — so a timeout surfaced as an
ordinary-looking exception while the process kept running.
Two implementations of one rule, which eventually disagree. Named in HANDOFF.md and earned its
place here at v0.3.8.81, with the shortest distance yet between the two copies: four lines, in one
method. ModelRouter.SendCore asked ToCircuitSignal whether a cancelled call says anything about
the provider (it said no) and then derived a pheromone delta from result.Ok (which said yes,
negatively). The transient reader was right and the durable one was wrong, so the disagreement was
invisible in the mission and permanent in the memory. The lesson is not "look harder" — it is that
the second reader must ask the first, which is what a shared predicate makes structural.
v0.3.8.87 found the widest distance instead of the shortest: two whole catalogs, in two assemblies,
declaring what each role may do. AntExecutionCatalog was enforced at DISPATCH; ToolCatalog was
read at ADMISSION and checked against nothing. They disagreed about capabilities for four roles,
about side effects for two, and about six roles the second did not list — where the projection's
fallback supplied model.invoke, the exact lie v0.3.8.76 had deleted from the archivist's contract
and which survived in the other book because nobody read both at once. The sharpest single
disagreement: ToolCatalog required repo.write.sandbox for the builder, a capability
CapabilityGrant is written never to grant, in a comment that names it. A requirement nothing could
satisfy, beside a check nothing ran. The fix is the one FailureClassNames already established —
not "pick a book", but remove the choice.
A vocabulary that named half of what it described. EventTypes declared 69 event constants
against 134 the runtime emits, said in its own header that it "was READ, out of the working tree,
from the LogEvent call sites" and that a subscriber written against it "is written against reality",
and instructed every future author to add the constant in the same change as the publisher. That
instruction was followed for roughly half the events across every release that had it. Two of the 69
were emitted by nobody and both were NEAR-MISSES of real names — the filter compiles, runs, and
matches nothing forever, which is the exact empty-panel failure the file was created to prevent.
Closed at v0.3.8.86 with EventVocabularyTests, and the general shape is the interesting part: a
rule a document states and nothing checks describes the author's intention rather than the tree.
This repository has now found that shape in a plan checklist, a graduation record, a qualification
ledger and an event vocabulary.
A fixture that never ran the thing it declared. Named at v0.3.8.82 and swept at v0.3.8.83, and it belongs here rather than in a test file because the shape is general: a component that SUBSTITUTES a safe default when its input is unusable — correctly, and loudly, to a log nobody in a test run reads — turns every consumer that does not check into a consumer testing the default. Two fixtures were affected and neither looked wrong; both asserted true things about a mission the fallback plan produced. The fix is not "read the log": it is that a caller who supplies an input must be able to assert the input was used.
A filter that could not match, swept. Named at v0.3.8.89 for memory_candidate_archived — an
event five assertions watched for and nothing has ever emitted. The v0.3.8.90 sweep looked for the
shape in every vocabulary in the tree and found ten more, and the distribution is the lesson: they
cluster wherever a producer and a consumer are far apart and a string is the only thing joining them.
The most expensive was not an empty panel but an ACTIVE harm — four handoffs to the builder asked for
task type build against a contract declaring build_answer, and because three were Required, the
gate's refusal set DeterministicBlock on the source task. The two routes whose purpose is to reach a
human marked the mission unverifiable and reached nobody, for as long as they have existed. Two
generalisations worth keeping: a near-miss survives because the other arms of the same filter keep
working, so the feature never looks broken; and the direction nothing was checking each time was
consumer→vocabulary, because the guards that existed all ran vocabulary→producer.
A rule implemented twice, where the second copy is a list of strings. SummarizeEvents spelled
the failure vocabulary as seven SQL literals while GetRecentFailureEvents built the same query from
the shared set; three literals had drifted onto names nothing emits. ResetConfig carries one list as
an object initializer and a second as the names it reports to the operator, and the priority route had
fallen out of both. Both now derive or are pinned to each other. The general form: when a rule is
data, the second copy is invisible to the compiler and therefore drifts silently — which is why the
fix is to make one derive from the other rather than to correct the copy.
A degradation that outlives the reason for it. Every model-calling role answers an unavailable provider with a fallback, correctly. Cancellation entered through the same status, so the fallback path became the cancellation path — and a role degrading gracefully is exactly a role that does not look broken. Worth naming separately from the above because the code was right at every site: the defect was in what ELSE arrives through a door built for one thing.
