MICROMOUND — DEVICES · v0.4.2.14
Safety and physical authority
Source: runtimes/micromound/docs/SAFETY.md, a path from the root of the product's repository — versioned with the code, rendered here at every release.
Canonical safety text. Where this document and any other file disagree, this document wins. Nothing in this repository may weaken a rule here without this file changing first, loudly.
Layer 0 — Independent safety systems (not ours)
Emergency stops, hardware watchdogs, interlocks, limit switches, thermal fuses, mechanical stops, RCDs and breakers. These belong to the electrical and mechanical design of each device, sit below all software in this repository, and are not addressable by anything here:
- No protocol envelope, charter field, manifest entry, routine, driver, or reasoning provider may configure, suppress, reset, or depend on defeating a Layer 0 device.
- Software treats Layer 0 trips as observed facts to report — with evidence — never as states to manage or recover from automatically. A Guard Ant reports an interlock trip; it does not clear one.
- A device whose Layer 0 protection is known-faulty is unfit for any charter above
observe.
Layer 1 — Deterministic enforcement on-device
The capability kernel is the single physical authority boundary. Every actuation, on every
hardware tier, passes through it — not by convention but by construction: drivers are reachable
only through ICapabilityExecutor, executors are held only by the kernel, and nothing hands one
out.
Limits intersect across three tiers, innermost first:
hardware/firmware ∩ device manifest ∩ charter = effective
Ceilings take the minimum, floors take the maximum. An outer tier can only narrow. A charter that asks for a longer run than the relay tolerates, or a shorter cooldown than the pump requires, does not get one — the request is intersected away at execution and the attempt is reported at validation.
Also at this layer:
- Software watchdog. Loss of the runtime's own heartbeat drops actuation and enters the
declared
safe_state. Safe states are de-energized or passive by construction. Enforced by the Guard Ant (Micromound.Runtime): a stale heartbeat or an observed safety trip makes it demand a safe state, and the coordinator engages the stop rather than continuing. A stale heartbeat is self-healing — a watchdog that latched on a scheduling hiccup is one nobody leaves enabled — but an observed trip is sticky and nothing in software clears it, because software that could clear a trip is software that could be asked to. - The safe-state walk is per driver, and every path takes the same one. Driving the hardware safe
means walking every driver, and one driver that throws must not decide the fate of the rest: each is
isolated, a failure becomes a sticky safety trip rather than a silent gap, and the whole walk is
serialised behind the host's safe-state gate so the service loop and the watchdog thread cannot
interleave on the same hardware. Every transition into stopped or quiesced — from a sync, a mission,
an expired lease, a cold start with a mission in flight, a shutdown, or the watchdog — goes through
that one method. Before
v0.9.29two of those paths walked the drivers themselves in a bare loop, so the first driver to throw left every later one energized during a stop, with no trip recorded; the fix was to route them through the isolated walk that already existed beside them. - A timed actuation is held, and its release is owed on every path. A digital actuator drives its
line active and holds it for the effective
on_s, so a real valve is open for its duration rather than pulsed; the hold is bounded (the requestedon_sis clamped to the intersected limit tiers and capped again at the effectivemax_on_s), and it is released on the service loop's cadence and by thesafe_stateon any stop, quiesce, shutdown, or trip. A line that will not de-energize is not swallowed: it keeps its hold pending and escalates to a sticky, persisted stop, because a line that cannot be proven safe is treated as unsafe. - A deadline is measured in elapsed time, not in clock readings. A hold bounds how long a line may stay hot, so nothing that merely takes time may buy it more: the service tick releases due holds before the blocking sync as well as after it, and the span the sync actually cost is measured monotonically and added to the tick's own clock, so a slow round trip or a timeout is time the hold sees. The hold itself carries two deadlines — a wall-clock instant and a monotonic duration — and releases on whichever comes first, because a wall clock can be stepped: an NTP correction jumping backwards must not postpone a release that is physically already due. Earliest wins is the fail-safe direction here; de-energizing early is safe, holding late is the failure the bound exists to prevent. The lease is re-checked after the sync for the same reason.
- An independent watchdog releases a held line behind a hung loop. Because a timed actuation is
held between ticks, a service loop that hangs would leave a line hot — the stale-heartbeat rule
refuses new actuations but cannot release a line already held. A hardware-independent watchdog on its
own thread (
LoopWatchdog/WatchdogThread, daemon--watchdog-s) notices the loop has stopped kicking and drives the mound to a de-energized, sticky, persisted stop without the loop's help. The concurrency is made correct rather than hoped for: the Guard is thread-safe, the host's safe-state path is serialised behind one gate with a consistent lock order and a bounded wait so the watchdog cannot itself wedge, and the loop answers the watchdog at the top of each tick — through the Guard's lock, a memory barrier — so a loop resuming from a hang stops itself before it could actuate on a stale, not-yet-stopped view of authority. Set the timeout generously so an ordinary GC or scheduling pause never trips it. - No single driver may hold the safe-state walk. Two ways a driver fails on the way to safe, and
they need different answers. One that THROWS is caught per driver, reported as a trip, and the walk
continues (
v0.9.29). One that BLOCKS cannot be caught at all: it stops the walk at itself, every driver after it in the manifest stays energized, and the caller holds the safe-state gate while it waits — so the independent watchdog cannot get in either, and one stuck I2C transaction on a sensor keeps a pump running. Each driver now gets a bounded call (v0.9.39, roadmap P0.8, default 5 s; a GPIO write is microseconds and an I2C transfer milliseconds, so five seconds is already pathological). Past the bound the mound stops WAITING: it trips, abandons that driver — permanently, because it has proved it does not answer and each retry would strand another thread — and makes the rest safe. Each driver's bounded calls run on a thread of that driver's own (v0.9.42), never a shared pool:v0.9.39used the thread pool, so a blocked driver occupied a pool thread and the next driver's call waited on the pool to inject another — about a second, longer than the bound. On a small machine the second driver timed out as well and its line stayed live. The guarantee inverted itself exactly where it matters most, because a Pi is a small machine. A driver with its own thread can only ever wedge itself. Stated exactly, because the difference matters: nothing interrupts the stuck driver or makes ITS line safe; a blocked call cannot be cancelled. The guarantee is that one blocked driver cannot keep unrelated outputs live, and that the gate is released in bounded time so the watchdog can take it. For the stuck line itself, process supervision (systemdRestart=, whose restart de-energizes at configure time) remains the backstop — but it is now the backstop for one line rather than for the whole mound. - A device never fakes its hardware by accident. A manifest that names physical ports (a pin, a
channel, a bus address) is refused by a daemon running on in-memory ports unless the operator says
--simulatein so many words (v0.9.17): in-memory readings and actuations look real and are neither, and a warning scrolling past a log is not consent.--check-hardwareclaims every port the manifest names and reads each sensor once — at the safe level, never actuating — so the wiring is checked before any authority exists. - A line comes up at its safe level, and stays there only while the daemon holds it. Both GPIO
backings now request a line already at
!active_high— the character device carries the initial value in the line request, sysfs writeshigh/lowas the direction — so an active-low relay is never energized for the instant between "becomes an output" and "is written safe" (v0.9.16). The flip side is stated plainly: when the daemon exits, crashes, or releases a line, the kernel returns it to its default state (usually an input, floating or with the board's pull), and the level is no longer held by anything. Whether that idle state is safe is a property of the board — a relay input with a pull-up idles off; one without may not — and it is exactly the case Layer 0 exists for. Process supervision restarts the daemon, which requests the line at the safe level again. - Clamp, don't lie. Where a limit narrows a request, the work proceeds and the outcome is
clamped, carrying both the requested and the effective parameters plus the limit responsible. A silent clamp is a false statement about what the mound did. - Model output is a proposal. Any reasoning provider produces proposals to the deterministic
layer and holds no actuation path. This is enforced structurally:
Micromound.Reasoningdoes not referenceMicromound.Capabilities, so a provider cannot call the kernel, hold an executor, or touch a driver.
Layer 2 — Authority (charters and leases)
- No charter →
observeonly. Expired lease →safe_state. Ambiguity → downward. - The expiry is checked on every service tick, whether or not anything is happening (
v0.9.27). A lease is a promise about TIME, so it runs out on an idle mound exactly as it does on a busy one, and the scenario the lease exists for is precisely the one where nobody is left to ask the mound anything. Before this it was checked only when a mission arrived or a process restarted, which left an idle moundcharteredwith its outputs live for as long as the silence lasted.MoundService.TickcallsMoundHost.QuiesceIfLeaseExpiredbefore the sync beat; crossing the expiry de-energizes every driver and persists the quiesce, so a restart comes back quiesced rather than briefly re-authorized. - Disconnection never widens authority; nothing on-device can extend a lease. Renewal happens only when the controller acknowledges a sync beat.
- Reconnection resumes nothing. A quiesced mound reports its state and waits for fresh authority.
- Registration-time refusals, because a misconfigured device should fail at startup rather than at
first use: a
sense.capability may not be classed aboveobserve; nothing may be registered ashazardous; a routine may not be classed below a capability it drives. - The audit path is bounded, and what it loses is counted. A queue that grows without limit ends in a full disk, and a mound that cannot write cannot record what it did. The uplink queue is therefore bounded by items and bytes, spills oldest-first when it must, and reports the count on the next beat — the chain makes a gap detectable, and the count makes it explicable.
- A mound that cannot record what it did must not do it. The bound above is now enforced BEFORE
the effect rather than after it: the kernel's fourteenth check refuses new physical work once the
audit path has no room for the record that work would produce (
no_record_capacity,v0.9.37). Enforcing it afterwards meant spilling — trading history the mound already owed for work it had not done yet, which is the wrong trade in a system whose whole claim is that every actuation is accounted for. Refusing loses only the work, and the controller can ask again. Both profiles do this now; the reduced-profile device used to refuse to record while still acting, which was the same defect wearing the opposite failure. The queue holds a slice of its bound back so the refusal can always itself be recorded — a mound that could refuse but not say so would have swapped one silent failure for another. Observation is exempt, for the same reason a stop does not blind the mound, which means sensing can still fill a queue past that reserve; the mound then loses readings rather than the account of what it physically did. - A clock that is STEPPED cannot hand back what the hardware owes. A duty cycle and a rate limit
are answers to "has enough time passed?", and both were computed by subtracting two readings of a
wall clock — which moves for reasons other than time passing. The case is not exotic: a Pi or an
ESP32 with no battery-backed RTC boots believing it is 1970 and steps forward by decades on its
first sync, at which instant every cooldown reads as elapsed and every rate window as empty. Each
recorded instant now carries a monotonic stamp alongside it, and an entry's age is the SMALLER of
what the two clocks claim (
v0.9.38, roadmap P0.5) — the mirror of the rule for releasing a hold, where the question is "is it time to de-energize?" and the LARGER elapsed wins. A step backward was already conservative and stays so. Where the guard does not reach: across a restart the monotonic counter reset with the process, so a restored budget is aged on the wall clock alone; the real gap is unknowable from inside the mound, and only the controller could close it. A board whose HAL offers no monotonic source is likewise wall-clock only, and says so by leaving the hook NULL rather than by pretending. - What the hardware owes survives a restart. A capability's minimum off-time and its rate budget are limits on the DEVICE, not on a session: they persist and are restored before anything may ask the hardware for more, so a reboot cannot hand back a cooldown that was already spent. And what the mound has already been told to do persists too — a controller redelivering a completed mission, whether as the same envelope or as a fresh one around the same mission id, is answered with a refusal rather than a second actuation. That ledger is bounded by a validity horizon rather than a count, so an entry can never expire into being executable again: an instruction older than the horizon is refused, because the mound can no longer prove it has not already run it.
hazardous-class actions — physical risk to people, property, or surroundings; fabrication tools, motion near people, building systems — require explicit per-action authorization from the controller, never a standing grant, expiring on use or timeout. Until that pipeline ships with tests, hazardous actions are refused unconditionally, andhazardousis not a legal charter ceiling.
Layer 3 — Controller oversight
- Every actuation is audited with evidence;
unverifiedactions gate missions as failures. An actuator's own report is never that evidence — a command is not evidence — so what makes an actuation verifiable at all is a second, independent observation: thedigital_sensorprimitive (a limit switch, an interlock contact, a float) read by a mission'sverifystep. A mission may promise that observation up front, which holds the verdict open until it arrives; a promise not kept demotes the record tounverifiedbefore it is published (PROTOCOL.md §9). - Stops: physical (Layer 0), per-mound, and global. Stop processing precedes all other downlink and needs no valid charter. Clearing a stop restores nothing — the mound returns to observe-only and waits for a fresh charter.
- Who may end a stop (
v0.9.46; roadmap E-17, decision F82, PROTOCOL.md §7).- Scope. A stop says whose it is. A controller's stop of one mound has scope
mound. Its colony-wide stop iscolony. A stop the mound took itself (a guard trip, the watchdog) isdevice. A stop that does not say iscolony. - From the controller. It can end only a
moundstop, with a signedstop_clearthat names the stop by number. A clear that names another stop, arrives past the replay horizon, or is replayed ends nothing. - At the device. Everything else is ended by a person with the device:
micromound --clear-stop, with the service stopped, on a Pi-class mound, and reprovisioning on the C mound. - Why the line is there. Before
v0.9.46nothing on the wire cleared a stop. That was the containment for a stolen controller key: it "can stop the fleet and cannot un-stop it". The owner kept that for the emergency stop and for the mound's own stops. A stolen key can now end a stop it could itself have given one mound, and nothing more. The mound it ends comes back observe-only.
- Scope. A stop says whose it is. A controller's stop of one mound has scope
- A stop is acted on when it arrives, not when the queue empties. An authenticated stop takes
effect on the exchange that delivered it and ends that drain, and a sync beat is bounded in how
many batches it will push. A deep backlog is exactly the situation an operator reaches for the stop
in, so the amount of queued work must never be an input to how fast the mound stops. This is true
of both profiles since roadmap E-20. Before it, the C mound acted on a stop on the exchange that
delivered it and then went on offering the rest of its queue in the same beat.
- Under a controller's standing stop that was worse than slow. Once the C mound's queue was full, its beat could not be queued and the drain had no end: each answer freed one slot and carried a fresh stop, whose acknowledgement took the slot. The outputs stayed safe, but the mound sensed nothing. And the board application writes the stop to storage only after a beat returns, so a board whose queue was already full when the stop first arrived held it in memory only. The ESP32 task watchdog then rebooted it unstopped (with no charter, so observe-only until its next beat was answered with the stop).
- Since roadmap E-20 no C beat goes past what was queued when it began, the beat or, when a full queue cannot take it, the newest envelope already queued (PROTOCOL.md §7). Every beat ends, and the stop is written through on the tick that received it.
- Since roadmap E-20 a Pi-class mound also drives its outputs safe and persists the stop at that moment, when an order begins a stop, before anything else in the beat is handled. Until then both waited for the rest of the beat. The walk and the persist after the beat still run. (The C mound has always entered its safe state there; it persists the stop after the beat.)
- A stop's reason is kept, and decides nothing (roadmap E-20). A Pi-class mound keeps the reason
its stop was taken for with the stop, across a restart, and
micromound --clear-stopshows it to the person ending the stop. For a controller's stop that is the order'sreason; for the mound's own it is the guard's words. No path reads it to decide anything: the stop, its scope and its number are the same with a reason and without one. The C mound reads no stop body and keeps none. - A stop survives the power, on every profile. "A restart never clears a stop" has been true of
the Pi-class host since it had a state store, and was NOT true of the C mound until
v0.9.35: the stop lived only in RAM there, so power-cycling a stopped board brought it back willing to actuate — and a fault that stops a mound is precisely the kind of event that also power-cycles it. It is now one byte in the same protected storage the identity uses — read back by PRESENCE rather than content, since nothing can accidentally create a key but a flipped bit can change one, and only "still stopped" is a safe direction to be wrong in — written on the tick the authority stops rather than on the beat that reports it, and read back before the first tick can run: the authority is stopped and the hardware driven safe during bring-up, so a board that comes up hot goes cold without waiting for anything. Nothing on the device clears it, and a fresh charter does not lift it either. Reprovisioning does, which is a person with the board in their hands. The identity-free port server is out of scope: it holds no authority at all, and its equivalent is the link watchdog that drops every line when the Pi goes quiet.- Why nothing clears it (corrected by roadmap E-20). Until then this said the storage abstraction
has no delete. It has no delete hook, but it can still erase a key:
mm_halstores a zero-length value as empty, the library reads empty the same as absent, and the ESP32 binding erases the key outright (nvs_erase_key). The library burns the enrollment token that way. - The stop holds because of what the library writes, not because of what the HAL lacks. The stop
key is written in one place,
persist_stopinmm_app.c. It always writes one byte there, and no code path writes that key empty. A change that did so would clear the stop, and it would have to change this rule first. - One path erases the whole store, the stop with it, and it is outside the library. A development
image built with
MM_ALLOW_NVS_ERASEerases the partition when NVS will not initialise. The rule below is why the autonomous images do not.
- Why nothing clears it (corrected by roadmap E-20). Until then this said the storage abstraction
has no delete. It has no delete hook, but it can still erase a key:
- An identity that exists is never silently replaced. The device's Ed25519 seed is what its whole
signed history hangs off. Until
v0.9.41any read of it that did not produce exactly 32 bytes fell through to minting a new one and overwriting the old, so a truncated read or a storage hiccup turned the mound into a different device — one that still signs, beats and enrolls, while the controller rejects everything it sends and the real mound's records are orphaned. The seed now carries a checksum of its own (not a reliance on the backing store's, because the HAL is an abstraction), and a bad checksum, a wrong length or a storage fault each halt the board with its outputs safe rather than guess. Minting is for a device with no identity at all, and the storage contract distinguishes "not there" from "there and unreadable" so that judgement can be made. - A storage fault does not become a new identity. ESP-IDF's stock boot recipe erases the whole
NVS partition when it will not initialise. On a mound that partition holds the Ed25519 seed the
identity is, the controller's key, and now the stop — so the stock recipe answers "the flash is
full" by minting a different mound and clearing a halt meant to survive anything, and it is the
reboot after a fault that is most likely to hit a full page. The autonomous images refuse and halt
with their outputs safe instead (
v0.9.35); erasing is a deliberate act by a person. - A stop ceases actuation; it does not blind the mound. Observation continues, as PROTOCOL.md
§7 has always specified, and the same section requires the stop acknowledgement to carry a
post-stop sensor snapshot — which a mound that refused to sense could never produce. Refusing
every capability would also darken the instruments at the exact moment an operator most needs to
see what the hardware is doing. The kernel decides this from the capability id's namespace,
before the registry is consulted, so stop still works when the registry, the charter and the
drivers are all broken.
- A Pi-class mound takes that snapshot since roadmap E-20, when an order begins a stop. It reads only after the safe-state walk and the persist, so no reading delays either. The walk gives each driver a bounded time: a driver that fails or does not answer in time is a trip, and its output may still be live while the sensors are read.
- It reads through the kernel, as observation: every id is in the
sense.namespace, which the registry pins to classobserve. Nothing in it can actuate. - It is evidence about the stop, never a condition of it. A snapshot that fails is named in the acknowledgement, and the acknowledgement still goes up. A read that never returns is the exception: it holds the acknowledgement, as a blocked observation holds any beat, until the host's loop watchdog (when armed) trips the mound. The outputs are driven safe and the stop is persisted before the first read. A snapshot that would take the room the uplink queue holds back for acknowledgements is not taken.
- An approval pipeline fronts all
controlledactions.
Prohibited by construction
- Unsigned protocol traffic; unsigned acceptance of charters or configuration.
- Mound-to-mound delegation or authority transfer.
- Any code path that retries a
hazardousaction without fresh authorization. - Any tool, endpoint, or envelope that reads back or exports a device private key.
- Any argument, field, or flag through which a caller can request an exception on grounds of urgency. Ambiguity resolves downward, and the way to guarantee that is to give ambiguity nowhere to enter.
- Silent failure. Every refusal, clamp, trip, and validation failure is reported and audited, with a specific reason. A refusal without a reason is itself a contract violation.
