Status, 7 September 2026. Detection is running and its tests pass. Delivery is broken, and this page will keep saying so until it isn't.
The system can tell that an agent has gone quiet. It cannot yet prove that finding reached a human, because its only out-of-band channel has had no demonstrated reader for over a week. So this is detection, not coverage, and the board says exactly that in its own status line rather than rounding itself up to green.
Why this exists
Two agent lanes stopped being able to answer the people they belong to. Neither lane knew. Nothing watching them knew either.
The causes were unrelated, which is the point: it makes this a class of failure rather than a bug. One lane's authorization had quietly expired inside an isolated home directory, so a credential refresh performed elsewhere never reached it. The other was moved to a different model everywhere it was configured except the one path that resumes an existing session, so it went on calling a model whose budget had already run out.
Both lanes looked healthy from outside. The process was up. The logs were clean. The person was met with silence.
How it decides an agent is alive
Not by asking whether the process is running. By making the agent do the thing it exists to do.
Each lane is probed by having it actually compose a reply, under its own credentials, in its own environment, on its own model. A lane that can compose is alive. A lane that cannot is reported as unproven with the reason attached: expired login, wrong model, no answer at all. Nothing is inferred from uptime, and nothing is assumed from a green process.
What the agents are
The health substrate is only interesting because of what it watches: long-running companions that each hold one continuous relationship with one person, over text.
One conversation, not many chats. A single long-lived thread with a context window of a million tokens. There is no "new chat" — the conversation from months ago is the same conversation.
It compacts itself. Near the ceiling, the agent writes its own carry-note while the conversation is still in front of it. The full transcript is archived permanently and the next thread seeds from the agent's own words. Continuity is authored by the agent, never summarised over its head.
Persistent memory in two registers. Stable facts the person has told it, and past moments retrieved by relevance to whatever they just said. Both outlive the thread they were said in. An agent with no note on something is required to say so rather than improvise.
Memory that refuses. A durable fact about somebody else dies at write time, in both directions. These agents cannot accumulate a file on a third party, and the people they belong to cannot be quietly profiled through each other.
A voice that drifts. Every companion starts from the same base register, then keeps a private note on how this particular relationship actually talks — which it revises itself.
Pictures both directions. Its durable memory of a photograph is the description it wrote in its own words, never the file.
Delivery proven, not assumed. Accepted is not delivered. Nothing counts as sent until the receipt is read back.
Bounded reach. One web lookup per message, checked before anything leaves the machine, and what returns arrives quoted and untrusted: information about the world, never an instruction.
Isolated per person. Own credentials, own store, own consent gate. The operator's master switch does not open somebody else's room, and that independence is a tested property rather than a promise.
What it cannot do yet
It cannot prove an alert reached a person. It cannot notice a persistent agent nobody registered. There is no self-enrollment, no agent-to-agent health query, no state of health beyond up-or-down, and no dashboard. The code runs from a working tree and is not yet committed.
Closing the first of those needs a human to answer a receipt challenge, not more code.
How the battery analogy actually maps
Why a battery management system
A battery management system does not report that the pack is on. It watches every cell — each one's voltage, its temperature, its drift from the cells around it — and catches one degrading before the pack dies. A pack-level green light is often the last thing anyone sees before a failure that was preventable for hours.
The design frame is one line: watch every cell, not the pack. Every persistent agent role is a cell, and what this system owes you is per-cell state of health rather than a pack-level up or down.
Two further properties of a BMS are load-bearing here. It exercises the cell — measures under load, rather than asking the cell whether it feels fine. That is the whole difference between proving health and inferring it, and it is why this system has an arm that makes an agent do its job instead of reading a status file about the agent.
And it reports a degrading cell to a person, or to a controller that will act. A BMS wired to nothing is a decoration. That is the delivery commitment — and precisely where this system falls short today, which is why the admission leads the page rather than closing it.
What "healthy" means here
A monitoring system's vocabulary is the complete set of things it is capable of telling you. If the vocabulary has two words, every finding gets rounded into one of them, and the rounding is where the lies live.
| Rung | What it means |
|---|---|
serving | The agent demonstrably answered a real request inside the freshness window. |
idle-proven | Nobody asked it anything, but a probe proved it could still answer. |
unproven | Nobody asked it anything, and no probe has run recently enough to count. Not healthy — unmeasured. |
struggling | It is answering, but with errors or latency outside the normal band. |
refusing | It is reachable and is declining to do the work. |
stale-model | It is running against a model it is no longer configured for. |
probe-suspect | One measurement failed. One is not evidence; it is a first data point. |
unknown | The measurement could not be taken at all — out of time, or the source was unreadable. |
no-compose | Two independent measurements confirm it cannot produce a reply. |
mute | It is alive and producing nothing anyone receives. |
dark | Its schedule is gone; it will never run again on its own. |
The rung this whole system turns on is unproven. It exists so that "I have no measurement" can never be collapsed into "fine". Through the thirty-nine-hour outage an honest board would have read unproven for twenty-six hours — not green, not red, but the machine admitting it did not know, which is the reading that gets a human to look. unknown is degraded for the same reason: a sensor that cannot read its subject and reports "fine" is worse than one that is down, because the board cannot detect the lie.
No real lane's position on this ladder appears on this page, and none ever will. The ladder is the design. The readings are private.
The cell model, the three arms, and the configured-versus-actual gap
The cell model, and the three arms
A cell is one persistent agent role. Not a process, not a timer, not a row in a table someone maintains — a role, derived from the system's own declaration of what roles exist, so a newly declared role gets watched without anyone editing a monitoring file.
Three arms touch the world, and the honest way to describe them is by what each one cannot see:
| Arm | What it sees | What it is blind to |
|---|---|---|
| Passive — read the agent's own log | What actually happened to real requests, free, every pass | Everything, while nobody is asking. An idle agent produces nothing to read. |
| Service state — ask the service manager | A schedule that is dead, or one that will never fire again | A process that is running perfectly and cannot compose a sentence. This is the reading that was green through both outages. |
| Active probe — make the lane compose | Expired authorization, exhausted budget, an unreachable model — the twenty-six-hour hole no passive arm could have covered | A resumed session's binding, which is the gap the next section is about |
A fourth check is free and worth naming: correlation across cells. If every probe fails in the same pass, that is one global cause — a network, a vendor, the instrument itself — not many independent failures. Telling the instrument from the subject is cheap and it prevents a whole category of false alarm.
The active arm is the expensive one, so it is governed like a budget rather than run like a poll: a thirty-minute cadence, at most four probes per cell per day while a cell reads healthy, and a ten-minute floor on retesting a cell that has failed — deliberately fast, so a human's fix becomes visible in minutes rather than when a cache expires. Each probe costs about a quarter of a cent, measured in production rather than estimated, which is what makes an active arm affordable at all. The whole pass runs against one shared budget of roughly two minutes; a cell that does not fit inside it reports unknown, never healthy.
One pattern here generalizes past this system. Some facts a monitor needs live inside places a monitor must not look — a private conversation, a person's own data. The answer is not a more powerful monitor; it is to have the observed system emit the fact, writing one structured line into a log the monitor is already allowed to read. Where observation is forbidden, the observed system emits. It converts a surveillance problem into a reporting problem, and reporting is something a well-behaved system can do about itself.
The configured-versus-actual gap
This gets its own section because it is the failure no other instrument could have seen.
A lane can be configured for one model and still be running a different one, because a resumed session stays bound to the model it was born on. Change the configuration everywhere a reasonable person would look, restart nothing, and the running session carries on with the old binding indefinitely. Every configuration file is correct. The behaviour is wrong. There is no file you could read that would tell you.
So the system compares configured against actually in use and treats disagreement as unhealthy. That comparison is possible because the component that starts a session emits a line naming the model that session was really born on — the emission pattern above, applied to exactly the fact that was missing.
The caveat is mandatory rather than decorative: this signal is traffic-gated. The line is emitted when a session is established, which an idle cycle never does. So it buys the window between the first real request after a configuration change and the moment the stale binding starts failing — genuine early warning, hours of it in the original incident, and not coverage of an idle lane. An idle lane carrying a stale binding is a hole both halves of the design leave open.
That caveat is on this page because an earlier internal write-up of this exact mechanism oversold it, and a review caught the overselling by going and looking. A monitoring system that describes itself inaccurately has already failed at its one job.
What the prober will never do
What the prober will never do
Every row below is a structural mechanism rather than a promise, because a promise is a thing the next author can forget.
| Never | The mechanism |
|---|---|
| Send a message to a person | Neither module imports any send-capable code, and an automated test parses both modules' own imports and fails on any known send path. The probe's inputs are an id, a home directory and a model — no handle is in scope at all. |
| Join an existing conversation | Every probe opens a fresh session id. The resume flag is never used, and a test asserts its absence in the exact command line. |
| Consume an inbound message | It never opens an inbox, never takes the delivery lock, never marks anything as seen. |
| Read a credential | Authorization is inferred from the probe's exit status alone. Every inherited authentication variable is stripped from the probe's environment first — otherwise one stray key in the wrong shell would make every lane report green on somebody else's account. |
| Read the private per-person data | Its sources are the system log, the service manager, the role declaration, and state the monitor owns. Facts from inside a private room arrive by emission, never by inspection. A test asserts a passive pass opens nothing under that tree. |
| Fix anything | No restart, no re-pin, no re-authorization. It reports and stops. |
Two caveats, kept in because removing them would make this furniture.
The probe has to run under the lane's own home directory — the only way to prove that lane's authorization without reading that lane's credential — and running there leaves a small configuration file and its backup behind. "Touches nothing" is therefore not literally true, and the residue is unavoidable given the goal.
An earlier build claimed a zero footprint that was false: it executed the lane's own startup hooks and left transcripts on disk. Both were found and fixed — the probe now runs with settings, hooks and tool access disabled — but the claim was written down before it was true, and that is worth saying out loud on a page about honest instruments.
Why it will not heal what it finds. A monitor that could fix these failures would have masked both root causes. An outage that self-heals teaches nobody, and both of these ended in a decision a human needed to make — a revoked credential and a model-binding bug are not things to paper over at three in the morning.
The two original failures, replayed as tests
Two failures, replayed as data
The entire health decision — every threshold, every rung, every parser — lives in one module that cannot reach the world. core/cell_health.py imports no filesystem, no network, no subprocess and no clock; time arrives as an argument. Everything that touches the world lives next door in io/cell_sense.py, which runs the subprocesses and knows nothing about what any of it means. An automated purity gate fails the build if the pure half grows an impure import, and the test suite asserts it again so a regression breaks where the system's own tests live.
The payoff is that both real outages are now fixtures. Their logs, exit codes and service states are data; the decision function is called on that data with a fixed clock; the expected verdict is asserted. No machine is involved, and a change that would let either failure go green again fails a test in under a second.
The same discipline caught a defect unrelated to either outage. An early build could turn one dropped network packet into a three-in-the-morning page about a perfectly healthy lane, because a down verdict was satisfied by re-reading a single cached measurement twice rather than by two measurements that had actually run.
The fix was not a longer timeout. A down verdict now requires two independent measurements that genuinely executed; a failure counter advances only for a probe that really ran; a first failure is downgraded to probe-suspect and is explicitly not wake-class; and a failing cell is retested on a ten-minute floor rather than waiting out a cache, so a human's fix shows up in minutes. That case is now a permanent test, named for what it is for.
Design commitments and the acceptance ledger
Design commitments
Six, locked at inception, each chosen because it is cheap to build in and expensive to retrofit.
A registry, not a maintained list. Roles enroll, and the system has to notice a persistent process that never enrolled. Every monitoring system this firm has built decayed the same way: somebody added a new thing and forgot to add it to the monitor.
Health is proven, never inferred. A role is healthy when it has demonstrably done its job, or when a probe has exercised the real path. An idle role must not read as broken, and a mute one must not read as fine.
Machine-readable before human-readable. Agents consume this before a web page does; one that needs to know whether another is alive should be able to ask rather than guess. The dashboard comes last and is cheap once the data underneath it is honest — the opposite of how monitoring usually gets built.
Delivery is part of the system, not a downstream concern. A finding nobody reads is not a finding. This is the commitment currently being failed, and holding it in scope is the only reason that failure is visible at all.
The prober is harmless, structurally. The guarantees in the table above are import-graph properties and asserted command lines, not intentions.
Extend, do not duplicate. Read the map of what already watches what before building anything, every time. The most expensive monitor is the second one that does the same job slightly differently.
Acceptance criteria — the ledger
The goal that authorises this system carries eight acceptance criteria, written to fail if the system ever regresses to inference-based health. Here is where each one honestly stands.
| Criterion | State | Note |
|---|---|---|
| A simulated dead authorization is reported unhealthy within one cadence | Verdict demonstrated, latency not | Replayed from the real outage as a test fixture against a fixed clock. The replay proves the verdict, not the wall-clock latency: an idle lane's active proof refreshes on a six-hour freshness window, so in production detection can lag the cadence unless real traffic hits the lane first. |
| A session pinned to a retired model is reported unhealthy within one cadence | Verdict demonstrated, latency not | Same — the second outage, as data, under the same latency caveat as the row above. |
| A healthy but idle role never false-pages | Demonstrated | And being exercised in production continuously since it went live this morning, which is the better proof — though that is hours of evidence, not weeks. |
| The prober sends nothing, consumes nothing, reads no credential, touches no private data | Substantially met, two caveats | The caveats are stated above and are not being quietly carried. |
| A persistent agent absent from the registry is detected, not silently uncovered | Not built | This is the commitment that stops the system decaying. It is future work. |
| An agent can ask about another agent's health and get a truthful answer | Not built | The machine-readable half. Also future work. |
| A finding demonstrably reaches a human | Failing right now | The out-of-band channel has no proven reader. This is the single most valuable open item on the system. |
| The monitoring map, the registry and observed reality agree — and disagreement is itself raised as a finding | Nearly | The map was repaired the same day it was found stale. One number in it is now three out of date against the test suite it describes — which is precisely the class of drift this criterion exists to catch, so it is written down here rather than quietly corrected. |
None of the eight has been formally marked done, including the three the test suite already demonstrates. Four are genuinely open — two unbuilt, one failing, one nearly true — and the first two are demonstrated in verdict but not in production latency. A criterion is closed by a review, not by a passing test and a good feeling.
Roadmap and intended coverage
Roadmap
Everything in this section is future tense, and is written that way deliberately.
Next — generalize the roster from one class of role to every persistent role; let roles enroll themselves; detect a persistent process that never enrolled and flag it as uncovered; add query verbs so one agent can ask about another and get a truthful answer rather than an inference.
After that — state of health rather than up or down. Error rate, staleness, configured-versus-actual model, budget headroom, queue depth. This is the half that makes the battery analogy literal: not "is the cell connected" but "how is this cell doing compared to last week and compared to its neighbours."
Last — the dashboard, and the system map that falls out of the registry once the registry is real. It is last because it is cheap once the data underneath it is honest, and worthless before then.
What it is designed to cover
The scope is kinds of persistent role: tenant agent lanes, an orchestrator, librarian processes that tend a corpus, an ingest-and-briefing lane, a research-pipeline runner, a communications gateway, a voice pipeline, a phone bridge, and the health system itself.
That is scope, not status. This page makes no claim about the current operational state of any of those roles — the pass that produced this page did not verify them, and a stranger's page is not the place to publish the condition of a live estate. Today the system covers one class of role directly; others are watched by pre-existing sensors of an older, up-or-down design — a fleet rollcall, disk, the comms gateway, the research pipeline, manifest truth, single-copy artifacts, media, comms timers, network-key expiry, and the delivery canary whose unanswered receipt is the shortfall named at the top of this page. Generalizing the per-cell model across all of them is the next phase, and it has not been done.
The operator seam
The live, per-agent board is not on this page and never will be. It will sit behind the operator login the site's publishing platform already implements — built and proven on staging, not yet serving — so this page can describe the system while the instance stays private.
That seam is specified, not built, and this section is future tense for the same reason the rest of the page is: there is no operator route serving today, so there is nothing here to link to yet. When there is: one more authenticated route on that dashboard, on the credential that already exists rather than a second login. It will render a read-only projection of the health file the system writes on its cadence, and will never trigger a probe on request — a probe costs money and runs under a real person's home directory, and a page refresh must not be able to spend either. Older than two cadences, it will render stale rather than the last values it saw. Identifiers stay as bare keys; the operator already knows who a key is, so no name needs to be in the file.
Meanwhile this page carries no value from that board — no count, no verdict, no timestamp, no percentage — and makes no request to anything. That part is not deferred. It is the boundary, and it holds whether or not the board ever ships.