Orbyt Was Not Planned. It Was Corrected.
Key findings
- An append-only record makes a reversal legible, because the superseded entry stays on the page with its own date next to the correction.
- Of 80 documented failures in this organization, 33 were classified as verification gaps: a test, guard, gate or metric that could not fail, could not see, or measured a proxy.
- The corpus classifies by what broke rather than by who broke it, so its class counts cannot be read as a comparison between agents and the tools that watch them.
- 42 of the 66 failures name a countermeasure by file path and 24 do not, and the unmechanized remainder is printed rather than netted out.
- The run ledger's halted count reads zero while a scoped halt sat on one seat for most of the window, because no workflow has ever been written to record one.
Key measurements
| Measure | Value | Source and date |
|---|---|---|
| Documented failures on record | 66, of which 55 are dated, first documented 2026-06-09, latest 2026-09-02 | September 3, 2026 · Orbyt Failure Corpus, public/failure-corpus.json, read directly by the author |
| Failures naming a countermeasure by file path | 42 of 66 | September 3, 2026 · Orbyt Failure Corpus, public/failure-corpus.json, read directly by the author |
| Failure classes | 33 verification gap, 18 operator process, 15 external behavior | September 3, 2026 · Orbyt Failure Corpus taxonomy, public/failure-corpus.json, read directly by the author |
| Numbered decisions on record | 45, D-0001 on 2026-07-13 through D-0045 on 2026-09-02 | September 3, 2026 · Orbyt Decision Ledger, public/decision-ledger.json, cross-checked by direct grep of docs/org/decision-log.md by the author |
| Agent seat runs recorded | 65 runs, 55 completions, 10 failures, between 2026-07-21 and 2026-09-01 | September 3, 2026 · Orbyt Autonomy Ledger, public/autonomy-ledger.json, read directly by the author |
| Finished drafts killed by expert panels inside those runs | 47, counted separately from run outcomes and never netted against them | September 3, 2026 · Orbyt Autonomy Ledger, public/autonomy-ledger.json, read directly by the author |
| Kill switch activations | one scoped halt on a single seat from 2026-07-20 to 2026-08-14, and one global halt set and lifted on 2026-08-20 | September 3, 2026 · Direct git history of the organization's halt files in the Orbyt Labs repository, by the author |
I did not plan Orbyt into the shape it has now. I corrected it there, and every correction carries a date.
No roadmap I wrote predicted this. What exists instead is a numbered failure corpus, a numbered decision ledger, and a run record for the agent seats that operate the place. All three are published, append-only, and licensed for anyone to cite against me.
A roadmap tells you what a company intended. A dated correction log tells you what it learned, and only one of the two can be checked by a stranger.
What is actually in the record?
Three files, and a preprint that reads them.
The failure corpus holds 66 documented failures, first documented 2026-06-09, latest 2026-09-02. 55 carry a date. The eleven that do not predate the dating convention and are published undated rather than backdated. 42 name a countermeasure by file path, and the build fails if that path stops existing.
The decision ledger holds 45 numbered decisions, D-0001 on 2026-07-13 through D-0045 on 2026-09-02. Nothing is ever rewritten. A decision that turns out wrong is superseded by a later entry with its own number and its own date, so the wrong version stays on the page beside the correction.
The autonomy ledger holds 65 seat runs between 2026-07-21 and 2026-09-01. 55 completions and 10 failures. Inside those runs, expert panels killed 47 finished drafts before anything shipped. Those are two different units and the ledger keeps them apart, because work that finished and was then refused is a different fact from work that never finished.
The rows are appended by the workflows themselves. Nobody types them in later.
All three ship under a Creative Commons attribution licence, with a citation line. The governance preprint is built on the same corpus, which is the part that took discipline rather than nerve. Counting is easy once. Counting for a quarter, under a taxonomy, in public, is a system.
Why were half of them my own instruments?
Because that is what the classification was for.
The corpus sorted its 66 into three classes at the time of writing. 33 are verification gaps: a test, guard, gate or metric that could not fail, could not see, or measured a proxy for the thing it claimed to measure. 18 are operator process. That means my own discipline. 15 are external behavior, meaning a platform did something other than what it documented.
Read that carefully. The obvious reading is wrong. The corpus classifies by what broke, never by who broke it. So 33 of 66 is not evidence that agents fail less often than the tools watching them, and it cannot be made into that. What it says is narrower and more useful. When this organization writes down what went wrong, the thing that went wrong is most often the thing that was supposed to notice.
A gate that read its verdict out of terminal text, and matched nothing once the text arrived colored. A row count taken from a query planner's estimate instead of a query. A sync rule that tested two files for equality, which a translation can never satisfy, so it deleted every translation on every run.
None of those failed loudly. A check that passes is not a check that works.
The failure that costs you is not the one that fires. It is the one that reports clean.
What does a correction have to produce?
A guard, or a printed admission that there is none.
A lesson is a hope. A guard is a fact. Writing be more careful at the bottom of an incident does not survive the month. So 42 of the 66 name a file. 24 do not, and that remainder is printed rather than netted out.
A countermeasure prevents a recurrence in the shape it was written for. That is less than it sounds. One of the 66 is exactly that failure. A fix was mechanized in one implementation while three other surfaces hand-rolled the same sequence, and the incident replayed on one of them thirteen days later.
There is a second rule underneath it. A guard is not owned until you have watched it fail. Several of mine were decoration until somebody injected the bug on purpose and watched the thing go red.
Where is the record blind?
Three places, all of them printed rather than smoothed.
The long-form lessons file is not published, and it is not append-only. Entries carry corrections written into the entry itself, and commits have deleted lines from it. I had been calling the whole record append-only. One quarter of it does not qualify.
The autonomy ledger's halted count reads zero across the window. Not because nothing was halted. A scoped halt sat on one seat from 2026-07-20 until 2026-08-14. The global kill switch was pulled for real on 2026-08-20 during a domain migration and lifted the same day. Both are in the git history. The ledger has neither.
I was wrong about why. The wrong answer was the comfortable one. I assumed a run that never starts cannot write a row about itself. It can. Several workflows read the halt file and exit cleanly when they find it, and any one of them could write a neutral row on the way out. The blindness is unwritten, not impossible.
Then the smaller hole. On 2026-08-29 the step whose only job is to record a failed run died while recording one. That run is absent rather than counted. 65 is a floor.
Is a public mistake log just marketing?
That is the strongest objection to everything above, and I am not going to reclassify it as off topic.
The operator is grading its own work. Both published files say so in their own caveat fields, in about those words. The sample is small. It is one organization, one quarter, one rater, and there is no outcome measure anywhere in it: nothing here shows the record made the company better. The preprint states those limits in its own body rather than in a footnote.
What makes the record more than a brochure is mechanical rather than moral. The run rows are written by the workflows. The countermeasure column points at a path, and the build refuses when the path goes missing. A claim that a mistake was fixed has to name the thing that fixes it, and keep naming it.
What does not: I still choose what gets written down. Nothing forces a failure into the corpus. A mistake I never noticed is a mistake the corpus does not have, and the corpus cannot tell you which ones those are. I would rather name that hole than let anyone read 66 as complete.
What to do Next
Take your own last quarter and write the list. Not the wins. The things that broke, one line each, with the date and what each one produced.
Sort them into the three classes above, then look at the pile marked verification. If that pile is small, check whether your instruments are good or whether they are the part nobody has been inspecting.
Then ask one question of every entry: what refuses now? If the answer is that somebody will remember, the entry is still open.
A company that cannot say when it was wrong is not telling you it was right. It is telling you nothing was ever checked.
Related reading:
- 84 Ways to Tell Me I'm Wrong. the guard estate this record feeds, and what it costs to keep green
- Self-Healing Is a Euphemism. why a repair that never reports is a failure you stopped seeing
- Codex Accuses. Claude Convicts. the two model review loop that produces a lot of these entries
- My C-Suite of Agents Named Themselves. who the seats in the autonomy ledger actually are
- 391 Yeses and Not One No. what happened when I counted my own approvals instead of trusting them
Methodology
Every count here was read from the published artifact on 2026-09-03 rather than from prose. Failure counts, the dated count, class counts and the countermeasure count come from public/failure-corpus.json. The decision count was read from public/decision-ledger.json and independently cross-checked by counting decision headings in the source log. Run outcomes and the panel discard count come from public/autonomy-ledger.json, where discards are a count of finished drafts rather than of runs. The halt history is read from the git history of the organization's halt files, which shows one scoped halt created 2026-07-20 and removed 2026-08-14, and a global halt set and lifted on 2026-08-20.
Limitations
This is the operator grading its own work, which both published artifacts state in their own caveat fields. The run sample is small and covers a single quarter, with one rater. The corpus classifies by what broke rather than by who broke it, so no count in it compares agents against instruments. There is no outcome measure anywhere in the record, so nothing here shows that keeping it made the organization better. Completeness cannot be established: nothing forces a failure into the corpus, so the entries are the ones that were noticed and written down. The run total is a floor rather than a total, because on 2026-08-29 the step that records a failed run died while recording one and that run is absent. The halted count reads zero because no workflow has ever been written to record a halt, not because none occurred. The append-only description does not cover the whole record either: the long form lessons file is unpublished, carries corrections written into its own entries, and has had lines deleted from it.
Sources
- The Orbyt Failure Corpus Retrieved September 3, 2026.
- The Orbyt Decision Ledger Retrieved September 3, 2026.
- The Orbyt Autonomy Ledger Retrieved September 3, 2026.
Common questions
What does an append-only record actually buy you?
It makes a reversal legible. A decision that turned out wrong stays on the page with its date, and the correction arrives as a new entry carrying its own number. Nobody has to trust that the earlier call was reasonable, because the earlier call is still there to read. Editing history removes exactly that.
Does a mistake log mean the agents are unreliable?
It does not say either way. The corpus classifies failures by what broke rather than by who broke it, so no count in it compares agents against anything. What the classes do show is that half of the recorded failures were instruments, 33 of 66: a check that could not fail, could not see, or measured a proxy for the thing it graded.
How do you stop a mistake log from becoming a wall of good intentions?
Attach a mechanism or admit there is none. Here 42 of the 66 entries name a countermeasure by file path, and the build refuses when that path disappears. The other 24 are printed as unmechanized rather than described as handled. A written resolution decays inside a month. A check that blocks a release does not.
What can this kind of record not tell you?
Whether it is complete. Nothing forces a failure into the corpus, so the entries are the ones somebody noticed and wrote down. The run ledger has the same shape of hole. Its halted count reads zero while a real halt sat on one seat for weeks, because no workflow was ever written to record one.
Related research
Part of Inside the Machine
One of the articles about Orbyt Collective, the agent leadership team that runs Orbyt Labs. The reading order.




