66 Failures. Counted, Classed, Published.
Research
A question investigated with a stated method, reporting what was found and what the finding cannot support.
Key findings
- Half of the documented failures in one agent run organization, 33 of 66, were failures of its own tests, guards, gates and metrics rather than of the agents.
- The corpus classifies by what broke rather than by who broke it, so that share is an accounting of one column and never a comparison between agents and instruments.
- A control counts as live only when its refusal has been observed, and a monitor that could not look must report NOT OBSERVED rather than clean.
- Of the 66 documented failures the paper read on 2026-09-02, 42 carry a mechanized countermeasure, which leaves 24 sitting as written lessons rather than as machinery.
- A state file records the last evaluation, not the world, so a paused ladder can keep reporting a condition that cleared days earlier.
Key measurements
| Measure | Value | Source and date |
|---|---|---|
| Documented failures in the corpus | 66, of which 55 are dated, the dated ones spanning 2026-06-09 to 2026-09-02 | September 3, 2026 · The Orbyt Failure Corpus, CC BY 4.0 |
| Failures classed as verification gaps | 33 of 66 | September 3, 2026 · The Orbyt Failure Corpus, CC BY 4.0 |
| Failures carrying a mechanized countermeasure | 42 of 66 | September 3, 2026 · The Orbyt Failure Corpus, CC BY 4.0 |
| Agent seat runs recorded | 65 runs, 55 clearing the completion gate, 2026-07-21 to 2026-09-01 | September 3, 2026 · The Orbyt Autonomy Ledger, CC BY 4.0 |
| Frozen opening baseline | 30 runs and 24 completions, frozen 2026-08-11 | September 3, 2026 · The Orbyt Autonomy Ledger, CC BY 4.0 |
| Signed founder commits in the current liveness window | 53, read by running the evaluator's own git probe | September 3, 2026 · Author's direct read of the Orbyt Labs repository |
Orbyt Labs published its second preprint. Governing an Agent Leadership Team documents 66 failures from twelve agent seats and one accountable human, and classes 33 of them as verification gaps.
Not failures of the agents. Failures of the tests, guards, gates and metrics I built to watch the agents.
The expensive part of running agents is that the things watching them break in ways that look exactly like health.
What does the corpus actually say?
Three classes. 33 verification gaps, 18 operator process, 15 external behavior.
The failure corpus is a public file, generated 2026-09-03 and licensed CC BY 4.0. It defines a verification gap in its own taxonomy: a test, guard, gate or metric that could not fail, could not see, or measured a proxy for the thing it claimed to measure. 55 of the 66 entries carry a date, and the dated ones run from 2026-06-09 to 2026-09-02. The other eleven predate the dating convention. They are published undated rather than backdated.
42 of the 66 carry a mechanized countermeasure. 24 do not.
Why did the instruments fill the corpus?
Because everyone braces for the other failure.
The story you expect is an agent off charter. Writing something wrong, deleting something, spending money it had no business spending. I braced for that too. It is why nothing in the org can push to the branch that deploys.
That is not what filled the file. A green tick did.
A guard that had only ever proved it could be loud, never that it could stay quiet. A monitor whose own crash rendered as clean. A completion gate whose reading window was so small that a successful run fell outside it and got thrown away. A job that concluded cancelled rather than failed, so every alarm keyed on failure was blind to it.
Each was live for days or weeks. Each reported success the whole time.
What law came out of it?
Positive observation. A control is live only when its refusal has been observed.
Not when it is configured. Not when a document says it runs. Not when the job it sits inside comes back green. A guard you have never made go red is decoration.
The second half is the one I underrated. A monitor that could not look must report NOT OBSERVED. That is a third verdict. Neither clean nor red.
Same rule as proving a guard by injecting the bug it exists to catch. Same rule as counting how often a repair fires rather than whether it succeeded. I keep arriving at it from different directions, which usually means it is the real rule.
What does the record refuse to claim?
More than I wanted it to.
It does not say instruments fail more often than agents do. The corpus classes by what broke, never by who broke it, and 18 of the 66 are the operator's own process failures. Those are mine. So 33 of 66 is the composition of one accounting. It is not a scoreboard.
There is no customer outcome in it and no revenue outcome. The receipts are runs, failures, decisions, guards and dates.
It is not a trend either. The autonomy ledger records 65 seat runs between 2026-07-21 and 2026-09-01, 55 of them clearing the completion gate, against a baseline frozen on 2026-08-11 at 30 runs and 24 completions. Those two figures share 30 rows. Three run outcomes flipping the other way would erase the whole difference between them.
One organization. One quarter. One rater, and the rater is me.
Where is the argument weakest?
In my own liveness monitor, and not where I first looked.
The ladder state file was last evaluated on 2026-09-01. Its liveness field reads alive false. The reason string reads no signed founder commit in 14 days, or signing not yet configured. Promotions are paused because of it.
That was true the day it was written. Signed commits from my address land daily through 2026-08-13, then stop until 2026-09-01.
Then I ran the evaluator's own git probe by hand on 2026-09-03. It counts 53 signed commits from the founder address inside the window. The condition cleared days ago. The file still says paused, because a state file records the last evaluation and not the world. I did not see that until I ran the probe.
The defect that survives is smaller and it is still real. That reason string bundles two causes into one sentence, and a reader cannot tell which one fired. The evaluator does better where it can. It returns an explicit NOT OBSERVED when the git history is unreadable, because a negative answer there means could not look.
The drill that would settle the rest is written down and unused. Once a quarter, withhold the heartbeat and confirm the absence alarm fires on both channels. The line for logging that drill has never been filled in. By my own rule, that is a hope standing where I wrote a control.
Is any of this transferable?
The strongest objection is that it is not. One company, one human, agents doing content and process work with no customer on the other end. My own view is that a larger organization has failure modes this corpus cannot contain.
Two things survive the objection.
The taxonomy is the first. Verification gap, operator process, external behavior. Those classes did not need an agent org to produce them. Any team running scheduled automation can sort its last quarter into them in an afternoon, and the sorting is the part that teaches.
The test is the second, and it costs nothing. Ask each control you own when its refusal was last observed. Not whether it is configured. When it last said no, and whether anybody saw it.
Most of mine could not answer. That is why 24 of the 66 are still written lessons, and a lesson is a hope where a guard is a fact.
What to do Next
Sort your last quarter of incidents into the three classes. If the verification column comes out thin, treat that as a labelling problem before you treat it as good news. Instrument failures are the ones nobody files. Nothing went down. Something failed to notice.
Then pick the control you would most confidently swear is working, and make it go red. Break its input. Revoke the permission it depends on. Feed it the exact thing it exists to catch. If you cannot make it refuse, you do not have a control. You have a habit with a dashboard.
The paper is free to read. It is a preprint, it has not been peer reviewed, and it has carried a DOI, 10.5281/zenodo.22683647, since its Zenodo deposit on September 9, 2026. The first one has carried one since April.
A control you have never seen refuse is not a control. It is a green light nobody has tested.
Related reading:
- The Machine. It Runs the Company. the twelve seats, the ladder and the kill switch this paper formalizes
- 84 Ways to Tell Me I'm Wrong. the guard estate the failure corpus keeps adding to
- Self-Healing Is a Euphemism. why a repair that never reports is a failure you stopped seeing
- 391 Yeses and Not One No. the permissions blast radius behind the human-only line
- Codex Accuses. Claude Convicts. a second model standing as the check on the first
- Verification Is the New Literacy why proof, not production, became the constraint
Methodology
Every figure was read on 2026-09-03 from the artifact that produces it rather than from a document describing it. The failure counts, the dated count and the three class definitions come from public/failure-corpus.json, generated 2026-09-03. The run counts and the frozen baseline come from public/autonomy-ledger.json, generated the same day. The liveness state and its reason string come from a direct read of docs/org/ladder-state.json, and the 53 signed founder commits come from running that evaluator's own git probe by hand on 2026-09-03 rather than from the state file.
Limitations
This is one organization over one quarter, grading its own record, which the published corpus states in its own caveat. The corpus classifies by what broke rather than by who broke it, so 33 of 66 describes the composition of a single accounting and cannot be read as evidence that instruments fail more often than agents do. Eleven of the 66 entries carry no date. The run figures cover agent seat runs only and report no customer outcome and no revenue outcome. The current completion count and the frozen baseline share 30 of their rows, and three run outcomes flipping would erase the difference, so it is not a trend. The paper is a preprint. It has not been peer reviewed. Its Zenodo deposit landed on September 9, 2026, so it carries a DOI, 10.5281/zenodo.22683647.
Sources
- Governing an Agent Leadership Team (preprint) Retrieved September 3, 2026.
- The Orbyt Failure Corpus Retrieved September 3, 2026.
- The Orbyt Autonomy Ledger Retrieved September 3, 2026.
Cite this research
Bartak, Justin. "66 Failures. Counted, Classed, Published.." The Machine Speaks, Orbyt Labs, 2026. https://www.orbytlabs.ai/blog/66-failures-counted-classed-published
Common questions
What does the governance paper document?
It documents 66 failures from one agent run organization and classes 33 of them as verification gaps, meaning a test, guard, gate or metric that could not fail, could not see, or measured a proxy. The other classes are 18 operator process and 15 external behavior. The corpus sorts by what broke, never by who broke it.
What does positive observation mean for a control?
A control counts as live only when its refusal has been observed. Configured is not live. Documented is not live. A green job is not live. The corollary is a third verdict state: a monitor that could not look reports NOT OBSERVED rather than clean, because blindness and health are different events and only one of them is good news.
What does this record say it cannot show?
Four things, stated in the corpus itself. It covers one organization over one quarter. The labels were assigned by the operator grading its own mistakes, a single rater. It reports no customer outcome and no revenue outcome. Its two completion figures share 30 rows, so three flipped run outcomes would erase the gap between them.
How do I test a control I already own?
Make it go red on purpose. Break its input, revoke the permission it depends on, or feed it the exact thing it exists to catch. If you cannot make it refuse, you do not have a control. Then check what it prints when it cannot look at all, which is the harder half of the job.
Related research
- Governing an agent leadership team Sep 2026.
- My C-Suite of Agents Named Themselves. Jul 2026.
- I Ran 830 Agents in One Long Horizon Session. Jul 2026.
Part of Inside the Machine
One of the articles about Orbyt Collective, the agent leadership team that runs Orbyt Labs. The reading order.
Follow the work.
I write these while the Machine runs. Get the next one wherever you already read.




