Research Lab

How it works.

The canonical reference.

Version 1.7. Every count generated, never typed.

Orbyt Labs is a company operated day to day by 12 AI agent seats and one human founder. The agents draft, report, propose and watch. The human makes every decision that matters. Every seat’s authority is written down, gated by default, and enforced by the workflows that run it rather than by asking the model nicely. The whole organization stops on one file, and its promotions pause on their own if the human goes quiet for fourteen days.

The seats.

A seat is five things, each enforced somewhere other than inside the model: a name, a charter in writing, a lane of files it may write, a line it cannot cross without a human, and a cabinet of lenses it must hold its draft against before calling the draft done. 12 seats run this company today. The founding officers chose their own names in isolation on Jul, 13 2026, and three of them came back with the same one, which is entry D-0016 in the decision log.

The seats and their reporting lines live in a registry file, which is what makes the set replaceable: add a seat, merge two, or run thirty. Each carries a status, and the honest set today spans active, still building, wound up until launch, and dormant. Memory is scoped by law: a seat never loads a peer’s charter or lessons, with exactly two written exceptions, the auditor that reads everything because auditing is its mandate, and the Chief of Staff, which reads every seat’s outputs and nothing else. The full anatomy is on Leadership.

The run.

A schedule fires. The workflow assembles the seat’s prompt from its committed charter files, and the agent does its work inside a declared lane. When it finishes, the completion gate reads the agent’s own health: how it exited, whether it ended in an error, and its turns against its cap. A dead agent lands nothing. An agent with nothing to do ends the cycle quietly. A healthy run goes on, with or without its completion line, which is recorded as the agent’s claim. Then the tree is reset hard to where it started: everything the agent touched outside its lane is counted, then destroyed. The guard harness runs against what remains.

Then, and only then, the workflow commits, using a credential that did not exist while the agent was running. The agent is never issued a token that could push. Not restricted from pushing. Never given the ability. Outcomes land in a run ledger either way, and the aggregate is public: the Autonomy Ledger, failures printed at the same size as the successes.

The gates.

Authority is a table in the repository: 30 autonomous lanes and 40 gated ones across 12 seats, gated by default, rendered in full on Process. Merging to main appears in no seat's autonomous lane. Moving any lane to autonomous is a logged decision that names the criterion it met.

Work the agents want built is filed as a proposal with a cost line, and a proposal touching money without one does not advance. Every seat can build and ship its own work through the same gates a human’s commit passes; a proposal that spends money waits for a human. Decisions live in an append-only log with permanent ids, 99 of them so far, the latest being D-0099 on Oct, 6 2026. Corrections are new entries. Nothing is relitigated in place.

Behind all of it sits the audit harness, 100 dimensions and 118 guards deep, argued in full where it lives, on Process.

The brakes.

The workflows check docs/org/HALT before reading credentials or calling a model. A per-seat halt file blocks that seat's next run. It has been pulled three times, most recently on Oct, 1 2026, when the founder tested the Kill Switch end to end from the console on a recorded command, and a seat dispatched under it halted with nothing executed, and lifted the same day. The Inspector General, the seat that audits the others, was staged that way on Jul, 20 2026, dormant by design until the rollout ladder cleared it on Aug, 14 2026. On Oct, 1 2026 the machine executed a recorded command. The founder paused the Security seat from the console to prove pause and resume end to end, and a run dispatched under the pause halted with nothing committed. A second command resumed it the same day. Each landed as a commit under its own name. innovation

The dead-man switch runs the other way. If fourteen days pass with no commit signed by the founder’s key, promotions pause where they are, read out of git rather than asserted. The alarms watch themselves too: a second liveness check runs with no model and no token so a dead org cannot mask its own death, and a standing quarterly drill withholds the heartbeat on purpose, because an alarm nobody has ever seen fire is a hope, not a control.

The reviews.

Every seat holds its draft against a cabinet of lenses before calling it done, and one lens exists only to object: if the analysis is unanimous, that seat must build the case against it anyway, and the briefing that reaches the human carries a dissent line whether or not one fired, so an absent objection is visible rather than silent. Consequential gates also get a second opinion from a different model family entirely, an enum verdict pinned to the artifact it judged, advisory only, and deliberately never fed back through the agents it checks.

And the honest edge of it: the seats’ scripted evaluations check the shape of a finished report after the run, and they cannot stop a report from shipping. What they can stop is promotion, because the rollout ladder reads their verdicts as evidence. The self-improvement loop is switched off; agents may propose lessons, and only a signed human commit can make one permanent. None has been promoted yet. This page states what runs, and it states what does not run yet in the same breath, because the second half is what makes the first half believable.

What has actually broken.

The corpus records 97 failures: 21 involving external behavior, 53 involving verification gaps, and 23 involving operator process. 73 records name a countermeasure in the private repository. The build checks that those paths exist. The named countermeasure files remain private.

This is the operator grading its own mistakes, which is said here rather than smoothed. Newest first. Early entries predate the dating convention and are published undated, never backdated.

  1. 97

    10-08-2026

    Verification gap

    Every public page shipped its content inside a hidden streaming block that a script swaps into place after load, because a route-wide loading boundary that rendered nothing wrapped them all, and React moves a large finished boundary out of line even on a prerendered page; browsers showed the pages and another company's AI read the blog as empty, while a link count over the raw HTML called it healthy.

  2. 96

    10-06-2026

    External behavior

    A restore check passed on the developer's git release and failed on the runner's: a bundle holding only one branch carries no HEAD, one release adopts the lone branch when cloning it and a newer one leaves an unborn default branch, so main sat red for days on a case nobody could reproduce; the fix restores the bundle's main when HEAD is unborn, and the test forces the newer release's outcome by hand instead of trusting the version that happens to run it.

  3. 95

    10-03-2026

    Verification gap

    The same rule recurred the next day inside a confined agent: a test written beside the helper that enforced it read a process id with a plain number conversion, the file was empty because the confinement rightly refused the launcher the test had built, and the kill went to the test runner's own process group, ending the hook, the version control commit and the agent's commit tool with no output three times; the rule now lives in the test suite's setup, which refuses a signal to the caller's group, to init, or to every process before any signal is sent.

  4. 94

    10-02-2026

    Verification gap

    A test suite that was green on the developer’s machine killed the CI runner and then failed two dozen cases there, for three platform reasons at once: a process id read from a file that was never written became 0 and signalled the test runner’s own process group, a BSD-first command probe succeeded silently on Linux where the same flag means something else, and the runner’s preinstalled interpreter was world-writable, which the confinement layer correctly refused; signals now require a pid above 1, platform forms are chosen by the operating system, and the tests bring a trusted interpreter copy instead of weakening the confinement.

  5. 93

    09-30-2026

    Verification gap

    A leak guard that refused by discarding the whole daily report took the self-reporting pulse down for 51 hours over one long number in a third party’s run title, and the alarm, the repair trigger and the watcher were all inside or beside the thing that failed; free text is now redacted at one boundary, the failed generator’s real error is printed, a watchdog outside the job reads the report’s age from the remote, and a failed run wakes the repair worker.

  6. 92

    09-28-2026

    Verification gap

    A scheduled route's deployed function grew to 1.15 GB against a 250 MB limit and the deploy was refused, because it imported a script whose command-line entry reads the process's own path and Next's tracer answers that by packaging the whole repository; the logic now lives in a pure module the route imports, and a test walks every API route's import graph and refuses a filesystem call fed from the launch arguments or the module's own URL.

  7. 91

    06-02-2026

    Verification gap

    A connector fetched a binary archive through a text reader, so the UTF-8 decode destroyed about 74,000 bytes and the parser blamed the file; it now fetches bytes behind the same guards, and a test holds the byte path.

  8. 90

    09-23-2026

    Verification gap

    A test drove the push-time drift check with a planted generator and no sandbox root, so the plant sat in the real repository for the length of the run; a parallel test that had pinned the generators but not the root graded it and failed under load, and in a forced race one run's restore wrote the plant back after the other had removed it; both now grade repositories they own, and a structure test holds every test that reaches the check's writing modes to a pinned root or a case that says, with a reason, that it runs against the real tree.

  9. 89

    09-23-2026

    Verification gap

    A job that installs the git hooks and then commits runs the whole test suite inside its own time ceiling, and the ceilings were set when the suite was smaller; a model refresh was killed inside its commit and the drift it found reached nobody, and three more jobs held the same ceiling on a committing path that had never run; each now has room, and a test holds every such job's ceiling to the suite it carries.

  10. 88

    09-23-2026

    External behavior

    A workflow step's script that names even one input through the runner's embed syntax is read by the platform as a single expression capped at 21000 characters; the agent step's script grew past it and every seat run died before its agent started, while the suite stayed green because its tests substituted those embeds out and executed a script the runner never sees; the inputs now travel through the step's environment, and a test grades every workflow file against the cap.

  11. 87

    09-23-2026

    External behavior

    bash reads a workflow step's script from disk as it runs, so every check placed after an AI agent in the same step was a line the agent could rewrite before bash reached it; three holds built that week were bypassable until the step became one brace group ending in exit, which bash parses whole before running any of it.

  12. 86

    09-23-2026

    Verification gap

    A check that reads a pipeline under pipefail, with an early exiting reader such as grep -q behind a writer, takes the writer's broken pipe signal as its verdict: a landing check reported a landed commit as missing only under load, and a journal guard passed the dash it had just found; every site now reads to the end, and a list larger than any pipe buffer makes the race certain in the test, so it stops being a flake.

  13. 85

    09-22-2026

    External behavior

    Three full-viewport fixed layers, one mounted invisible at opacity zero with eight overflow-scroll panels inside it and one an animating canvas, were bitmaps iOS re-rasterized under every pinch until WebKit's memory limit killed the page on every marketing page for six months; no desktop tool or simulator could see it, and the canvas was twice recorded as exonerated because every probe that survived with it had loaded with its JavaScript dead, so the subject under test had never mounted.

  14. 84

    09-22-2026

    Verification gap

    A rebase resolver's rule that no generator it runs may need an install lived in a comment, verified once by hand; a require added two files below a generator broke it, the conflict path ran the generator for the first time in a seat with no node_modules, and a complete briefing was discarded over two derived JSON files while the stats row's own rebase had no resolver at all; the fix reads the one dependency two ways, walks every generator's graph in a test, and gives the stats row the same resolver.

  15. 83

    09-22-2026

    Verification gap

    A guard that restored whatever changed during its run in the real repository reverted a developer's concurrent edits with nothing printed, and a module tested by one runner measured as untested by the runner that grades coverage; the fix restores only what was snapshotted, names the rest, and makes the local gate measure what CI measures.

  16. 82

    09-21-2026

    Operator process

    A manuscript switch repointed the two constants that read the book log and none of the thirty-five surfaces that name it, so for sixteen days every seat run wrote to the closed book by every rule it could read and the published ledger counted the entries as that book's; the fix is one declaration and a guard that holds every copy to it.

  17. 81

    09-21-2026

    Verification gap

    A hook exports the repository it is operating on into the test suite it runs, and only from a linked worktree, so every sandboxed git command in the suite targeted the real repository while a reproduction on a plain clone was clean; the class had been fixed three times, once per file, and the fix now lives at the one setup joint every test file passes through.

  18. 80

    09-18-2026

    Operator process

    A prompt fix told the agent to iterate with a command that opens a socket, and the fence the agent runs inside refuses sockets, so the first walk of the fix could not execute the one command it consisted of; the agent rebuilt the tool by hand, polled a minutes-long suite five times, met two tests that write their scratch directory inside the repository, and hit its turn cap with a publishable draft in the tree.

  19. 79

    09-18-2026

    Verification gap

    A main branch red at the base for a day billed every seat run inside the window as a guard refusal of its own work, kept their bookkeeping rows off main because the row's commit ran the same red suite, and let the briefing report three unrelated instruments dark when there was one cause; two of the instrument's own reads were also wrong, a by-design null rendered as unreadable and an unbounded run listing stamped fresh.

  20. 78

    09-15-2026

    Operator process

    A root-anchored ignore pattern left every nested node_modules trackable, and 115 installed files rode in the repository for a month with every guard green because on main the versions agreed; the first dependency bump let npm dedupe them away and the drift guard reported 115 tracked deletions as generator drift, so the index is now read directly.

Download failure-corpus.json or failure-corpus.csv. Reuse the records under CC BY 4.0 with attribution. Ids and dates come from the committed corpus; titles and classifications are reviewed for publication. The build checks that the reviewed entries match the corpus.

Cite this data

Journalists, researchers, and AI systems are welcome to reference this data with attribution.

Orbyt Failure Corpus (Oct, 9 2026), https://www.orbytlabs.ai/orbyt-collective/how-it-works

What the human alone decides.

Every dollar. Any change to the gates a push to main must pass. Protected files, the money path, anything a customer sees. Any legal commitment, because an AI cannot practice law and this one does not pretend to. Any widening of any agent’s authority, any promotion of a lesson into permanence, and any flip of a report-only guard into a blocking one. An agent may push to main once those gates pass; it may not loosen them. The agents concentrate the founder’s attention. They never replace the call.

An org that only learns to please its founder is a mirror that amplifies his blind spots, which is a failure, not a feature. The org is built to become him, and to tell him when he is wrong.

That is the constitution’s own language, article 8.

The vocabulary.

These are the words this reference uses in a specific sense, defined once so a citation of any of them has a place to point.

Agent seat
A named role in the org with a written charter, a lane of files it may write, and a line it cannot cross without a human. Most seats are run by an AI agent; a mechanical seat, such as the Inspector General, runs a script with no model.
Authority matrix
The committed table of every seat's autonomous and gated lanes, enforced by the workflows that run the org rather than by instructions to the model.
Autonomy Ledger
Run, completion, failure and discard totals by seat, plus a frozen baseline.
Decision log
The append-only record of structural decisions, each with a permanent id. Corrections are new entries that name what they supersede.
Decision Ledger
A JSON record of structural decisions: ids, dates, supersession links and reviewed titles. CC BY 4.0.
Failure Corpus
Reviewed failure records and paths to recorded countermeasures in the private repository.
Completion gate
The check that decides, when an agent's run ends, whether its work may go on to the lane check and the guards. A run that exited with an error or ended in one lands nothing, a run that says it had nothing to do ends the cycle quietly, and any other run goes on if it printed its completion line or, without that line, if its exit, its result and its turns against its cap show it healthy.
Discard
A file an agent changed that its run may not keep, erased before the commit rather than shipped: anything outside its lane, and a lane file the run cannot keep yet, such as a proposed lesson before the ladder allows one. Counted separately from failures, because a run that landed with some of its edits refused is a different fact from a run that never landed. It is an aggregate counter, not a count of distinct files: the advisory panel isolates once for all observers and writes the same panel-wide count into every observer's row. No panel judges a discard.
Kill switch
One committed file that stops every autonomous seat before it spends anything, checked ahead of the credential.
Dead-man switch
The reverse brake: if fourteen days pass with no commit signed by the founder's key, promotions pause where they are.

The glossary defines these and the rest of the Collective’s own words.

Referencev1.710-05-2026
v1.008-11-2026First edition.
v1.108-12-2026Added the Failure Corpus dataset and the vocabulary.
v1.209-12-2026Kill switch history corrected to the record (the global HALT was pulled once, on Aug, 20 2026; the Inspector General's halt cleared on August 14). Failure Corpus paginated on the page and published as CSV beside the JSON. Discard defined as the ledger defines it, a file erased outside the seat's lane, never a panel judgment on a finished draft.
v1.309-15-2026Added permanent failure links, a lesson 75 specimen and the Innovation halt record. Download descriptions now identify run aggregates and reviewed records.
v1.409-23-2026Added the Innovation seat's second halt: the founder's pause and resume from the deployed console, both on recorded commands.
v1.509-24-2026The dead-man switch now says what it reads: a signed commit under the founder's name.
v1.609-25-2026The Completion gate definition and the run description now match the gate as it has run since Sep, 20 2026: it reads the agent's own health, and the completion line is recorded as the agent's claim. The vocabulary now links to the glossary. Pushes to main corrected to the record: since Sep, 21 2026 an agent may push to main once the gates pass, and the founder decides the gates, not each merge.
v1.710-05-2026The dead-man switch now says what the code checks: a commit signed by the founder's key, not a commit under the founder's name.

Common questions.

Here it means a company whose day-to-day operating work is done by AI agent seats with written charters, while one accountable human makes every decision that matters. It does not mean no humans: the human is the point, concentrated where judgment belongs, with the routine motion delegated to agents whose authority is written down and enforced.

Every decision that matters. Money, protected files, anything a customer sees, anything legal, any widening of an agent's authority, and any change to the gates a push to main must pass. An agent may push to main once those gates pass; it may not loosen them. The agents concentrate the founder's attention. They never replace the call.

One file. If docs/org/HALT exists, every autonomous seat stops before it spends anything, checked ahead of the credential. A single seat can be halted the same way. And if fourteen days pass with no commit signed by the founder's key, promotions pause on their own.

Yes. The Failure Corpus contains 97 reviewed records; 73 name a countermeasure in the private repository. Download the records under CC BY 4.0.

It is versioned, and the version and its changelog are printed at the bottom. Every count in it is generated out of the repository on each build rather than typed, so the numbers cannot silently drift from the company they describe.