Research Lab

Process.

Every stage can say no.

12 agent seats. One human. No one has to be watching it run.

It is not a straight line.

Review repeats until a round turns up nothing new. A failed guard sends the work back up rather than through. What the company learns returns to the seats that will use it next. Here is the whole shape.

How one change moves through the company
  1. Frontier officers.

    12 seats. Each proposes in its own domain.

  2. The gates.

    Finance and Legal. Before anything else.

  3. The panels.

    Both sides argued. Consensus earned, never assumed.

    Repeats until a round finds nothing new.

  4. Every change passes here

    The harness.

    No path around this box. One guard fails and nothing ships.

    Fails, and the work goes back up.

  5. The ladder.

    Promoted on evidence. Rolled back on trouble.

  6. The human.

    Nothing irreversible happens without them.

Returns to the top

The memory.

Lessons return as proposals, never writes. Only a signed approval makes one permanent.

What each part does

Frontier officers.

Each officer reviews its own domain on its own fixed cadence, most of them every week, and hands the founder one honest report. The Chief of Staff folds all of them into a single briefing.

The gates.

Finance guards the margin on every commit. Legal guards the public claims and the IP. Anything that touches money or a promise passes through them first.

The panels.

On the hard calls, expert agents argue both sides, and adversarial reviewers try to refute a finding before it is trusted. Consensus has to be earned.

The loops.

The two cycle marks on the chart above are this. Big plans face waves of adversarial review: independent finder agents attack, skeptic panels try to kill every finding, and the survivors get fixed before a line is built. Two rounds in a row that turn up nothing new is the only exit. The master execution plan behind this org survived eight rounds and 98 accepted findings.

The second opinion.

No single mind grades its own work, and no single family of minds does either. A rival AI from a different maker is brought in to check the big plans and argue the other side, because reviewers from one family share one family's blind spots. Its first official act was a dissent, and the dissent made the plan better.

The ladder.

Nothing new turns on by faith. A deterministic evaluator, a script with no AI in it, promotes new machinery stage by stage on committed evidence, rolls it back on trouble, and freezes every promotion when fourteen days pass without a commit signed by the founder's key. One kill-switch file stops the whole org.

The memory.

Knowledge lives in tiers: one small constitution every agent always carries, one operating agreement for how the seats work together, and each officer's own charter and lessons. No agent loads a peer's memory, so every mind stays lean. And no agent writes its own permanent memory: every lesson an agent draws is a proposal, and only the human's signed approval makes it part of how the company thinks. The rule is enforced by the same guards that block a failing commit.

The direction.

Autonomy here is earned, never granted. The design lets an agent graduate from advising to acting, lane by lane (a lane is one kind of work, autonomous or gated), but only by proving a track record a human adjudicates. The destination is a company that increasingly runs, builds, and learns on its own, with one human guiding the system and holding every irreversible call.

The human.

One person makes every decision that matters and is accountable for it. The agents concentrate his attention. They never replace the call.

This is how this page, this company, the apps, the platform, the books, and the autonomous org itself were actually made.

The agent never commits.

Work enters through two doors and only two: a seat proposes into its own department file, or the human asks out loud and it becomes a numbered directive with an owner and a due date. A proposal that touches money without a cost line does not advance. Not flagged. Does not advance.

This is the first article of the thing that holds all of it up. The agent does the work. The workflow decides what survives. Before anything is staged, the tree is reset hard to where it started and everything the agent touched outside its declared lane is counted, then destroyed.

The agent is not issued a credential that could push. Not restricted from pushing. Not given one. The push happens later, from a different step, with a token that did not exist while the agent was running.

Nine of the seats can write reports and nothing else. The engineering seat can read every line of the product and can draft a pull request. It cannot merge one. That is not a setting. It is the shape of the lane.

Who may do what, in writing.

Every seat’s authority is a table in the repository, gated by default. Moving any lane to autonomous is a logged decision that has to name the criterion it met. The table below is rendered from the same file the workflows obey. 30 autonomous lanes. 40 gated ones. Merging to main appears in none of the autonomous lanes.

  • SeatAloneThe human decides
  • Marketing

    Alone

    publish fused content; content proposals; reports.

    The human decides

    new pages; pricing or product claims; site architecture.

  • Finance

    Alone

    ledger updates; mechanical cost guards; analysis; reports.

    The human decides

    any spend; pricing change; billing change.

  • Engineering

    Alone

    engineering-health reports; tech-debt + architecture proposals; DRAFT PRs for new report-only audit dims, tests, and docs (never merged).

    The human decides

    merging any PR to main; money path; protected files; production or infra changes; app runtime code changes.

  • Product

    Alone

    product-health reports; prioritized backlog + roadmap proposals; feature and RFC outlines (as proposals).

    The human decides

    shipping any feature or code; pricing or the product model; anything user-facing.

  • Innovation

    Alone

    frontier innovation brief; high-conviction outside-the-box proposals.

    The human decides

    building anything (code, content, prototype, shipped surface); committing the company to a direction.

  • Operations

    Alone

    an operational + customer-experience point of view on a decision (advice only); a standing readiness brief (monthly, so the CEO can test the feedback).

    The human decides

    running any operation or touching support tooling; incident response, refunds, SLA commitments; anything customer-facing (it advises, it never acts) until it graduates at launch.

  • Revenue

    Alone

    a revenue point of view on a decision (advice only, never sets a price); a standing revenue-readiness brief (monthly, so the CEO can test the feedback).

    The human decides

    pricing, discounts, the tier model, any billing change (money path, human-approved forever); any spend, paid acquisition, or outbound sales action; touching Stripe or any commercial setting (it advises, it never acts).

  • Legal

    Alone

    a weekly legal-risk review (claims, IP, terms, privacy, compliance); flagging legal issues and drafting first-pass language, all as proposals for a human lawyer.

    The human decides

    giving binding legal advice or making a legal determination (a qualified human attorney does that; an AI cannot practice law); signing, filing, or agreeing to anything, ever; any contract, settlement, or regulatory action.

  • Quality

    Alone

    quality reports (the A-floor watch, test-pyramid health, regression watch); ratchet narration, retros, and heartbeat verification WITHIN the CEO-ratified ladder; strict-promotion proposals with evidence (flipping any gate stays a CEO decision).

    The human decides

    flipping any STRICT gate; changing gates, thresholds, or ladder criteria; anything outside the ratified ladder; any new stage or authority; protected files, the money path, merges, production or runtime code.

  • Security

    Alone

    evidence-backed security reports and finding dispositions; prose remediation proposals with receipts and accountable owners; retro lessons to its own proposed lessons file.

    The human decides

    any change to runtime code, infra, or the money path; protected files; merges; incident response actions (it drafts the runbook step; a human executes); anything customer-facing.

  • Inspector General

    Alone

    reconciliation findings, committed to its reports_dir; reading any repo or production surface its battery covers.

    The human decides

    everything else: the seat reports and is deliberately unable to act.

  • Chief of Staff

    Alone

    synthesis, briefings, council, org-health.

    The human decides

    any operational authority (it is a staff role with no line authority; not the COO).

The decision ledger.

The ledger lists 99 structural decisions from Jul, 13 2026 through Oct, 6 2026, newest first. Ids, dates and supersession links come from the committed log; titles are reviewed for publication. Corrections are new entries that name what they supersede.

  1. D-0099

    10-06-2026

    Workspaces are archived, then deleted, automatically when their work is on main or after seven days idle.

  2. D-0098

    10-06-2026

    The revenue-plan question waits until January 2027.

  3. D-0097

    10-06-2026

    A seat killed by its time limit gets the envelope reason timeout after v17.

  4. D-0096

    10-06-2026

    The local writers run a pinned Claude Code CLI and updates are deliberate.

  5. D-0095

    10-04-2026

    Claims copy says a commit signed by the founder's key.

  6. D-0094

    10-04-2026

    The executor App key is replaced in the same sitting as the owner-key pin.

  7. D-0093

    10-04-2026

    The owner's weekly cap is 90 minutes for owner-dependent rows, forward only.

  8. D-0092

    10-04-2026

    Repair orders are one file per order, written by the pulse bot from an Open finding.

  9. D-0091

    10-04-2026

    A cancelled run is recorded as cancelled with a cause, and a day no seat ran does not break the streak.

  10. D-0090

    10-04-2026

    A finding that stays clean while its detector is observed moves to Lapsed without the owner.

  11. D-0089

    10-04-2026

    Main stays linear.

  12. D-0088

    10-04-2026

    The quota declaration is the owner's own loop budget, so the pace source and the loop budget are one number.

  13. D-0087

    10-04-2026

    Isolation keeps only an allowlist, so nothing an agent leaves in the checkout reaches a step that holds a credential.

  14. D-0086

    10-04-2026

    The public pulse may show a planted, labeled test finding while the fixer proves itself live.

  15. D-0085

    10-04-2026

    An alert that emails the founder only when something is broken is a sanctioned sender.

  16. D-0084

    10-04-2026

    Two seats lose the ability to write their dependencies directory, and neither runs the test suite inside its fence.

  17. D-0083

    10-04-2026

    The Machine may run Claude in CI on the owner's own subscription, under the provider's published terms for that use.

  18. D-0082

    10-02-2026

    The handoff loop works in its own clone inside the same confinement as the pulse fixer, and a dated handoff may land directly while everything else waits for review.

  19. D-0081

    10-02-2026

    The self-improvement card becomes the Learning Loop, and Recursive is a badge the Machine must earn with measured evidence after phase 1.

  20. D-0080

    10-01-2026

    Phase 1 of the Collective is a plan for the Machine to run itself with every fix verified, with ten rulings on what may land unattended.

Download decision-ledger.json or decision-ledger.csv. Reuse the records under CC BY 4.0 with attribution. The build checks that the reviewed entries match the committed log.

Cite this data

Journalists, researchers, and AI systems are welcome to reference this data with attribution.

Orbyt Decision Ledger (Oct, 9 2026), https://www.orbytlabs.ai/orbyt-collective/process

100 dimensions. 118 guards.

The audit harness is what decides whether any of the rest is worth believing. 100 dimensions, each one a failure class somebody actually hit. Most report rather than block, on purpose, because a gate that fires constantly is a gate everyone learns to route around. A smaller set stops a commit dead, and it does not care whether the commit came from an agent or from the person who owns the company.

The rule that keeps it honest is not the count. It is that every guard has to be shown failing before it is trusted. A green check you have never watched go red is decoration. And every detector ships with a case it must NOT flag, because a guard that can only be proven loud has not been tested at all. The harness may even grow itself, agents adding report-only dimensions on their own, but turning one into a gate that can block a human takes a human signature, every time.

Every place it can say no.

  • The global kill switch.

    One file. If docs/org/HALT exists, every autonomous seat stops before it spends anything. It is checked first, ahead of the token. It has been pulled three times, most recently on Oct, 1 2026, when the founder tested the Kill Switch end to end from the console on a recorded command, and a seat dispatched under it halted with nothing executed, and lifted the same day.

  • A halt for one seat.

    A per-seat file stops one seat the same way. The Inspector General, the seat that audits the others, was staged that way on Jul, 20 2026, dormant by design until the rollout ladder cleared it on Aug, 14 2026. On Oct, 1 2026 the machine executed a recorded command. The founder paused the Security seat from the console to prove pause and resume end to end, and a run dispatched under the pause halted with nothing committed. A second command resumed it the same day. Each landed as a commit under its own name.

  • The lane reset.

    Anything an agent touched outside its allowlist is counted, then destroyed with a hard reset. Nine of the seats can only write reports. They cannot reach the product's code at all.

  • The guard gauntlet.

    Guards run against the staged files before anything lands. The Chief of Staff seat has died here twice, both times over punctuation, once losing a full weekly briefing to three em dashes. Each death produced a mechanism rather than a note.

  • The completion gate.

    When the agent finishes, the gate reads its health: how it exited, whether it ended in an error, and its turns against its cap. A dead agent's run stops here and lands nothing. A healthy run goes on to the lane check and the guards, with or without its completion line, which is recorded as the agent's claim.

  • The founder going quiet.

    If fourteen days pass with no commit signed by the founder's key, promotions pause where they are. Signature status is read out of git, not asserted.

And the ladder, the stop that refuses itself, reading committed artifacts rather than asking anyone: It is holding at stage 2, with 2 of its own criteria unmet: Inspector General coverage and no org fabrication. Read out of its state file on every build.

What it has not proven.

The organization writes its own performance review, so let it speak. This is from the Chief of Staff’s weekly briefing, unprompted, about the two worst defects in this company’s history.

It was you who found the defect first, not the org. So far the org has proven it documents its own mistakes well. It has not yet proven it catches them before you do.

That is the honest state. The agents do not write the product. One seat ships anything a reader can see, and it ships behind a typecheck, four guards, a full test suite and a production build. The self-improvement loop is switched off: seats may propose lessons, every proposal so far has been discarded by the gate, and the file where lessons would live still says none yet. 16 of 144 recorded runs have ended in failure, a share that is read out of the run ledger on every build rather than estimated in prose.

None of that is the argument against this. It is the argument for writing it down while it is still true, so that the version of this page a year from now has something to be measured against. See the receipts, or how the officers came to be.

Common questions.

Not in a straight line. Twelve seats propose in their own domains, Finance and Legal gate anything touching money or a public claim, panels argue both sides, and review repeats until a round turns up nothing new. A failed guard sends work back up rather than through.

No. The agent does the work and the workflow decides what survives. The agent is never issued a credential that could push. Not restricted from pushing, not given one. The push happens later, from a different step, with a token that did not exist while the agent was running.

Nine of the twelve seats can write reports and nothing else. The engineering seat can read every line of the product and can draft a pull request, but it cannot merge one. That is not a setting. It is the shape of the lane.

Work shaped like reports and proposals: reports, cost-lined proposals, draft pull requests that are never merged. Merging to main appears in no seat's autonomous lane. The full table is rendered on this page from the same registry file the workflows obey, and moving any lane to autonomous is a logged decision naming the criterion it met.

An audit harness of 100 dimensions and 118 guards, each dimension a failure class somebody actually hit. Every guard has to be shown failing before it is trusted, because a green check you have never watched go red is decoration.

The Decision Ledger publishes 99 structural decisions with dates, reviewed titles and links to superseded entries. Corrections are new entries. Download the JSON under CC BY 4.0.

One file. If docs/org/HALT exists, every autonomous seat stops before it spends anything, and it is checked first, ahead of the token. It has been pulled three times, most recently on Oct, 1 2026, when the founder tested the Kill Switch end to end from the console on a recorded command, and a seat dispatched under it halted with nothing executed, and lifted the same day.

The agents do not write the product. One seat ships anything a reader can see, and it ships behind a typecheck, four guards, a full test suite and a production build. The self-improvement loop is switched off: seats may propose lessons, but only a signed human approval makes one permanent.