Governing an Agent Leadership Team: Guardrail, Kill-Switch, and Liveness Patterns from a One-Human Autonomous Organization.
By Justin Bartak. ORCID 0009-0005-2615-3624
DOI: Zenodo deposit pending. A DOI will be added on deposit.
Abstract
Half of the documented engineering failures at one company running an agent organization, 33 of 66 in a single-rater, agent-coded corpus, were failures of its own instruments: tests, guards, gates or metrics that could not fail, could not see, or measured a proxy. We operate the organization studied: twelve named seats, nine running a language model, hold a reporting, proposing and publishing cadence under one accountable human. The corpus (`public/failure-corpus.json`, 2026-09-02) spans 2026-06-09 to 2026-09-02 with eleven undated items; 42 of 66 carry a named countermeasure, and verification gaps are the largest class at 33.
The ledger (`public/autonomy-ledger.json`, 2026-09-02) records 65 seat runs over six weeks, 2026-07-21 to 2026-09-01, at 84.6 percent completion, six of the 55 completions being panel-written observer rows; the frozen 80.0 percent baseline shares its 30 rows, and the 4.6-point gap is within two runs' movement. Forty-five decisions went to an append-only log (`docs/org/decision-log.md`); the organization kill switch was pulled once, on 2026-08-20. The corpus has no agent-failure class and one rater assigned every class, so the share is exploratory rather than a comparison.
The design law that survived is positive observation: a control is live only when its refusal has been observed, and a monitor that could not look reports NOT OBSERVED rather than clean. That verdict is printed; at the layer that escalates, a monitor that could not look and one that saw health are the same event. We contribute a three-tier guardrail model, positive-observation liveness, the published corpus, ten design principles, and a retrofit checklist.
Key findings
Five results from the record, ranked by load-bearing weight on the paper’s thesis. Every number carries the date it was measured.
Reproducibility package
The paper reads three published data files, each generated from the organization’s own records and served from this site at a stable URL under CC BY 4.0. The source repository is private, so these three files are the whole surface a reader can re-count without access, and every table in the paper names the command that produced it.
Cite this paper
Permissive (CC BY 4.0) attribution. Cite the site URL for now. A DOI will be added here and in both citation forms on Zenodo deposit; the URL does not change.
Plain text (APA-style)
Bartak, J. (2026). Governing an Agent Leadership Team: Guardrail, Kill-Switch, and Liveness Patterns from a One-Human Autonomous Organization. Preprint. https://www.orbytlabs.ai/research/agent-leadership-governance
BibTeX
@misc{bartak2026governing,
author = {Bartak, Justin},
title = {Governing an Agent Leadership Team: Guardrail, Kill-Switch,
and Liveness Patterns from a One-Human Autonomous Organization},
year = {2026},
note = {Preprint},
url = {https://www.orbytlabs.ai/research/agent-leadership-governance}
}Where this paper lives
This page is the canonical owned URL with full Schema.org structured data. The permanent DOI arrives with the Zenodo deposit, which is a human step because a DOI cannot be withdrawn.
License
Released under Creative Commons Attribution 4.0 International (CC BY 4.0). You may copy, redistribute, remix, transform, and build on the paper for any purpose, including commercially, with attribution.
The three published data files the paper reads are also CC BY 4.0. Quote, redistribute, build on, with attribution to Bartak, J. (2026) for the paper and Orbyt Labs for the data.
Related on this site
Read the full paper.
The full preprint covers the machine, the guardrail tiers, the watch chain, the empirical record, ten design principles and a retrofit checklist, with three appendices.