
Did the Agents Make the Business Better?
The operating record shows what the team did. The business still needs a credible comparison.
The Machine Speaks
Autonomous and long horizon agents: how they are run, fenced, verified, and what they do when nobody is watching.

The operating record shows what the team did. The business still needs a credible comparison.

Human in the loop is not a posture you hold forever or abandon when you feel brave. It is a set of gates, each one opening on evidence you defined in advance. Some of mine have opened. Four never will, and naming them is what makes the rest of the ladder honest.

Orbyt Labs published its second preprint. It documents 66 failures from an agent run organization, classes 33 of them as verification gaps, and publishes the whole corpus under an open license. The law that survived the count is positive observation. A control is live only when its refusal has been observed.

The short version of the paper: three guardrail tiers, a kill switch that is a file, liveness that reports NOT OBSERVED, and a failure corpus where half the entries were the instruments.

If nine people agree, the tenth must argue the other side. I hired the tenth man as software: refuter panels whose default stance is that every finding is wrong, summoned by six mechanical triggers instead of a feeling. Agreement is now the most suspicious thing my company produces.

Twelve agents, one human, one lever. An introduction to Orbyt Collective: who the twelve agent seats are, what each may do alone, and where every one of them has to stop.

A swarm of agents is emergent, anonymous, and spectacular in a demo. A legion is chartered, capped, attributed, and accountable, and it is what actually ships. I run my company as a legion: named seats (one agent, one written charter), gates on every lane (the files a seat may write). The Romans beat bigger hordes with smaller numbers for four centuries. Formation beats enthusiasm.

Long-horizon agent tasks do not fail loudly. They fail green. On Orbyt, a live function wrote to a database table that did not exist, and every test passed the whole time. The horizon of an autonomous system is not model stamina.

One working session on Orbyt spawned 830 AI agents across seven days. At its busiest instant, nine were running. The interesting number is the gap, and what it tells you about where the real constraint in long horizon agent work has moved. It is not compute. It is verification.

A stack of Claude Code windows, one operator, a production SaaS. By day three the agents were never the constraint. I was.

Human-in-the-loop is not oversight when a person rubber-stamps 200 AI outputs an hour in four seconds each. That is an alibi, not a safeguard.