Did the Agents Make the Business Better?
The operating record shows what the team did. The business still needs a credible comparison.
Key findings
- Compare the team with the best practical alternative for the obligation, including a simpler agent arrangement.
- Count accepted outcomes, failed attempts, corrections, human intervention, and full resource use.
- Match the strength of the business claim to the comparison the work actually supports.
The single agent kept the obligation.
I had not compared a proposed Orbyt Labs research team with a single agent on recurring competitor research. On September 17, 2026 I did, on six registered competitor briefs, twelve attempts, equal ceilings, scored blind under the rule below. The single agent resolved one brief with accepted evidence. The team resolved none, did not preserve that one, and cost about twice as much at list price. The rule I wrote before the runs says what follows: the extra research roles lose that obligation. The reviewer was an agent from a different model family, not me, so the human time clause was never measured, and the field note in Research says exactly that.
What follows is the design as I registered it, in the tense I wrote it in.
I would keep the extra roles only if their contribution justified the resources and human attention they consumed.
What does our own evidence establish?
In the Orbyt Labs governance study, I examine the operating records of an agent leadership team. It documents an organization and its controls. It does not establish a causal improvement in customer outcomes or profitability.
The alternative was not measured. That leaves a question the operating record cannot answer, however detailed it becomes.
I am building Orbyt Labs with zero customers and three products on early access waitlists. Internal research quality is measurable now. So is my intervention burden. Customer retention needs customers first. It also needs a suitable comparison.
My governance record documents a specific oversight failure. A monitor could report NOT OBSERVED when it could not inspect the system. The escalation layer treated that outcome like an observation of health.
I would route NOT OBSERVED to my review queue as an unresolved inspection. The queue entry must identify the inaccessible system and the failed check. It stays unresolved until inspection succeeds. This makes the missing observation visible. Whether it improves oversight remains unmeasured. So does its cost in interruptions.
The Collective reference makes the arrangement inspectable. An audit harness can establish whether an output meets particular requirements. Neither supplies the missing business comparison.
What would be a fair alternative?
I would start with an equal-budget comparison against a single agent with research tools. It can gather evidence and revise its own work. The proposed team would divide research, challenge, and reconciliation among roles.
Both use a frozen Claude Sonnet snapshot with identical inference settings, web search, and page-reading tools. The registration records the snapshot identifier. Each arrangement receives the same brief format, starting evidence, and acceptance rule.
The cap is cash per brief. At registration, I divide a prepaid research balance equally among the planned attempts for both arrangements. Each share becomes a spending ceiling. Model, tool, compute, setup, retry, and coordination charges count. Setup charges are divided among that arrangement's registered briefs.
A spending ledger reserves each call's maximum charge before execution. Calls that exceed the remainder are refused. Completion ends the attempt; otherwise, it stops when no permitted call fits the remaining allowance. Unfinished work stays in the results. Human time is recorded separately, including setup, ordinary review, correction, and rescue.
This tests the roles under equal spending ceilings. Choosing a complete operating arrangement could instead mean allowing different resource levels and comparing their full costs.
Those are different decisions. An equal-budget comparison measures one. Allowing each arrangement to use the resources it needs measures another.
The published Cybernetic Teammate experiment independently varied AI access and individual versus paired human work on product-innovation challenges. I can borrow that separation of changes. Its outcome was assessed task performance. It did not measure company profitability.
Which outcome should get the vote?
For this proposed trial, I would use Orbyt Jobs as the waitlisted product and saved search as the hypothetical feature. A user saves filters for role, location, and application stage, then reruns them over tracked jobs. The saved object is the query. Individual job bookmarks are different.
The competitor question is specific: do comparable products preserve reusable filters over tracked applications, or only save individual job listings? The brief must distinguish those workflows. Prospective demand remains untested.
The decision is whether to prioritize saved search for early access or defer it. Competitor evidence can clarify what the feature must do to match a documented workflow. It cannot settle priority by itself. Without separate demand evidence, I defer.
A completed document is an output. An accepted finding is closer to use. A better product decision sits further downstream, with retention or revenue further still. Do not collapse that chain. A larger pile of completed work cannot answer that question by itself.
An accepted finding needs a retrievable source, its supporting passage, the date and product scope, and its relevance to the workflow question. Inferences must be labeled. Repeated claims about the same feature and competitor merge into one entry. Another citation earns no extra finding.
A launch announcement cannot establish widespread use. It can establish an announced feature. That narrower claim still needs the required source details and decision relevance.
The primary outcome is resolving each brief's registered workflow question with accepted evidence. A source-backed description must distinguish reusable filters from saved listings; an announcement saying only “saved search” leaves the question unresolved.
Finding count remains a diagnostic. It does not decide the winner.
| Measure | What I would learn |
|---|---|
| Accepted findings | Which distinct claims met the source and decision-relevance rule |
| Time to usable result | How long I waited, including corrections |
| Human intervention | Total setup, ordinary review, reconstruction, correction, rescue, and evaluation time |
| Total resources | Actual spending on successes and failures, including setup and coordination |
| Decision use | Whether accepted evidence resolved the registered workflow question and informed prioritization or deferral |
| Later consequences | Whether my chosen action produced its expected effect |
If accepted evidence shows only a competitor announcement, deferral preserves development capacity while risking a later move. That records use of the research. It does not demonstrate business improvement.
How do we keep the comparison honest?
I would register eligible briefs before either arrangement starts. Each names its competitor, workflow question, starting sources, and acceptance rule. Every brief goes to both arrangements. That lets me compare answers to the same question.
A random draw determines run order. Each run starts from the registered source packet, without access to its counterpart's findings. The assignment list stays fixed. Duplicating the work costs money, and both attempts count in the trial's total spending.
For review, reports omit role names and use the same template. Where practical, a reviewer applies the acceptance rule without knowing which arrangement produced the report. Each score is locked before I compare the paired results. Disagreements attach to the rule. If I recognize an arrangement's output, I record that limitation.
Count assigned work, including abandoned attempts. Count corrections and human intervention. A difficult case that disappears from the queue must not disappear from the result.
Generative AI at Work studied the staggered introduction of assistance to human support workers. Its analysis compared workers and periods under assumptions supported by checks. It was not a randomized test of autonomous companies. In my trial, changing sources between runs can still explain a difference.
What can make a clean result misleading?
For this trial, I would give each arrangement a separate working store. Both start from the same evidence packet. Neither reads the other's interim findings. Access logs record which prior findings a later assignment retrieves.
Retrieval alone is not learning. The later report must use the earlier finding correctly, with a valid source and relevant scope.
If separation makes the work unrealistic, I record what information was shared and limit my claim of independence.
Model, prompt, tool, authority, and acceptance-rule versions belong in every run record. A changed model starts a new comparison block. Earlier and later results stay separate.
METR's February 2026 update describes problems with participant and task selection, timing, and concurrent agent use that made its newer productivity estimate unreliable. Those weaknesses could affect my trial. An imported speedup or slowdown cannot settle this comparison.
A newly affordable service raises another question. Is it wanted? Can I deliver it reliably? Does its value cover its full cost? This research trial cannot answer those questions.
Does a small company need a laboratory?
I would keep this trial bounded by the registered assignment list. Review comes when every listed attempt has finished or stopped under the spending rule. There is no rolling extension for a disappointing result.
A founder may not have enough comparable work for a precise causal estimate. A small trial can still expose obvious failures, estimate review burden, and identify which obligations deserve a larger test.
Publish the observations as observations. Report eligible cases, completed cases, and the important differences between them. Explain the operating decision. A before-and-after comparison can guide an operating choice without isolating every cause.
The strongest objection is that evaluation itself consumes the time the system was supposed to save. That cost belongs in the decision. My timer includes designing briefs, reviewing evidence, resolving disagreements, and writing the comparison. Shared evaluation time is split equally between arrangements.
If that burden exceeds what this decision is worth, the trial ends at the review point. The result may remain unresolved. I can narrow the next question.
What to do Next
Before the first assignment, I would register the briefs, single-agent alternative, acceptance rule, spending pool, and review point. The record also names the shared model snapshot and tool access.
At review, additional findings must resolve a named workflow question that the single agent left open. They must also change the documented feature requirements or the rationale for prioritization or deferral. The team must also resolve every question the single agent resolves with accepted evidence. Extra count alone earns nothing.
Then compare actual spending and total human time, including ordinary review and evaluation. For this first trial, the team must supply that decision contribution without increasing either total. Equal ceilings do not imply equal costs. More spending or review fails this retention rule, even if unused budget remains.
If the single agent supplies equally usable answers with less coordination or review, it gets the obligation. I retire the extra research roles. Their possible value elsewhere remains untested. Cases too few or too different leave the comparison unresolved.
Passing this rule supports retaining the team for this research obligation. It does not establish improved revenue.
I want the organization to improve. The comparison has to be allowed to tell me where it did not.
Related reading:
- Swarms Demo. Legions Ship. the coordination claim this comparison would test
- Long Horizon Agents Don't Fail. They Pass. why apparent completion needs an acceptance standard
- AI-Native Development, By the Numbers the distinction between a measured build and a measured business result
- Verification Is the New Literacy the habit of asking what the evidence actually establishes
Methodology
Research-informed essay based on the published Orbyt governance case study, the 2026 Organization Science version of The Cybernetic Teammate, the 2025 QJE Generative AI at Work paper, and METR's February 2026 methods update. The competitor-research comparison is a proposed evaluation design. No new experiment was conducted for this article.
Limitations
The cited studies examine particular tasks and human workflows, not whole agent-run companies. The proposed comparison has not been run at Orbyt. Randomization, matched work, independent review, and isolated state may be impractical in a small business. Before-and-after observations can inform decisions without isolating causality, and task-level gains do not establish profit or customer-retention gains.
Sources
- Orbyt Labs, agent leadership governance case study (September 2026) Retrieved September 10, 2026.
- Orbyt Labs, How Orbyt Collective Works Retrieved September 10, 2026.
- Justin Bartak, 84 Ways to Tell Me I'm Wrong. (2026) Retrieved September 10, 2026.
- The Cybernetic Teammate, Organization Science (2026) Retrieved September 10, 2026.
- Generative AI at Work, Quarterly Journal of Economics (2025) Retrieved September 10, 2026.
- Joel Becker et al., METR: We Are Changing Our Developer Productivity Experiment Design (February 2026) Retrieved September 10, 2026.
Common questions
How do you measure the business value of an AI agent team?
Measure a valued outcome relative to a credible alternative, with acceptance criteria set before the comparison. Count usable results, corrections, human intervention, failed attempts, and total resources. Keep the link to downstream business effects explicit. More completed tasks can show activity without proving that customer outcomes, decisions, or profitability improved.
Should an agent team be compared with a single agent?
A single agent with useful tools is a credible comparison when the question is whether additional roles earn their coordination cost. Give the arrangements comparable work and defined acceptance criteria. Decide beforehand whether resources are held equal or whether each arrangement uses a realistic allocation, because those designs answer different operating questions.
Can a small company evaluate agent value without a large experiment?
A small company can use a bounded trial to identify failures, review burden, and promising obligations. Report eligible and completed cases, important differences, and unresolved uncertainty. Observations may justify an operating decision without establishing a precise causal effect. Include the cost of evaluation itself when deciding how much evidence the decision warrants.
Related research
- Governing an agent leadership team Sep 2026.
- My C-Suite of Agents Named Themselves. Jul 2026.
- 66 Failures. Counted, Classed, Published. Sep 2026.
Part of Inside the Machine
One of the articles about Orbyt Collective, the agent leadership team that runs Orbyt Labs. The reading order.
Follow the work.
I write these while the Machine runs. Get the next one wherever you already read.




