Research.

What we are learning by building a company that runs on agents, a harness of checks, and governance.

Papers and field notes for engineers building the same thing.

Orbyt Labs is a company where AI agents do most of the engineering. They write the code, run the audits, review each other, and get stopped by guardrails when they reach for something they should not have. We are keeping the receipts: the autonomy ledger, the decision ledger and the failure corpus, all public.

This section is where that work gets published. Not the product, and not career advice. The measurements, the failures, and the patterns that turned out to hold. Every claim carries the date it was measured, because a number without one is a number that quietly stops being true.

Papers come with a method and a DOI. Field notes come with instructions. Where a result depends on our private repository, we publish the pattern and the measurement and keep the configuration to ourselves.

The machine, as measured on Oct, 9 2026

118

automated guards

290

org documents

97

logged failures

100

audit dimensions

Counted in the Orbyt Labs repository on the date above. These figures move as the system grows and are not restated as current. Each paper below carries its own measurement date.

Papers

Formal work with a method, a measurement date, and a DOI.

Field notes

Reproducible practices, written so you can run them on your own project.

Open questions.

Untested:
Not yet tested.
In trial:
Protocol frozen; results pending.
Measured:
A published study answers the question within its stated limits.
  • Whether any of it holds at scale. Everything you can measure here was measured with almost no customers on it.

    Untested. As of .

    Nothing measured yet.

    What would change our view: Measurements with customers using the products at scale would test whether the findings still hold.

  • Whether agents keep a standard when the work is boring. They hold it on the interesting days. The test is the thousandth unglamorous one.

    Untested. As of .

    Nothing measured yet.

    What would change our view: Repeated routine work would test whether agents maintain the same standard beyond interesting days.

  • Whether taking myself out of the middle makes the work better or only faster. Those are not the same result and I cannot yet tell them apart.

    Untested. As of .

    Nothing measured yet.

    What would change our view: Better work with less founder involvement, rather than faster work alone, would resolve the distinction.

  • Whether a company that can tell its founder no will actually be listened to. By me. That one is not a question about AI.

    Untested. As of .

    Nothing measured yet.

    What would change our view: Following a recorded refusal from the company would show the founder listening.

  • Whether anyone but me and these agents could maintain it. Software built this fast is only an asset if a second person can pick it up, and no second person has tried.

    Untested. As of .

    Nothing measured yet.

    What would change our view: Someone other than the founder and these agents successfully maintaining the software would answer the handoff question.

  • Does a newer model improve one agent's work under a written charter on the same task fixed before the trial?

    In trial. As of .

    The first comparison block is closed: read blind by an agent reader, 0 of 3 claude-sonnet-4-5-20250929 reports and 0 of 3 claude-sonnet-5 reports were accepted, and the candidate did not win. A second comparison block, the current model against claude-opus-5-5 on the same frozen packet under the same rule, is pre registered and partly run, and none of its reports has been read, so it has no result yet.

    What would change our view: If both models pass each substantive check on each run and the newer model needs less total founder review time, that would support keeping the upgrade for that task. A repeated documented failure would keep the current model.

    Founder's role: His plan, written before the runs and queued for publication, contains the acceptance rule; he chose an agent of another model family to read the reports blind.

  • Does a team of agent roles resolve research questions a single agent leaves open, with the same tools and spending limit per brief and no increase in spending or human time?

    Measured. As of .

    Read blind by an agent reader. Of 6 briefs, the single agent resolved 1 and the team 0. The team resolved 0 briefs the single agent left open, so it was not retained under the pre registered rule. The effect of the decision the result informed is not yet observed.

    What would change our view: A team resolving questions the single agent leaves open under the same tools and ceiling, without more spend or human time, would support the team's added value.

    Founder's role: His plan, queued for publication, includes the briefs and the win condition; registration came before the first run, and he chose an agent of another model family to read the reports blind.

Independent use and critique.

None recorded. We list independently authored citations, implementations, replications and critiques in the other party's own words. We exclude search indexing, our own deposits and our discovery probes.

Inside the Machine: essays and explainers.

Editorial writing about the system these papers measure. Interpretation, not evidence.

All Collective articles, in reading order

The Machine itself.

The running system these papers measure.

Orbyt Collective

The agent leadership team, the audit harness, and the governance layer that decides what an agent is allowed to do. This is what the papers measure.

Take a look

Common questions.

What does Orbyt Labs publish research about?

Orbyt Labs publishes work on building and operating an AI-run engineering organization: how agents are given work, how their output is verified, what it costs to run them at scale, and the ways they fail. It is written for engineers building the same thing, not for customers of Orbyt Labs' products.

Is Orbyt Labs research peer reviewed?

Papers are self-published preprints with DOIs, not journal submissions. Each carries its method, its measurement date, and the data needed to check it. The governance paper publishes the files it read. The datasets paper's evaluation code is in a private repository, available on request. Field notes are reproducible practices rather than formal papers, and they are labeled as such.

Can I reproduce Orbyt Labs research?

That is the intent. Papers ship with a reproduction repository and a dataset DOI where one applies, and field notes are written so a reader can run the practice on their own project. Where a result depends on a private repository, the pattern and the measurement are published and the configuration is not.

How is Orbyt Labs research different from the Orbyt blog?

Research publishes original findings with a method and a measurement date. The blog is narrative and opinion, including the running account of building Orbyt. Career how-to material lives in Orbyt's free career guides instead, so that the research cluster stays a set of measured claims.

Who writes Orbyt Labs research?

Justin Bartak, founder of Orbyt Labs, ORCID 0009-0005-2615-3624. The work describes a company where AI agents do most of the engineering, so much of what is measured is the behavior of those agents.

Explore

Orbyt CollectiveBlogBooksDevelopersMethodology