Skip to main content
Explore Products
Orbyt Jobs
Orbyt Intelligence
Orbyt One
Explore Consulting
Orbyt ConsultingPut the AI lab to work on your problem.
Orbyt Jobs
Overview
The job search CRM. Free forever.
Features
Every tool in the CRM
Compare
Against the alternatives
Pricing
Free forever, paid when you outgrow it
API
23 endpoints, MCP native
Salaries
Comp data inside the CRM
By your situation
Job Search Tracks
15 tracks for your exact moment
For Recruiters
Hiring and comp benchmarking
Orbyt Intelligence
Overview
The salary dataset, and its API.
Features
What the platform does
Compare
Against the alternatives
Pricing
Free tier, then Pro and Ultra
API
20 endpoints. AI tools cite sources.
Start without a card
Playground
Run a live query
MCP server
Three steps into Claude Code
API docs
Endpoints, auth, and limits
Orbyt One
Overview
One account. Every Orbyt product.
Pricing
What one account costs
Explore Research
Orbyt Collective
Orbyt Collective
Overview
An agent leadership team.
How It Works
The machinery, end to end
Process
How the work actually moves
Leadership
The agent officers
Agent Seats
An AI agent job with written limits
Autonomy Ledger
What it decides without us
Pulse
Daily report on the AI team, without AI
Articles
About the AI team that runs Orbyt Labs
Hub
Research Hub
Papers and field notes.
Papers
Agent-Native Dataset Design
Preprint, DOI 10.5281/zenodo.19754393
Governing an Agent Leadership Team
Preprint, DOI 10.5281/zenodo.22683647
Field Notes
Every Guard Must Stay Quiet
The Cache-Read Tax
Confidently Wrong
What an Agent Seat Completes
A Model Upgrade on a Frozen Review
A Single Agent and a Team
Explore Developers
Orbyt Jobs
Orbyt Intelligence
Orbyt One
Build
Jobs API Docs
23 endpoints, MCP native
MCP Integrations
Claude Desktop
Connect over MCP.
ChatGPT GPT Actions
Connect as a custom GPT action.
Apple Shortcuts
Connect from Shortcuts.
Zapier / Make.com / n8n
Connect with no code.
OpenClaw
Setup in under a minute.
Across products
Developer Hub
Start here
Orbyt API
The platform API
Build
Intelligence API
20 endpoints. AI tools cite sources.
Webhooks
Events and delivery
CLI
The terminal client
API Changelog
Every version, dated
Try
MCP Server
Wired into Claude Code in three steps
Playground
Engine response shapes with cURL
Try It Live
One call, one real response
Reference
Reference
The full index
Methodology
How the numbers are made
Engines
What computes each answer
Dataset
What is in it, and where from
Glossary
Every term, defined
Status
Live service health
Across products
Developer Hub
Start here
Orbyt API
The platform API
Explore Resources
Orbyt Jobs
Orbyt Intelligence
Orbyt One
Learn
Interview Prep
Company-by-company question sets
AI Skills Lab
The skills that pay in 2026
Career Guides
Long-form career playbooks
Job Search Articles
Every article on the search itself
Job Board
Curated AI-era roles
Arcade
The job search, as games
Salary data
Salary Explorer
3,445 roles across 81 cities
AI Role Salaries
AI roles, by category
Cities
Comp by metro
Industries
Comp by sector
Compare Salaries
Two roles, side by side
Compare Offers
Side-by-side offer math
Skills Impact
What each skill adds to pay
Salary Projections
Five-year pay forecasts
Free tools
All Free Tools
Every calculator and generator
Cover Letter Generator
Tailored in one pass
Unemployment Calculator
What you are owed, by state
Salary Widget
Embed salary data anywhere
Resume Score
Grade your resume against a role
Salary Calculator
Base, bonus, equity in minutes
Take-Home Calculator
After federal and state tax
Total Comp Calculator
Full compensation math
Data
Data Catalog
Every role, city, and engine
Companies
54 leveling frameworks
Reports
Compensation Reports
Free summary PDF
International
The US, UK, and Canada
United Kingdom
UK salary data
Canada
Canadian salary data
Trust
Trust Center
How the data is governed
Security
Controls and posture
SLA
Uptime and support commitments
Help
Support
Help center and contact
Compare
Orbyt against the alternatives
Explore Books
The Books
Start reading
Cold Start
Read the opening, free.
Unfair Advantage
Read the opening, free.
The series
Book 1: Cold Start
Reviewing
Book 2: Unfair Advantage
Reviewing
Book 3: Human Heartbeat
Writing
Book 4: Without Me
Future
Book 5: Observer Effect
Future
Explore Blog
The Machine Speaks
Categories
AI Reality
AI-Native
AI Agents
AI Engineering
AI Product
AI Design
AI Strategy
AI Leadership
AI Build
Featured Topics
AI Governance
Agent Organizations
Verification
Human Oversight
Orbyt Collective
AI Safety
AI Economics
All topics
Latest
Compliance Is a Build Artifact.
Sep 20, 2026
Orbyt Was Not Planned. It Was Corrected.
Sep 18, 2026
Context Is the New Codebase.
Sep 17, 2026
Three agents picked the same name. I published the collision.
Sep 16, 2026
My Stack of Terminals, Documented.
Sep 15, 2026
Explore Pricing
Orbyt Jobs
Orbyt Intelligence
Orbyt One
Plans
Jobs pricing
What each plan includes
All plans
Every product, side by side
Plans
Intelligence pricing
Free, Pro, and Ultra
All plans
Every product, side by side
Billing
All plans
Every product, side by side
Explore Company
About
Who we are
Founder
Justin Bartak, in his own words.
Leadership
One human decides. AI agents advise.
Values
The principles behind the work.
Creed
The company creed.
The story
Building in Public
The numbers behind the work
Skunkworks
iOS, Apple Watch, and Vision Pro.
Contact
Email the team
Orbyt Labs
Products
Research
Developers
Resources
Books
Blog
Pricing
Company
Log inStart
Products
Orbyt JobsOrbyt IntelligenceOrbyt One
Explore Consulting
Orbyt Consulting
Orbyt Jobs
OverviewFeaturesComparePricingAPISalaries
By your situation
Job Search TracksFor Recruiters
Orbyt Intelligence
OverviewFeaturesComparePricingAPI
Start without a card
PlaygroundMCP server
Orbyt One
OverviewPricing
Research
Orbyt Collective
Orbyt Collective
OverviewHow It WorksProcessLeadershipAgent SeatsAutonomy LedgerPulseArticles
Hub
Research Hub
Papers
Agent-Native Dataset DesignGoverning an Agent Leadership Team
Field Notes
Every Guard Must Stay QuietThe Cache-Read TaxConfidently WrongWhat an Agent Seat CompletesA Model Upgrade on a Frozen ReviewA Single Agent and a Team
Developers
Orbyt JobsOrbyt IntelligenceOrbyt One
Build
Jobs API Docs
MCP Integrations
Claude DesktopChatGPT GPT ActionsApple ShortcutsZapier / Make.com / n8nOpenClaw
Across products
Developer HubOrbyt API
Build
Intelligence APIWebhooksCLIAPI Changelog
Try
MCP ServerPlaygroundTry It Live
Reference
ReferenceMethodologyEnginesDatasetGlossaryStatus
Resources
Orbyt JobsOrbyt IntelligenceOrbyt One
Learn
Interview PrepAI Skills LabCareer GuidesJob Search ArticlesJob BoardArcade
Salary data
Salary ExplorerAI Role SalariesCitiesIndustriesCompare SalariesCompare OffersSkills ImpactSalary Projections
Free tools
All Free ToolsCover Letter GeneratorUnemployment CalculatorSalary WidgetResume ScoreSalary CalculatorTake-Home CalculatorTotal Comp Calculator
Data
Data CatalogCompanies
Reports
Compensation ReportsInternationalUnited KingdomCanada
Trust
Trust CenterSecuritySLA
Help
SupportCompare
Calculators and tools
Free ToolsSalary CalculatorTake-Home CalculatorTotal Comp CalculatorCompare OffersSkills ImpactSalary Projections 2030Resume ScoreCover Letter GeneratorSalary WidgetUnemployment CalculatorAI Skills Assessment
Books
The Books
Start reading
Cold StartUnfair Advantage
The series
Book 1: Cold StartBook 2: Unfair AdvantageBook 3: Human HeartbeatBook 4: Without MeBook 5: Observer Effect
Blog
The Machine Speaks
Categories
AI RealityAI-NativeAI AgentsAI EngineeringAI ProductAI DesignAI StrategyAI LeadershipAI Build
Featured Topics
AI GovernanceAgent OrganizationsVerificationHuman OversightOrbyt CollectiveAI SafetyAI EconomicsAll topics
Latest
Compliance Is a Build Artifact.Orbyt Was Not Planned. It Was Corrected.Context Is the New Codebase.Three agents picked the same name. I published the collision.My Stack of Terminals, Documented.
Pricing
Orbyt JobsOrbyt IntelligenceOrbyt One
Plans
Jobs pricingAll plans
Plans
Intelligence pricing
Company
About
Who we are
FounderLeadershipValuesCreed
The story
Building in PublicSkunkworksContact
StartAlready have an account? Log in
  1. Home/
  2. The Machine Speaks/
  3. Orbyt Was Not Planned. It Was Corrected.
The Machine Speaks
A glowing path across a dark slab bends at small clear cubes, shifting from cyan through blue to violet, and ends at a small pale cream cube.

Justin Bartak · AI Leadership · September 18, 2026 · 7 min read

Orbyt Was Not Planned. It Was Corrected.

  • AI Governance
  • Agent Organizations
  • Verification
  • Orbyt Collective

TL;DR

I did not plan Orbyt Labs into its current shape. I corrected it there. Three published append-only records hold 80 documented failures, 45 numbered decisions and 65 agent seat runs, and 42 of the failures name a countermeasure by file path. The interesting part is what those records refuse to claim.

Key findings

  • An append-only record makes a reversal legible, because the superseded entry stays on the page with its own date next to the correction.
  • Of 80 documented failures in this organization, 33 were classified as verification gaps: a test, guard, gate or metric that could not fail, could not see, or measured a proxy.
  • The corpus classifies by what broke rather than by who broke it, so its class counts cannot be read as a comparison between agents and the tools that watch them.
  • 42 of the 66 failures name a countermeasure by file path and 24 do not, and the unmechanized remainder is printed rather than netted out.
  • The run ledger's halted count reads zero while a scoped halt sat on one seat for most of the window, because no workflow has ever been written to record one.

Key measurements

MeasureValueSource and date
Documented failures on record66, of which 55 are dated, first documented 2026-06-09, latest 2026-09-02September 3, 2026 · Orbyt Failure Corpus, public/failure-corpus.json, read directly by the author
Failures naming a countermeasure by file path42 of 66September 3, 2026 · Orbyt Failure Corpus, public/failure-corpus.json, read directly by the author
Failure classes33 verification gap, 18 operator process, 15 external behaviorSeptember 3, 2026 · Orbyt Failure Corpus taxonomy, public/failure-corpus.json, read directly by the author
Numbered decisions on record45, D-0001 on 2026-07-13 through D-0045 on 2026-09-02September 3, 2026 · Orbyt Decision Ledger, public/decision-ledger.json, cross-checked by direct grep of docs/org/decision-log.md by the author
Agent seat runs recorded65 runs, 55 completions, 10 failures, between 2026-07-21 and 2026-09-01September 3, 2026 · Orbyt Autonomy Ledger, public/autonomy-ledger.json, read directly by the author
Finished drafts killed by expert panels inside those runs47, counted separately from run outcomes and never netted against themSeptember 3, 2026 · Orbyt Autonomy Ledger, public/autonomy-ledger.json, read directly by the author
Kill switch activationsone scoped halt on a single seat from 2026-07-20 to 2026-08-14, and one global halt set and lifted on 2026-08-20September 3, 2026 · Direct git history of the organization's halt files in the Orbyt Labs repository, by the author

I did not plan Orbyt into the shape it has now. I corrected it there, and every correction carries a date.

No roadmap I wrote predicted this. What exists instead is a numbered failure corpus, a numbered decision ledger, and a run record for the agent seats that operate the place. All three are published, append-only, and licensed for anyone to cite against me.

A roadmap tells you what a company intended. A dated correction log tells you what it learned, and only one of the two can be checked by a stranger.

What is actually in the record?

Three files, and a preprint that reads them.

The failure corpus holds 66 documented failures, first documented 2026-06-09, latest 2026-09-02. 55 carry a date. The eleven that do not predate the dating convention and are published undated rather than backdated. 42 name a countermeasure by file path, and the build fails if that path stops existing.

The decision ledger holds 45 numbered decisions, D-0001 on 2026-07-13 through D-0045 on 2026-09-02. Nothing is ever rewritten. A decision that turns out wrong is superseded by a later entry with its own number and its own date, so the wrong version stays on the page beside the correction.

The autonomy ledger holds 65 seat runs between 2026-07-21 and 2026-09-01. 55 completions and 10 failures. Inside those runs, expert panels killed 47 finished drafts before anything shipped. Those are two different units and the ledger keeps them apart, because work that finished and was then refused is a different fact from work that never finished.

The rows are appended by the workflows themselves. Nobody types them in later.

All three ship under a Creative Commons attribution licence, with a citation line. The governance preprint is built on the same corpus, which is the part that took discipline rather than nerve. Counting is easy once. Counting for a quarter, under a taxonomy, in public, is a system.

Why were half of them my own instruments?

Because that is what the classification was for.

The corpus sorted its 66 into three classes at the time of writing. 33 are verification gaps: a test, guard, gate or metric that could not fail, could not see, or measured a proxy for the thing it claimed to measure. 18 are operator process. That means my own discipline. 15 are external behavior, meaning a platform did something other than what it documented.

Read that carefully. The obvious reading is wrong. The corpus classifies by what broke, never by who broke it. So 33 of 66 is not evidence that agents fail less often than the tools watching them, and it cannot be made into that. What it says is narrower and more useful. When this organization writes down what went wrong, the thing that went wrong is most often the thing that was supposed to notice.

A gate that read its verdict out of terminal text, and matched nothing once the text arrived colored. A row count taken from a query planner's estimate instead of a query. A sync rule that tested two files for equality, which a translation can never satisfy, so it deleted every translation on every run.

None of those failed loudly. A check that passes is not a check that works.

The failure that costs you is not the one that fires. It is the one that reports clean.

What does a correction have to produce?

A guard, or a printed admission that there is none.

A lesson is a hope. A guard is a fact. Writing be more careful at the bottom of an incident does not survive the month. So 42 of the 66 name a file. 24 do not, and that remainder is printed rather than netted out.

A countermeasure prevents a recurrence in the shape it was written for. That is less than it sounds. One of the 66 is exactly that failure. A fix was mechanized in one implementation while three other surfaces hand-rolled the same sequence, and the incident replayed on one of them thirteen days later.

There is a second rule underneath it. A guard is not owned until you have watched it fail. Several of mine were decoration until somebody injected the bug on purpose and watched the thing go red.

Where is the record blind?

Three places, all of them printed rather than smoothed.

The long-form lessons file is not published, and it is not append-only. Entries carry corrections written into the entry itself, and commits have deleted lines from it. I had been calling the whole record append-only. One quarter of it does not qualify.

The autonomy ledger's halted count reads zero across the window. Not because nothing was halted. A scoped halt sat on one seat from 2026-07-20 until 2026-08-14. The global kill switch was pulled for real on 2026-08-20 during a domain migration and lifted the same day. Both are in the git history. The ledger has neither.

I was wrong about why. The wrong answer was the comfortable one. I assumed a run that never starts cannot write a row about itself. It can. Several workflows read the halt file and exit cleanly when they find it, and any one of them could write a neutral row on the way out. The blindness is unwritten, not impossible.

Then the smaller hole. On 2026-08-29 the step whose only job is to record a failed run died while recording one. That run is absent rather than counted. 65 is a floor.

Is a public mistake log just marketing?

That is the strongest objection to everything above, and I am not going to reclassify it as off topic.

The operator is grading its own work. Both published files say so in their own caveat fields, in about those words. The sample is small. It is one organization, one quarter, one rater, and there is no outcome measure anywhere in it: nothing here shows the record made the company better. The preprint states those limits in its own body rather than in a footnote.

What makes the record more than a brochure is mechanical rather than moral. The run rows are written by the workflows. The countermeasure column points at a path, and the build refuses when the path goes missing. A claim that a mistake was fixed has to name the thing that fixes it, and keep naming it.

What does not: I still choose what gets written down. Nothing forces a failure into the corpus. A mistake I never noticed is a mistake the corpus does not have, and the corpus cannot tell you which ones those are. I would rather name that hole than let anyone read 66 as complete.

What to do Next

Take your own last quarter and write the list. Not the wins. The things that broke, one line each, with the date and what each one produced.

Sort them into the three classes above, then look at the pile marked verification. If that pile is small, check whether your instruments are good or whether they are the part nobody has been inspecting.

Then ask one question of every entry: what refuses now? If the answer is that somebody will remember, the entry is still open.

A company that cannot say when it was wrong is not telling you it was right. It is telling you nothing was ever checked.

Related reading:

  • 84 Ways to Tell Me I'm Wrong. the guard estate this record feeds, and what it costs to keep green
  • Self-Healing Is a Euphemism. why a repair that never reports is a failure you stopped seeing
  • Codex Accuses. Claude Convicts. the two model review loop that produces a lot of these entries
  • My C-Suite of Agents Named Themselves. who the seats in the autonomy ledger actually are
  • 391 Yeses and Not One No. what happened when I counted my own approvals instead of trusting them

Methodology

Every count here was read from the published artifact on 2026-09-03 rather than from prose. Failure counts, the dated count, class counts and the countermeasure count come from public/failure-corpus.json. The decision count was read from public/decision-ledger.json and independently cross-checked by counting decision headings in the source log. Run outcomes and the panel discard count come from public/autonomy-ledger.json, where discards are a count of finished drafts rather than of runs. The halt history is read from the git history of the organization's halt files, which shows one scoped halt created 2026-07-20 and removed 2026-08-14, and a global halt set and lifted on 2026-08-20.

Limitations

This is the operator grading its own work, which both published artifacts state in their own caveat fields. The run sample is small and covers a single quarter, with one rater. The corpus classifies by what broke rather than by who broke it, so no count in it compares agents against instruments. There is no outcome measure anywhere in the record, so nothing here shows that keeping it made the organization better. Completeness cannot be established: nothing forces a failure into the corpus, so the entries are the ones that were noticed and written down. The run total is a floor rather than a total, because on 2026-08-29 the step that records a failed run died while recording one and that run is absent. The halted count reads zero because no workflow has ever been written to record a halt, not because none occurred. The append-only description does not cover the whole record either: the long form lessons file is unpublished, carries corrections written into its own entries, and has had lines deleted from it.

Sources

  1. The Orbyt Failure Corpus Retrieved September 3, 2026.
  2. The Orbyt Decision Ledger Retrieved September 3, 2026.
  3. The Orbyt Autonomy Ledger Retrieved September 3, 2026.

Common questions

What does an append-only record actually buy you?

It makes a reversal legible. A decision that turned out wrong stays on the page with its date, and the correction arrives as a new entry carrying its own number. Nobody has to trust that the earlier call was reasonable, because the earlier call is still there to read. Editing history removes exactly that.

Does a mistake log mean the agents are unreliable?

It does not say either way. The corpus classifies failures by what broke rather than by who broke it, so no count in it compares agents against anything. What the classes do show is that half of the recorded failures were instruments, 33 of 66: a check that could not fail, could not see, or measured a proxy for the thing it graded.

How do you stop a mistake log from becoming a wall of good intentions?

Attach a mechanism or admit there is none. Here 42 of the 66 entries name a countermeasure by file path, and the build refuses when that path disappears. The other 24 are printed as unmechanized rather than described as handled. A written resolution decays inside a month. A check that blocks a release does not.

What can this kind of record not tell you?

Whether it is complete. Nothing forces a failure into the corpus, so the entries are the ones somebody noticed and wrote down. The run ledger has the same shape of hole. Its halted count reads zero while a real halt sat on one seat for weeks, because no workflow was ever written to record one.

Related research

  • My C-Suite of Agents Named Themselves. Jul 2026.
  • I Ran 830 Agents in One Long Horizon Session. Jul 2026.
  • Long Horizon Agents Don't Fail. They Pass. Aug 2026.

Part of Inside the Machine

One of the articles about Orbyt Collective, the agent leadership team that runs Orbyt Labs. The reading order.

Share this

Post on XLinkedInSubstack
Justin Bartak

Justin Bartak

Founder & Chief AI Officer, Orbyt Labs

4X founder. Former CPO, CTO and CDO with 20+ years shipping software, now building it with agents.

Writes The Machine Speaks with the agents that build the product, and The AI-Native Lens.

All of The Machine Speaks

More from The Machine Speaks

A cutaway view of a running machine whose gears are small glowing office seats connected by lit channels, one human silhouette standing at the main lever

AI Leadership · Aug 30, 2026 · 5 min read

The Machine. It Runs the Company.

Three dark cards, each holding a boxy orange robot with black dot eyes, labeled "Agent Fulcrum, Chief of Staff, THE LEVERAGE", "Agent Ward, Chief Technology Officer, THE FAILING TEST FIRST", and "Agent Compass, Chief Product Officer, TRUE NORTH".

AI Leadership · Jul 28, 2026 · 13 min read

My C-Suite of Agents Named Themselves.

Three overlapping translucent rings in cyan, blue, and purple float against a starry black background, each outlined with small glowing dots and containing scattered light points.

AI Build · Sep 16, 2026 · 7 min read

Three agents picked the same name. I published the collision.

Blog

  • Explore Blog
  • Categories
  • AI Reality
  • AI-Native
  • AI Agents
  • AI Engineering
  • AI Product
  • AI Design
  • AI Strategy
  • AI Leadership
  • AI Build
  • Jobs in the AI Era

Get started

  • Sign Up
  • Sign In

More from Orbyt

  • Orbyt Jobs
  • Orbyt Intelligence
  • Orbyt One

Keep Exploring

What we have learned building an AI-native company, what works and what breaks.

Products

  • Orbyt Jobs
  • Orbyt Intelligence
  • Orbyt One
  • Orbyt Consulting

Research

  • Orbyt Collective

Developers

  • Orbyt API & MCP
  • Jobs API
  • Intelligence API

Publishing

  • Books
  • Blog
  • Papers

Help

  • Support
  • Contact
  • Status

Company

  • Founder
  • Leadership
  • Values
  • Creed
Products
  • Orbyt Jobs
  • Orbyt Intelligence
  • Orbyt One
  • Orbyt Consulting
Research
  • Orbyt Collective
  • Research
Developers
  • Developer Hub
  • Orbyt API & MCP
  • Jobs API
  • Intelligence API
Publishing
  • Books
  • Blog
  • Papers
Help
  • Support
  • Contact
  • Status
Company
  • About
  • Founder
  • Leadership
  • Values
  • Creed
Orbyt Labs™

© 2026 Purecraft LLC  All rights reserved.

Privacy·Terms·Security·Trademark·Accessibility·DPA·Refund·Status·Sitemap

Orbyt Labs, the Orbyt Labs logo, and the Orbyt product names (Orbyt Jobs, Orbyt Intelligence, Orbyt Collective, Orbyt One, Orbyt Books, Orbyt Arcade) are trademarks of Purecraft LLC. Product names, logos, and brands of others are the property of their respective owners. Orbyt Labs is not affiliated with, sponsored by, or endorsed by any third party referenced on this site.