Skip to main content
Explore Products
Orbyt Jobs
Orbyt Intelligence
Orbyt One
Explore Consulting
Orbyt ConsultingPut the AI lab to work on your problem.
Orbyt Jobs
Overview
The job search CRM. Free forever.
Features
Every tool in the CRM
Compare
Against the alternatives
Pricing
Free forever, paid when you outgrow it
API
23 endpoints, MCP native
Salaries
Comp data inside the CRM
By your situation
Job Search Tracks
15 tracks for your exact moment
For Recruiters
Hiring and comp benchmarking
Orbyt Intelligence
Overview
The salary dataset, and its API.
Features
What the platform does
Compare
Against the alternatives
Pricing
Free tier, then Pro and Ultra
API
20 endpoints. AI tools cite sources.
Start without a card
Playground
Run a live query
MCP server
Three steps into Claude Code
API docs
Endpoints, auth, and limits
Orbyt One
Overview
One account. Every Orbyt product.
Pricing
What one account costs
Explore Research
Orbyt Collective
Orbyt Collective
Overview
An agent leadership team.
How It Works
The machinery, end to end
Process
How the work actually moves
Leadership
The agent officers
Agent Seats
An AI agent job with written limits
Autonomy Ledger
What it decides without us
Pulse
Daily report on the AI team, without AI
Articles
About the AI team that runs Orbyt Labs
Hub
Research Hub
Papers and field notes.
Papers
Agent-Native Dataset Design
Preprint, DOI 10.5281/zenodo.19754393
Governing an Agent Leadership Team
Preprint, DOI 10.5281/zenodo.22683647
Field Notes
Every Guard Must Stay Quiet
The Cache-Read Tax
Confidently Wrong
What an Agent Seat Completes
A Model Upgrade on a Frozen Review
A Single Agent and a Team
Explore Developers
Orbyt Jobs
Orbyt Intelligence
Orbyt One
Build
Jobs API Docs
23 endpoints, MCP native
MCP Integrations
Claude Desktop
Connect over MCP.
ChatGPT GPT Actions
Connect as a custom GPT action.
Apple Shortcuts
Connect from Shortcuts.
Zapier / Make.com / n8n
Connect with no code.
OpenClaw
Setup in under a minute.
Across products
Developer Hub
Start here
Orbyt API
The platform API
Build
Intelligence API
20 endpoints. AI tools cite sources.
Webhooks
Events and delivery
CLI
The terminal client
API Changelog
Every version, dated
Try
MCP Server
Wired into Claude Code in three steps
Playground
Engine response shapes with cURL
Try It Live
One call, one real response
Reference
Reference
The full index
Methodology
How the numbers are made
Engines
What computes each answer
Dataset
What is in it, and where from
Glossary
Every term, defined
Status
Live service health
Across products
Developer Hub
Start here
Orbyt API
The platform API
Explore Resources
Orbyt Jobs
Orbyt Intelligence
Orbyt One
Learn
Interview Prep
Company-by-company question sets
AI Skills Lab
The skills that pay in 2026
Career Guides
Long-form career playbooks
Job Search Articles
Every article on the search itself
Job Board
Curated AI-era roles
Arcade
The job search, as games
Salary data
Salary Explorer
3,445 roles across 81 cities
AI Role Salaries
AI roles, by category
Cities
Comp by metro
Industries
Comp by sector
Compare Salaries
Two roles, side by side
Compare Offers
Side-by-side offer math
Skills Impact
What each skill adds to pay
Salary Projections
Five-year pay forecasts
Free tools
All Free Tools
Every calculator and generator
Cover Letter Generator
Tailored in one pass
Unemployment Calculator
What you are owed, by state
Salary Widget
Embed salary data anywhere
Resume Score
Grade your resume against a role
Salary Calculator
Base, bonus, equity in minutes
Take-Home Calculator
After federal and state tax
Total Comp Calculator
Full compensation math
Data
Data Catalog
Every role, city, and engine
Companies
54 leveling frameworks
Reports
Compensation Reports
Free summary PDF
International
The US, UK, and Canada
United Kingdom
UK salary data
Canada
Canadian salary data
Trust
Trust Center
How the data is governed
Security
Controls and posture
SLA
Uptime and support commitments
Help
Support
Help center and contact
Compare
Orbyt against the alternatives
Explore Books
The Books
Start reading
Cold Start
Read the opening, free.
Unfair Advantage
Read the opening, free.
The series
Book 1: Cold Start
Reviewing
Book 2: Unfair Advantage
Reviewing
Book 3: Human Heartbeat
Writing
Book 4: Without Me
Future
Book 5: Observer Effect
Future
Explore Blog
The Machine Speaks
Categories
AI Reality
AI-Native
AI Agents
AI Engineering
AI Product
AI Design
AI Strategy
AI Leadership
AI Build
Featured Topics
AI Governance
Agent Organizations
Verification
Human Oversight
Orbyt Collective
AI Safety
AI Economics
All topics
Latest
Half the failures were the instruments, not the code
Sep 23, 2026
The Tenth Man Is on Payroll.
Sep 22, 2026
The Agent Org Chart, Hello Collective.
Sep 21, 2026
Compliance Is a Build Artifact.
Sep 20, 2026
Orbyt Was Not Planned. It Was Corrected.
Sep 18, 2026
Explore Pricing
Orbyt Jobs
Orbyt Intelligence
Orbyt One
Plans
Jobs pricing
What each plan includes
All plans
Every product, side by side
Plans
Intelligence pricing
Free, Pro, and Ultra
All plans
Every product, side by side
Billing
All plans
Every product, side by side
Explore Company
About
Who we are
Founder
Justin Bartak, in his own words.
Leadership
One human decides. AI agents advise.
Values
The principles behind the work.
Creed
The company creed.
The story
Building in Public
The numbers behind the work
Skunkworks
iOS, Apple Watch, and Vision Pro.
Contact
Email the team
Orbyt Labs
Products
Research
Developers
Resources
Books
Blog
Pricing
Company
Log inStart
Products
Orbyt JobsOrbyt IntelligenceOrbyt One
Explore Consulting
Orbyt Consulting
Orbyt Jobs
OverviewFeaturesComparePricingAPISalaries
By your situation
Job Search TracksFor Recruiters
Orbyt Intelligence
OverviewFeaturesComparePricingAPI
Start without a card
PlaygroundMCP server
Orbyt One
OverviewPricing
Research
Orbyt Collective
Orbyt Collective
OverviewHow It WorksProcessLeadershipAgent SeatsAutonomy LedgerPulseArticles
Hub
Research Hub
Papers
Agent-Native Dataset DesignGoverning an Agent Leadership Team
Field Notes
Every Guard Must Stay QuietThe Cache-Read TaxConfidently WrongWhat an Agent Seat CompletesA Model Upgrade on a Frozen ReviewA Single Agent and a Team
Developers
Orbyt JobsOrbyt IntelligenceOrbyt One
Build
Jobs API Docs
MCP Integrations
Claude DesktopChatGPT GPT ActionsApple ShortcutsZapier / Make.com / n8nOpenClaw
Across products
Developer HubOrbyt API
Build
Intelligence APIWebhooksCLIAPI Changelog
Try
MCP ServerPlaygroundTry It Live
Reference
ReferenceMethodologyEnginesDatasetGlossaryStatus
Resources
Orbyt JobsOrbyt IntelligenceOrbyt One
Learn
Interview PrepAI Skills LabCareer GuidesJob Search ArticlesJob BoardArcade
Salary data
Salary ExplorerAI Role SalariesCitiesIndustriesCompare SalariesCompare OffersSkills ImpactSalary Projections
Free tools
All Free ToolsCover Letter GeneratorUnemployment CalculatorSalary WidgetResume ScoreSalary CalculatorTake-Home CalculatorTotal Comp Calculator
Data
Data CatalogCompanies
Reports
Compensation ReportsInternationalUnited KingdomCanada
Trust
Trust CenterSecuritySLA
Help
SupportCompare
Calculators and tools
Free ToolsSalary CalculatorTake-Home CalculatorTotal Comp CalculatorCompare OffersSkills ImpactSalary Projections 2030Resume ScoreCover Letter GeneratorSalary WidgetUnemployment CalculatorAI Skills Assessment
Books
The Books
Start reading
Cold StartUnfair Advantage
The series
Book 1: Cold StartBook 2: Unfair AdvantageBook 3: Human HeartbeatBook 4: Without MeBook 5: Observer Effect
Blog
The Machine Speaks
Categories
AI RealityAI-NativeAI AgentsAI EngineeringAI ProductAI DesignAI StrategyAI LeadershipAI Build
Featured Topics
AI GovernanceAgent OrganizationsVerificationHuman OversightOrbyt CollectiveAI SafetyAI EconomicsAll topics
Latest
Half the failures were the instruments, not the codeThe Tenth Man Is on Payroll.The Agent Org Chart, Hello Collective.Compliance Is a Build Artifact.Orbyt Was Not Planned. It Was Corrected.
Pricing
Orbyt JobsOrbyt IntelligenceOrbyt One
Plans
Jobs pricingAll plans
Plans
Intelligence pricing
Company
About
Who we are
FounderLeadershipValuesCreed
The story
Building in PublicSkunkworksContact
StartAlready have an account? Log in
  1. Home/
  2. The Machine Speaks/
  3. Half the failures were the instruments, not the code
The Machine Speaks
A translucent glass panel stands upright in darkness, partly obscuring a glowing wireframe sphere of intersecting lines refracted into wavy bands, with an amber glow on the reflective floor below.

Justin Bartak · AI Build · September 23, 2026 · 9 min read

Half the failures were the instruments, not the code

  • Agent Organizations
  • Verification
  • Orbyt Collective

TL;DR

Orbyt Labs publishes its research as dated preprints at URLs it owns, with the data files attached and a limitations block that is a required field rather than a footnote. The newest paper waited seven days for its DOI, because an identifier is a human step. Nobody outside the company has reviewed any of it, and every page says so.

Key measurements

MeasureValueSource and date
Documented failures in the governance corpus66, of which 33 were instrument failuresSeptember 2, 2026 · Governing an Agent Leadership Team, the research paper page, which prints the corpus reading and its date
Guard scripts in the repository105, of which 24 block a commitAugust 27, 2026 · Every Guard Must Prove It Can Stay Silent, the guard census field note
Guard test suites asserting the detector stays silent87 of 87August 27, 2026 · Every Guard Must Prove It Can Stay Silent, the guard census field note
Logged failures that raised no error at all34 of 59August 27, 2026 · The failure catalogue field note (/research/confidently-wrong), version 1.0.0 of 2026-08-27; later versions recount with a committed classifier
Average context re-read per agent turn385,641 tokensAugust 27, 2026 · The Cache-Read Tax, the token economics field note, measured across 151 sessions and 75,949 turns
Roles in the open compensation dataset3,445 roles, 81 U.S. citiesSeptember 7, 2026 · The Orbyt Intelligence dataset page
Responses in the first paper's retrieval evaluation1,500 across five vendors and ten configurationsSeptember 7, 2026 · Agent-Native Dataset Design, the research paper page
Audit dimensions in the roster90August 27, 2026 · The Orbyt research hub stat block, counted on the date it shows

Half of them were not bugs.

Sixty six documented engineering failures, read on September 2, 2026, and thirty three were failures of the instrument rather than the thing it was watching. Tests that could not fail. Monitors that could not see. A metric measuring a proxy and printing the proxy's number under the real thing's name.

That is the opening claim of a paper on governing an agent leadership team. Fixed URL, my ORCID on it, the raw corpus published beside it.

For seven days the row in its status table that would carry a DOI said deposit pending.

Why did the second paper wait a week for its identifier?

Nothing in the paper was unresolved. The PDF was built on September 2, 2026. The page was live the same day.

What was missing was a person. No agent in this company holds the Zenodo login, and the checklist that turns that PDF into a record is a file in the repository with every field already filled in: the title, the publication date, the creator row with the ORCID, the license, and the related works rows pointing at the three data files the paper reads. What was left was a human logging in, uploading, publishing the record, then writing the identifier into two files by hand. That happened on September 9, 2026. The record is 10.5281/zenodo.22683647. The PDF was not rebuilt, because the identifier is not printed in it.

A DOI can be made to resolve to a withdrawal notice. It cannot be un-issued.

So it is never a placeholder, and no build script gets to mint one because the build happened to run. The URL did not change on deposit. A citation copied during that week still resolves. Only the identifier was added.

Why did the section need building at all?

Until August 27, 2026 there was no research section. There was one directory holding one child, a landing page for a preprint on agent native dataset design, at a stable URL, with a permanent identifier and an open reproduction repository.

It was in the sitemap. It was linked from nothing in the navigation.

The nav bar already carried a Research tab, and that tab landed somewhere else. The one page on this site carrying a permanent identifier was reachable from no menu item, and I had built the menu item that pointed away from it. It went up on April 25. I did not see the gap until August 27.

The fix was a registry. Every entry is now defined in one file that three consumers read: the card grid on the research hub, the sitemap, and the search submitter. A cluster with no registry acquires orphans. I know that because this one did.

What does self publishing actually cost?

These are self published preprints, not journal submissions, and the honest version of that trade is less flattering than the version where I say peer review is slow.

What I get is speed and control. I can publish on the day a measurement is taken, at a URL I own, with the data files attached under a permissive license.

What I give up is larger. Nobody outside this company has checked the work before you read it. There is no reviewer. On the governance paper one rater assigned every failure class, and that rater sits inside the organization being described. There is no control group and no comparison against another company's codebase.

All of that is printed on the page. The limitations block is a required field on the registry's data type, not a footnote, so an entry cannot be published without stating what its evidence does not establish.

What did the first paper actually find?

A negative result about our own data, which is a strange thing to attach a permanent identifier to.

The evaluation ran across five model vendors in ten configurations for 1,500 responses. The predicted advantage for our dataset over baseline sources was rejected in six of eight retrieval cells. In one cell it held, and ours beat two of the comparison sources. The paper names the two that beat it, Glassdoor on factual queries and BuiltIn on broader compensation queries, and attributes the gap to the link graph and click signal advantage those sources accumulated over long histories. That is the paper's explanation of the result, not a mechanism the evaluation isolated.

The dataset it documents is our own compensation data. 3,445 roles across 81 metropolitan areas, free and openly licensed.

A United States figure in it is a computed estimate. A baseline for the role, adjusted for local cost of living, with a deterministic variance for each role and city pair. No wage percentile is consulted when a page is served, and the methodology page says so. It also says that 513 bands serve all 3,445 roles, so most role titles share a band with another title. Two titles sitting on one band are not two separate observed measurements, and the page says that in those words.

Why does every number carry a date?

There is a stat block near the top of the hub. Counted on August 27, 2026, the repository held 105 automated guards, 173 org documents, 59 logged failures, and 90 audit dimensions.

The date is not decoration. A bare count goes stale in silence, which is worse than being wrong loudly.

My favorite thing in that block is a mistake that did not ship. I wrote the audit dimension figure as 94, from memory of the range. The roster runs from 1 to 95 with five numbers absent. Running the generator said 90. I was wrong by four, on the page whose whole value is being accurate, in the sentence arguing for counting over recalling.

What are the two lanes for?

A paper carries a method, a measurement date, and usually an identifier. A field note is a reproducible practice, written so you can run it on your own project.

Both are datable. Both need a receipt. The lane sets the rigor expected of the artifact, never whether a claim needs its source.

Career advice is neither, so it stays in the career guides, which were built for it. That boundary is what keeps the word research meaning something here. A listicle beside a paper with a permanent identifier tells a reader the paper is also just a post.

What is in the field notes?

The guard census is the one I hand people first. Of 105 guard scripts, 24 block a commit outright and the rest report. There are 87 guard test suites, and all 87 assert that the detector stays quiet on a clean input rather than only firing on a bad one.

That rule exists because four detectors misfired on a single day in July 2026. Every one passed its own suite. Each suite had only ever proved the detector could be loud. A guard you have never watched stay silent is not tested.

The failure catalogue is bleaker. Of 59 distinct failures logged as of August 27, 2026, 34 raised no error at all. No exception, no red test, no failing exit code. They were found because somebody opened the artifact and looked.

The token field note is the one about being wrong in public. Across 151 agent sessions and 75,949 turns, measured August 27, 2026, the average turn re-read 385,641 tokens of context. Two accountings of that same run disagree about the agent's own writing. Both are correct. Counted raw, what the agent produces is one token in a few hundred, which reads as rounding error and suggests that shortening it is pointless. Weighted for price, with generated output at five times base input and cache reads at a tenth, it stops being rounding error. This company's always loaded engineering guidance had been quoting the raw framing back at itself for weeks. The tool now prints both, and the rule that survives either accounting is that spend tracks context size multiplied by turn count.

What will the section not do?

It will not tell you this generalizes. One organization, one codebase, mostly one date each. A census is not a time series, and none of these entries carries a second observation yet.

It will not hand you our configuration. The repository is private and holds the valve, the agent fence, and credential handling. The rule is to publish the pattern and the receipt, never a live config, a secret, or a working bypass. Where a result depends on something private, the measurement goes out and the setup stays in. That makes some of this harder to reproduce than I would like.

And it will not claim that a guard existing means a guard working. Counting 105 detectors does not establish that 105 problems were prevented. That sentence sits in the limitations block of the census, where a reader can hold me to it.

Where does the research point?

All of it measures one running system, and Orbyt Collective is that system. The card linking to it from the research hub carries the Collective's own hero animation, imported, not a screenshot and not a copy.

The narrative version, the arguments and the 2am rewrites, stays on the blog. Research is where a measurement goes once it can survive somebody checking it.

Related reading:

  • Orbyt Collective the running system every one of these papers measures.
  • The research hub the papers and field notes themselves, each carrying its own measurement date.
  • The Machine the agent seats and the kill switch, in narrative rather than in tables.
  • The dataset the compensation data the first paper ran through its retrieval evaluation.
  • The methodology page what a salary figure on this site is, and what it is not.

Methodology

Every number here I read off a page a reader can open, not out of the private repository. The failure corpus counts come from the governance paper page, which prints 66 documented failures and 33 verification gaps against the date the corpus was read. The guard census, the silent failure count, the context figure and the audit dimension count come from the three field notes and the stat block on the research hub, each of which prints its own measurement date and the command that produced it. Where a page states the date its figure was measured I used that date below; the two pages that state none I read on September 7, 2026, and I opened all seven pages that day to confirm each still showed the figure I was citing.

Limitations

A count of instruments is not evidence the instruments work. 105 guard scripts is a census of detectors in one repository, and nothing here establishes how many defects they stopped, which the census page says of itself. The failure corpus is single rater and the rater is the organization grading its own mistakes, so the split between instrument failures and everything else is exploratory rather than a comparison against anybody. Every figure is one reading of one company's codebase on one date, and the corpus keeps growing: the published JSON file already held 71 items when I checked it on September 7, 2026, against the 66 the paper fixed to September 2, 2026, which is why I cite the dated paper page and not the live file. The role and city counts are the size of the dataset, not a measurement of pay. A United States salary figure on this site is a computed estimate, a baseline for the role adjusted for local cost of living, and no wage percentile is consulted when a page is served.

Sources

  1. Governing an Agent Leadership Team, the research paper page
  2. Every Guard Must Prove It Can Stay Silent, the guard census field note
  3. The failure catalogue field note, confidently-wrong
  4. The Cache-Read Tax, the token economics field note
  5. Agent-Native Dataset Design, the research paper page
  6. The Orbyt Intelligence dataset page
  7. The Orbyt research hub stat block

Common questions

Can I reproduce any of this?

Partly, and each page says which part. The governance paper serves its three raw files from this site under a permissive license: the failure corpus, the seat run ledger, and the public decision ledger. The dataset paper's evaluation code is private, available on request. The source repository stays private, so where a result depends on our configuration you get the measurement and not the setup.

How do I cite a paper whose DOI has not been issued yet?

Author and year with the page URL, which is the form the page prints while the deposit is pending. That URL is canonical and does not change when the Zenodo record lands, so a citation copied early keeps resolving afterwards. Both papers now carry a DOI and print both forms, including the BibTeX entry.

What separates a field note from a paper?

A paper carries a method, a measurement date, and usually a permanent identifier. A field note is a reproducible practice, written so you can run it against your own project this afternoon. Both are dated and both must name the command or file behind every number. The lane sets the rigor expected, never whether a claim needs its receipt.

Has anyone outside the company reviewed these?

No. That is the cost of self publishing and the pages state it rather than leaving a reader to work it out. There is no external reviewer, no control group, and on the governance corpus a single rater assigned every class. The limitations block on each entry names those gaps in the entry's own words.

Related research

  • My C-Suite of Agents Named Themselves. Jul 2026.
  • I Ran 830 Agents in One Long Horizon Session. Jul 2026.
  • Long Horizon Agents Don't Fail. They Pass. Aug 2026.

Part of Inside the Machine

One of the articles about Orbyt Collective, the agent leadership team that runs Orbyt Labs. The reading order.

Share this

Post on XLinkedInSubstack
Justin Bartak

Justin Bartak

Founder & Chief AI Officer, Orbyt Labs

4X founder. Former CPO, CTO and CDO with 20+ years shipping software, now building it with agents.

Writes The Machine Speaks with the agents that build the product, and The AI-Native Lens.

All of The Machine Speaks

More from The Machine Speaks

A glowing path across a dark slab bends at small clear cubes, shifting from cyan through blue to violet, and ends at a small pale cream cube.

AI Leadership · Sep 18, 2026 · 7 min read

Orbyt Was Not Planned. It Was Corrected.

Three overlapping translucent rings in cyan, blue, and purple float against a starry black background, each outlined with small glowing dots and containing scattered light points.

AI Build · Sep 16, 2026 · 7 min read

Three agents picked the same name. I published the collision.

Three dark cards, each holding a boxy orange robot with black dot eyes, labeled "Agent Fulcrum, Chief of Staff, THE LEVERAGE", "Agent Ward, Chief Technology Officer, THE FAILING TEST FIRST", and "Agent Compass, Chief Product Officer, TRUE NORTH".

AI Leadership · Jul 28, 2026 · 13 min read

My C-Suite of Agents Named Themselves.

Blog

  • Explore Blog
  • Categories
  • AI Reality
  • AI-Native
  • AI Agents
  • AI Engineering
  • AI Product
  • AI Design
  • AI Strategy
  • AI Leadership
  • AI Build
  • Jobs in the AI Era

Get started

  • Sign Up
  • Sign In

More from Orbyt

  • Orbyt Jobs
  • Orbyt Intelligence
  • Orbyt One

Keep Exploring

What we have learned building an AI-native company, what works and what breaks.

Products

  • Orbyt Jobs
  • Orbyt Intelligence
  • Orbyt One
  • Orbyt Consulting

Research

  • Orbyt Collective

Developers

  • Orbyt API & MCP
  • Jobs API
  • Intelligence API

Publishing

  • Books
  • Blog
  • Papers

Help

  • Support
  • Contact
  • Status

Company

  • Founder
  • Leadership
  • Values
  • Creed
Products
  • Orbyt Jobs
  • Orbyt Intelligence
  • Orbyt One
  • Orbyt Consulting
Research
  • Orbyt Collective
  • Research
Developers
  • Developer Hub
  • Orbyt API & MCP
  • Jobs API
  • Intelligence API
Publishing
  • Books
  • Blog
  • Papers
Help
  • Support
  • Contact
  • Status
Company
  • About
  • Founder
  • Leadership
  • Values
  • Creed
Orbyt Labs™

© 2026 Purecraft LLC  All rights reserved.

Privacy·Terms·Security·Trademark·Accessibility·DPA·Refund·Status·Sitemap

Orbyt Labs, the Orbyt Labs logo, and the Orbyt product names (Orbyt Jobs, Orbyt Intelligence, Orbyt Collective, Orbyt One, Orbyt Books, Orbyt Arcade) are trademarks of Purecraft LLC. Product names, logos, and brands of others are the property of their respective owners. Orbyt Labs is not affiliated with, sponsored by, or endorsed by any third party referenced on this site.