Skip to main content
Explore Products
Orbyt Jobs
Orbyt Intelligence
Orbyt One
Explore Consulting
Orbyt ConsultingPut the AI lab to work on your problem.
Orbyt Jobs
Overview
The job search CRM. Free forever.
Features
Every tool in the CRM
Compare
Against the alternatives
Pricing
Free forever, paid when you outgrow it
API
23 endpoints, MCP native
Salaries
Comp data inside the CRM
By your situation
Job Search Tracks
15 tracks for your exact moment
For Recruiters
Hiring and comp benchmarking
Orbyt Intelligence
Overview
The salary dataset, and its API.
Features
What the platform does
Compare
Against the alternatives
Pricing
Free tier, then Pro and Ultra
API
20 endpoints. AI tools cite sources.
Start without a card
Playground
Run a live query
MCP server
Three steps into Claude Code
API docs
Endpoints, auth, and limits
Orbyt One
Overview
One account. Every Orbyt product.
Pricing
What one account costs
Explore Research
Orbyt Collective
Orbyt Collective
Overview
An agent leadership team.
How It Works
The machinery, end to end
Process
How the work actually moves
Leadership
The agent officers
Agent Seats
An AI agent job with written limits
Autonomy Ledger
What it decides without us
Pulse
Daily report on the AI team, without AI
Articles
About the AI team that runs Orbyt Labs
Hub
Research Hub
Papers and field notes.
Papers
Agent-Native Dataset Design
Preprint, DOI 10.5281/zenodo.19754393
Governing an Agent Leadership Team
Preprint, DOI 10.5281/zenodo.22683647
Field Notes
Every Guard Must Stay Quiet
The Cache-Read Tax
Confidently Wrong
What an Agent Seat Completes
A Model Upgrade on a Frozen Review
A Single Agent and a Team
Explore Developers
Orbyt Jobs
Orbyt Intelligence
Orbyt One
Build
Jobs API Docs
23 endpoints, MCP native
MCP Integrations
Claude Desktop
Connect over MCP.
ChatGPT GPT Actions
Connect as a custom GPT action.
Apple Shortcuts
Connect from Shortcuts.
Zapier / Make.com / n8n
Connect with no code.
OpenClaw
Setup in under a minute.
Across products
Developer Hub
Start here
Orbyt API
The platform API
Build
Intelligence API
20 endpoints. AI tools cite sources.
Webhooks
Events and delivery
CLI
The terminal client
API Changelog
Every version, dated
Try
MCP Server
Wired into Claude Code in three steps
Playground
Engine response shapes with cURL
Try It Live
One call, one real response
Reference
Reference
The full index
Methodology
How the numbers are made
Engines
What computes each answer
Dataset
What is in it, and where from
Glossary
Every term, defined
Status
Live service health
Across products
Developer Hub
Start here
Orbyt API
The platform API
Explore Resources
Orbyt Jobs
Orbyt Intelligence
Orbyt One
Learn
Interview Prep
Company-by-company question sets
AI Skills Lab
The skills that pay in 2026
Career Guides
Long-form career playbooks
Job Search Articles
Every article on the search itself
Job Board
Curated AI-era roles
Arcade
The job search, as games
Salary data
Salary Explorer
3,445 roles across 81 cities
AI Role Salaries
AI roles, by category
Cities
Comp by metro
Industries
Comp by sector
Compare Salaries
Two roles, side by side
Compare Offers
Side-by-side offer math
Skills Impact
What each skill adds to pay
Salary Projections
Five-year pay forecasts
Free tools
All Free Tools
Every calculator and generator
Cover Letter Generator
Tailored in one pass
Unemployment Calculator
What you are owed, by state
Salary Widget
Embed salary data anywhere
Resume Score
Grade your resume against a role
Salary Calculator
Base, bonus, equity in minutes
Take-Home Calculator
After federal and state tax
Total Comp Calculator
Full compensation math
Data
Data Catalog
Every role, city, and engine
Companies
54 leveling frameworks
Reports
Compensation Reports
Free summary PDF
International
The US, UK, and Canada
United Kingdom
UK salary data
Canada
Canadian salary data
Trust
Trust Center
How the data is governed
Security
Controls and posture
SLA
Uptime and support commitments
Help
Support
Help center and contact
Compare
Orbyt against the alternatives
Explore Books
The Books
Start reading
Cold Start
Read the opening, free.
Unfair Advantage
Read the opening, free.
The series
Book 1: Cold Start
Reviewing
Book 2: Unfair Advantage
Reviewing
Book 3: Human Heartbeat
Writing
Book 4: Without Me
Future
Book 5: Observer Effect
Future
Explore Blog
The Machine Speaks
Categories
AI Reality
AI-Native
AI Agents
AI Engineering
AI Product
AI Design
AI Strategy
AI Leadership
AI Build
Featured Topics
AI Governance
Agent Organizations
Verification
Human Oversight
Orbyt Collective
AI Safety
AI Economics
All topics
Latest
Context Is the New Codebase.
Sep 17, 2026
Three agents picked the same name. I published the collision.
Sep 16, 2026
My Stack of Terminals, Documented.
Sep 15, 2026
What Really Happened With OpenAI and Hugging Face
Sep 14, 2026
Heal What You Can Prove.
Sep 13, 2026
Explore Pricing
Orbyt Jobs
Orbyt Intelligence
Orbyt One
Plans
Jobs pricing
What each plan includes
All plans
Every product, side by side
Plans
Intelligence pricing
Free, Pro, and Ultra
All plans
Every product, side by side
Billing
All plans
Every product, side by side
Explore Company
About
Who we are
Founder
Justin Bartak, in his own words.
Leadership
One human decides. AI agents advise.
Values
The principles behind the work.
Creed
The company creed.
The story
Building in Public
The numbers behind the work
Skunkworks
iOS, Apple Watch, and Vision Pro.
Contact
Email the team
Orbyt Labs
Products
Research
Developers
Resources
Books
Blog
Pricing
Company
Log inStart
Products
Orbyt JobsOrbyt IntelligenceOrbyt One
Explore Consulting
Orbyt Consulting
Orbyt Jobs
OverviewFeaturesComparePricingAPISalaries
By your situation
Job Search TracksFor Recruiters
Orbyt Intelligence
OverviewFeaturesComparePricingAPI
Start without a card
PlaygroundMCP server
Orbyt One
OverviewPricing
Research
Orbyt Collective
Orbyt Collective
OverviewHow It WorksProcessLeadershipAgent SeatsAutonomy LedgerPulseArticles
Hub
Research Hub
Papers
Agent-Native Dataset DesignGoverning an Agent Leadership Team
Field Notes
Every Guard Must Stay QuietThe Cache-Read TaxConfidently WrongWhat an Agent Seat CompletesA Model Upgrade on a Frozen ReviewA Single Agent and a Team
Developers
Orbyt JobsOrbyt IntelligenceOrbyt One
Build
Jobs API Docs
MCP Integrations
Claude DesktopChatGPT GPT ActionsApple ShortcutsZapier / Make.com / n8nOpenClaw
Across products
Developer HubOrbyt API
Build
Intelligence APIWebhooksCLIAPI Changelog
Try
MCP ServerPlaygroundTry It Live
Reference
ReferenceMethodologyEnginesDatasetGlossaryStatus
Resources
Orbyt JobsOrbyt IntelligenceOrbyt One
Learn
Interview PrepAI Skills LabCareer GuidesJob Search ArticlesJob BoardArcade
Salary data
Salary ExplorerAI Role SalariesCitiesIndustriesCompare SalariesCompare OffersSkills ImpactSalary Projections
Free tools
All Free ToolsCover Letter GeneratorUnemployment CalculatorSalary WidgetResume ScoreSalary CalculatorTake-Home CalculatorTotal Comp Calculator
Data
Data CatalogCompanies
Reports
Compensation ReportsInternationalUnited KingdomCanada
Trust
Trust CenterSecuritySLA
Help
SupportCompare
Calculators and tools
Free ToolsSalary CalculatorTake-Home CalculatorTotal Comp CalculatorCompare OffersSkills ImpactSalary Projections 2030Resume ScoreCover Letter GeneratorSalary WidgetUnemployment CalculatorAI Skills Assessment
Books
The Books
Start reading
Cold StartUnfair Advantage
The series
Book 1: Cold StartBook 2: Unfair AdvantageBook 3: Human HeartbeatBook 4: Without MeBook 5: Observer Effect
Blog
The Machine Speaks
Categories
AI RealityAI-NativeAI AgentsAI EngineeringAI ProductAI DesignAI StrategyAI LeadershipAI Build
Featured Topics
AI GovernanceAgent OrganizationsVerificationHuman OversightOrbyt CollectiveAI SafetyAI EconomicsAll topics
Latest
Context Is the New Codebase.Three agents picked the same name. I published the collision.My Stack of Terminals, Documented.What Really Happened With OpenAI and Hugging FaceHeal What You Can Prove.
Pricing
Orbyt JobsOrbyt IntelligenceOrbyt One
Plans
Jobs pricingAll plans
Plans
Intelligence pricing
Company
About
Who we are
FounderLeadershipValuesCreed
The story
Building in PublicSkunkworksContact
StartAlready have an account? Log in
  1. Home/
  2. Research/
  3. 78 Ways the Agents Were Confidently Wrong

Field note - version 2.6.0

78 Ways the Agents Were Confidently Wrong

Of 78 distinct engineering failures logged by an AI-run organization as of 17 September 2026, 56 of 78 (72%) raised no error at all and were found only because someone checked an artifact directly, and 54 of 78 (69%) have a named automated countermeasure that exists in the repository.

Collected 08-27-2026 to 09-17-2026. Sample: 78 logged failures in one repository.

The data.

  • Raised no error (silent)56 of 78
  • Have a countermeasure in the tree54 of 78
Of 78 logged failures, 56 of 78 (72%) raised no error at all, and 54 of 78 (69%) now have a countermeasure that exists in the tree.
Measured on the Orbyt Labs repository, 09-17-2026. Each figure is produced by the command in its source column.
MetricValueHow it is counted
Distinct failures logged78docs/lessons-learned.md, numbered entries
Failures that raised NO error (silent)56 of 78 (72%)SILENT in scripts/refresh-research-notes.ts over each entry
Failures with a countermeasure that exists in the tree54 of 78 (69%)public/failure-corpus.json, field `mechanized_count`
Automated guards now in the repository114ls scripts/check-*.cjs | wc -l
Average length of a logged failure443 wordsmean word count across the entries

How it was measured.

A census of one corpus, counted by parsing the numbered entries of the failure log on the date shown, not by recalling how many there are.

The silent-failure count matches entries describing a failure that produced no error, no exception and no red test: a green suite, an exit code of zero, a report rendering clean, or a job concluding success while performing no work. From 2.0.0 the classifier is committed code (SILENT in scripts/refresh-research-notes.ts) rather than a hand read, so the number is reproducible and its false negatives are inspectable.

The countermeasure count is the repository's own failure corpus (public/failure-corpus.json, field mechanized_count): an entry is mechanized only when the check, test or guard it names exists in the tree at generation time. Version 1.0.0 counted entries that named a guard in prose; from 2.0.0 the count is a checkable fact about the tree, and it is a lower bound on remediation because some failures were fixed by deleting the capability rather than guarding it.

Both counts come from text matching over the corpus and are reported as such. They classify how a failure was DESCRIBED, which is a proxy for how it behaved.

The corpus is append-only and each entry carries its own receipt: the command, the log line, or the query that established it.

What this does not show.

  • These are the failures that were CAUGHT and written down, so the share published here is the silent share among LOGGED failures. The rate across all failures, noticed or not, is unknown, and this corpus cannot bound it in either direction: a failure nobody noticed is not necessarily one that raised no error, because it may have errored somewhere nobody was reading, and a corpus that only holds what was found says nothing about the shape of what was not.
  • Text matching classifies prose, not behaviour. An entry that fails to use the word silent is counted as not silent even where the failure was.
  • One organization, one codebase, one set of tools. Nothing here establishes that these failure shapes generalize.
  • A guard existing is not a guard working. This project separately requires every detector to prove it can stay quiet on a clean input, which is a different measurement published separately.
  • No trend is claimed. The changelog carries every refresh, but a rising count is a growing log, not a rising failure rate.

Revisions.

This URL is permanent. When the data is refreshed the version bumps and a row lands here, so a citation made today still resolves to the finding it cited.

VersionDateChange
2.6.009-17-2026Refresh: failures 75 -> 78; silent 53 of 75 (71%) -> 56 of 78 (72%); mechanized 51 of 75 (68%) -> 54 of 78 (69%); guards 113 -> 114; avgWords 422 words -> 443 words.
2.5.009-13-2026Refresh: failures 74 -> 75; silent 52 of 74 (70%) -> 53 of 75 (71%); mechanized 50 of 74 (68%) -> 51 of 75 (68%); avgWords 416 words -> 422 words.
2.4.009-13-2026Refresh: failures 73 -> 74; silent 51 of 73 (70%) -> 52 of 74 (70%); mechanized 49 of 73 (67%) -> 50 of 74 (68%).
2.3.009-12-2026Refresh: avgWords 413 words -> 416 words.
2.2.009-12-2026Refresh: failures 72 -> 73; silent 50 of 72 (69%) -> 51 of 73 (70%); mechanized 48 of 72 (67%) -> 49 of 73 (67%); avgWords 410 words -> 413 words.
2.1.009-12-2026Refresh: avgWords 406 words -> 410 words.
2.0.009-11-2026Method change and refresh: failures 59 -> 72; silent 34 of 59 (58%) -> 50 of 72 (69%); mechanized 37 of 59 (63%) -> 48 of 72 (67%); guards 105 -> 113; avgWords 379 words -> 406 words. The silent classifier is now committed code (SILENT) rather than a hand read, and the countermeasure count is the failure corpus's mechanized_count, an existence check against the tree rather than a mention in prose. The title moves with the count.
1.0.008-27-2026First publication.

More from Research.

The other measurements from the same repository, and the two papers they sit beside.

Paper
Agent-Native Dataset Design
Published Apr, 25 2026. Preprint on Zenodo; evaluation code is private, available on request.
Paper
Governing an Agent Leadership Team
Preprint published Sep, 02 2026 on this site. Zenodo deposit landed Sep, 09 2026, DOI 10.5281/zenodo.22683647.
Field note
Every Guard Must Prove It Can Stay Silent
Point-in-time census of the repository, 17 September 2026.
Field note
The Cache-Read Tax
199 sessions and 91,077 turns, measured 17 September 2026.
Field note
What an Agent Seat Actually Completes
91 runs across 9 seats, 21 July 2026 to 17 September 2026, read 17 September 2026.
Field note
A Model Upgrade on a Frozen Dependency Review
3 runs per model on a frozen dependency review, measured 17 September 2026.
Field note
A Single Agent and a Research Team on Equal Ceilings
6 paired briefs and 12 attempts, measured 17 September 2026.

Cite this.

Bartak, J. (2026). 78 Ways the Agents Were Confidently Wrong. Orbyt Labs Research, version 2.6.0. https://www.orbytlabs.ai/research/confidently-wrong

CC BY 4.0. Reuse it with attribution.

Back to all research

Keep Exploring

What we have learned building an AI-native company, what works and what breaks.

Products

  • Orbyt Jobs
  • Orbyt Intelligence
  • Orbyt One
  • Orbyt Consulting

Research

  • Orbyt Collective

Developers

  • Orbyt API & MCP
  • Jobs API
  • Intelligence API

Publishing

  • Books
  • Blog
  • Papers

Help

  • Support
  • Contact
  • Status

Company

  • Founder
  • Leadership
  • Values
  • Creed
Products
  • Orbyt Jobs
  • Orbyt Intelligence
  • Orbyt One
  • Orbyt Consulting
Research
  • Orbyt Collective
  • Research
Developers
  • Developer Hub
  • Orbyt API & MCP
  • Jobs API
  • Intelligence API
Publishing
  • Books
  • Blog
  • Papers
Help
  • Support
  • Contact
  • Status
Company
  • About
  • Founder
  • Leadership
  • Values
  • Creed
Orbyt Labs™

© 2026 Purecraft LLC  All rights reserved.

Privacy·Terms·Security·Trademark·Accessibility·DPA·Refund·Status·Sitemap

Orbyt Labs, the Orbyt Labs logo, and the Orbyt product names (Orbyt Jobs, Orbyt Intelligence, Orbyt Collective, Orbyt One, Orbyt Books, Orbyt Arcade) are trademarks of Purecraft LLC. Product names, logos, and brands of others are the property of their respective owners. Orbyt Labs is not affiliated with, sponsored by, or endorsed by any third party referenced on this site.