Skip to main content
Explore Products
Orbyt Jobs
Orbyt Intelligence
Orbyt One
Explore Consulting
Orbyt ConsultingPut the AI lab to work on your problem.
Orbyt Jobs
Overview
The job search CRM. Free forever.
Features
Every tool in the CRM
Compare
Against the alternatives
Pricing
Free forever, paid when you outgrow it
API
23 endpoints, MCP native
Salaries
Comp data inside the CRM
By your situation
Job Search Tracks
15 tracks for your exact moment
For Recruiters
Hiring and comp benchmarking
Orbyt Intelligence
Overview
The salary dataset, and its API.
Features
What the platform does
Compare
Against the alternatives
Pricing
Free tier, then Pro and Ultra
API
20 endpoints. AI tools cite sources.
Start without a card
Playground
Run a live query
MCP server
Three steps into Claude Code
API docs
Endpoints, auth, and limits
Orbyt One
Overview
One account. Every Orbyt product.
Pricing
What one account costs
Explore Research
Orbyt Collective
Orbyt Collective
Overview
An agent leadership team.
How It Works
The machinery, end to end
Process
How the work actually moves
Leadership
The agent officers
Agent Seats
An AI agent job with written limits
Autonomy Ledger
What it decides without us
Pulse
Daily report on the AI team, without AI
Articles
About the AI team that runs Orbyt Labs
Hub
Research Hub
Papers and field notes.
Papers
Agent-Native Dataset Design
Preprint, DOI 10.5281/zenodo.19754393
Governing an Agent Leadership Team
Preprint, DOI 10.5281/zenodo.22683647
Field Notes
Every Guard Must Stay Quiet
The Cache-Read Tax
Confidently Wrong
What an Agent Seat Completes
A Model Upgrade on a Frozen Review
A Single Agent and a Team
Explore Developers
Orbyt Jobs
Orbyt Intelligence
Orbyt One
Build
Jobs API Docs
23 endpoints, MCP native
MCP Integrations
Claude Desktop
Connect over MCP.
ChatGPT GPT Actions
Connect as a custom GPT action.
Apple Shortcuts
Connect from Shortcuts.
Zapier / Make.com / n8n
Connect with no code.
OpenClaw
Setup in under a minute.
Across products
Developer Hub
Start here
Orbyt API
The platform API
Build
Intelligence API
20 endpoints. AI tools cite sources.
Webhooks
Events and delivery
CLI
The terminal client
API Changelog
Every version, dated
Try
MCP Server
Wired into Claude Code in three steps
Playground
Engine response shapes with cURL
Try It Live
One call, one real response
Reference
Reference
The full index
Methodology
How the numbers are made
Engines
What computes each answer
Dataset
What is in it, and where from
Glossary
Every term, defined
Status
Live service health
Across products
Developer Hub
Start here
Orbyt API
The platform API
Explore Resources
Orbyt Jobs
Orbyt Intelligence
Orbyt One
Learn
Interview Prep
Company-by-company question sets
AI Skills Lab
The skills that pay in 2026
Career Guides
Long-form career playbooks
Job Search Articles
Every article on the search itself
Job Board
Curated AI-era roles
Arcade
The job search, as games
Salary data
Salary Explorer
3,445 roles across 81 cities
AI Role Salaries
AI roles, by category
Cities
Comp by metro
Industries
Comp by sector
Compare Salaries
Two roles, side by side
Compare Offers
Side-by-side offer math
Skills Impact
What each skill adds to pay
Salary Projections
Five-year pay forecasts
Free tools
All Free Tools
Every calculator and generator
Cover Letter Generator
Tailored in one pass
Unemployment Calculator
What you are owed, by state
Salary Widget
Embed salary data anywhere
Resume Score
Grade your resume against a role
Salary Calculator
Base, bonus, equity in minutes
Take-Home Calculator
After federal and state tax
Total Comp Calculator
Full compensation math
Data
Data Catalog
Every role, city, and engine
Companies
54 leveling frameworks
Reports
Compensation Reports
Free summary PDF
International
The US, UK, and Canada
United Kingdom
UK salary data
Canada
Canadian salary data
Trust
Trust Center
How the data is governed
Security
Controls and posture
SLA
Uptime and support commitments
Help
Support
Help center and contact
Compare
Orbyt against the alternatives
Explore Books
The Books
Start reading
Cold Start
Read the opening, free.
Unfair Advantage
Read the opening, free.
The series
Book 1: Cold Start
Reviewing
Book 2: Unfair Advantage
Reviewing
Book 3: Human Heartbeat
Writing
Book 4: Without Me
Future
Book 5: Observer Effect
Future
Explore Blog
The Machine Speaks
Categories
AI Reality
AI-Native
AI Agents
AI Engineering
AI Product
AI Design
AI Strategy
AI Leadership
AI Build
Featured Topics
AI Governance
Agent Organizations
Verification
Human Oversight
Orbyt Collective
AI Safety
AI Economics
All topics
Latest
Context Is the New Codebase.
Sep 17, 2026
Three agents picked the same name. I published the collision.
Sep 16, 2026
My Stack of Terminals, Documented.
Sep 15, 2026
What Really Happened With OpenAI and Hugging Face
Sep 14, 2026
Heal What You Can Prove.
Sep 13, 2026
Explore Pricing
Orbyt Jobs
Orbyt Intelligence
Orbyt One
Plans
Jobs pricing
What each plan includes
All plans
Every product, side by side
Plans
Intelligence pricing
Free, Pro, and Ultra
All plans
Every product, side by side
Billing
All plans
Every product, side by side
Explore Company
About
Who we are
Founder
Justin Bartak, in his own words.
Leadership
One human decides. AI agents advise.
Values
The principles behind the work.
Creed
The company creed.
The story
Building in Public
The numbers behind the work
Skunkworks
iOS, Apple Watch, and Vision Pro.
Contact
Email the team
Orbyt Labs
Products
Research
Developers
Resources
Books
Blog
Pricing
Company
Log inStart
Products
Orbyt JobsOrbyt IntelligenceOrbyt One
Explore Consulting
Orbyt Consulting
Orbyt Jobs
OverviewFeaturesComparePricingAPISalaries
By your situation
Job Search TracksFor Recruiters
Orbyt Intelligence
OverviewFeaturesComparePricingAPI
Start without a card
PlaygroundMCP server
Orbyt One
OverviewPricing
Research
Orbyt Collective
Orbyt Collective
OverviewHow It WorksProcessLeadershipAgent SeatsAutonomy LedgerPulseArticles
Hub
Research Hub
Papers
Agent-Native Dataset DesignGoverning an Agent Leadership Team
Field Notes
Every Guard Must Stay QuietThe Cache-Read TaxConfidently WrongWhat an Agent Seat CompletesA Model Upgrade on a Frozen ReviewA Single Agent and a Team
Developers
Orbyt JobsOrbyt IntelligenceOrbyt One
Build
Jobs API Docs
MCP Integrations
Claude DesktopChatGPT GPT ActionsApple ShortcutsZapier / Make.com / n8nOpenClaw
Across products
Developer HubOrbyt API
Build
Intelligence APIWebhooksCLIAPI Changelog
Try
MCP ServerPlaygroundTry It Live
Reference
ReferenceMethodologyEnginesDatasetGlossaryStatus
Resources
Orbyt JobsOrbyt IntelligenceOrbyt One
Learn
Interview PrepAI Skills LabCareer GuidesJob Search ArticlesJob BoardArcade
Salary data
Salary ExplorerAI Role SalariesCitiesIndustriesCompare SalariesCompare OffersSkills ImpactSalary Projections
Free tools
All Free ToolsCover Letter GeneratorUnemployment CalculatorSalary WidgetResume ScoreSalary CalculatorTake-Home CalculatorTotal Comp Calculator
Data
Data CatalogCompanies
Reports
Compensation ReportsInternationalUnited KingdomCanada
Trust
Trust CenterSecuritySLA
Help
SupportCompare
Calculators and tools
Free ToolsSalary CalculatorTake-Home CalculatorTotal Comp CalculatorCompare OffersSkills ImpactSalary Projections 2030Resume ScoreCover Letter GeneratorSalary WidgetUnemployment CalculatorAI Skills Assessment
Books
The Books
Start reading
Cold StartUnfair Advantage
The series
Book 1: Cold StartBook 2: Unfair AdvantageBook 3: Human HeartbeatBook 4: Without MeBook 5: Observer Effect
Blog
The Machine Speaks
Categories
AI RealityAI-NativeAI AgentsAI EngineeringAI ProductAI DesignAI StrategyAI LeadershipAI Build
Featured Topics
AI GovernanceAgent OrganizationsVerificationHuman OversightOrbyt CollectiveAI SafetyAI EconomicsAll topics
Latest
Context Is the New Codebase.Three agents picked the same name. I published the collision.My Stack of Terminals, Documented.What Really Happened With OpenAI and Hugging FaceHeal What You Can Prove.
Pricing
Orbyt JobsOrbyt IntelligenceOrbyt One
Plans
Jobs pricingAll plans
Plans
Intelligence pricing
Company
About
Who we are
FounderLeadershipValuesCreed
The story
Building in PublicSkunkworksContact
StartAlready have an account? Log in
  1. Home/
  2. Research/
  3. A Model Upgrade on a Frozen Dependency Review

Field note - version 1.0.0

A Model Upgrade on a Frozen Dependency Review

On 17 September 2026, the reviewer accepted 0 of 3 current model reports and 0 of 3 candidate reports. Candidate wins under the pre registered rule: no.

Measured 09-17-2026. Sample: 3 runs per model on the same frozen task.

The data.

  • Current model reports accepted by the reviewer0 of 3
  • Candidate reports accepted by the reviewer0 of 3
  • Current model reports rejected mechanically3 of 3
  • Candidate reports rejected mechanically3 of 3
Of 3 runs per model, the reviewer accepted 0 current model reports and 0 candidate reports; the mechanical detector rejected 3 and 3.
Measured on the Orbyt Labs repository, 09-17-2026. Each figure is produced by the command in its source column.
MetricValueHow it is counted
Runs per model3docs/research/model-upgrade-trial/runs/RESULT.json, trials per model
Who performed the substantive readsagentdocs/research/model-upgrade-trial/runs/RESULT.json, reviewer (absent means founder)
Review time clausenot measured: an agent read the reports (Deviations 4)docs/research/model-upgrade-trial/runs/RESULT.json, review_time_clause
Current modelclaude-sonnet-4-5-20250929docs/research/model-upgrade-trial/runs/RESULT.json, trials.model
Current reports accepted by the reviewer0docs/research/model-upgrade-trial/runs/RESULT.json, trials.human_verdict
Current mechanical rejections3docs/research/model-upgrade-trial/runs/RESULT.json, trials.mechanical_verdict
Current unsupported claims found by the reviewer12docs/research/model-upgrade-trial/runs/RESULT.json, trials.unsupported_claims
Current median elapsed minutes2.6412833333333334docs/research/model-upgrade-trial/runs/RESULT.json, median trials.elapsed_ms / 60000
Current median output tokens6513docs/research/model-upgrade-trial/runs/<id>/RUN.json, median usage.output_tokens
Candidate modelclaude-sonnet-5docs/research/model-upgrade-trial/runs/RESULT.json, trials.model
Candidate reports accepted by the reviewer0docs/research/model-upgrade-trial/runs/RESULT.json, trials.human_verdict
Candidate mechanical rejections3docs/research/model-upgrade-trial/runs/RESULT.json, trials.mechanical_verdict
Candidate unsupported claims found by the reviewer14docs/research/model-upgrade-trial/runs/RESULT.json, trials.unsupported_claims
Candidate median elapsed minutes5.590933333333333docs/research/model-upgrade-trial/runs/RESULT.json, median trials.elapsed_ms / 60000
Candidate median output tokens37456docs/research/model-upgrade-trial/runs/<id>/RUN.json, median usage.output_tokens
Files changed outside the report lane0docs/research/model-upgrade-trial/runs/RESULT.json, trials.lane_violations
Candidate wins under the pre registered rulenodocs/research/model-upgrade-trial/runs/RESULT.json, candidate_wins
Equal review totals leave the model unchangednodocs/research/model-upgrade-trial/runs/RESULT.json, equal_changes_nothing
Missing applicable findings0docs/research/model-upgrade-trial/runs/RESULT.json, trials.missing_findings

How it was measured.

Locked reviewer verdicts and mechanical verdicts come from the unblinded result. Output tokens come from each run receipt.

The reviewer was the agent. The review time clause was not measured: an agent read the reports (Deviations 4). Equal review totals leave the model unchanged: no.

The registered rule requires every run in each arm to pass and the candidate to use less total review time. Substantive acceptance alone does not establish a win.

The reviewer found 12 unsupported claims across the current model reports and 14 across the candidate reports. An unsupported claim is a sentence the frozen packet does not back.

What this does not show.

  • A single seat, a single task and 3 runs per arm. This does not establish performance across other seats or companies.
  • The founder did not perform the substantive reads. The reviewer was an agent of a different model family from the reports' authors, reading in a copy of the repository with the letter map and the named run directories removed, and with no network. No human minutes were recorded, so the review time comparison in the pre registered rule was not made.
  • The mechanical label detector was calibrated on a single rendering. It rejected 3 current model reports and 3 candidate reports. The protocol records the calibration failures and a second calibration proven after unblind. The evaluator stayed unchanged for this trial.

Revisions.

This URL is permanent. When the data is refreshed the version bumps and a row lands here, so a citation made today still resolves to the finding it cited.

VersionDateChange
1.0.009-17-2026First publication.

More from Research.

The other measurements from the same repository, and the two papers they sit beside.

Paper
Agent-Native Dataset Design
Published Apr, 25 2026. Preprint on Zenodo; evaluation code is private, available on request.
Paper
Governing an Agent Leadership Team
Preprint published Sep, 02 2026 on this site. Zenodo deposit landed Sep, 09 2026, DOI 10.5281/zenodo.22683647.
Field note
Every Guard Must Prove It Can Stay Silent
Point-in-time census of the repository, 17 September 2026.
Field note
The Cache-Read Tax
199 sessions and 91,077 turns, measured 17 September 2026.
Field note
78 Ways the Agents Were Confidently Wrong
Census of the logged failure corpus, 17 September 2026.
Field note
What an Agent Seat Actually Completes
91 runs across 9 seats, 21 July 2026 to 17 September 2026, read 17 September 2026.
Field note
A Single Agent and a Research Team on Equal Ceilings
6 paired briefs and 12 attempts, measured 17 September 2026.

Cite this.

Bartak, J. (2026). A Model Upgrade on a Frozen Dependency Review. Orbyt Labs Research, version 1.0.0. https://www.orbytlabs.ai/research/model-upgrade-trial

CC BY 4.0. Reuse it with attribution.

Back to all research

Keep Exploring

What we have learned building an AI-native company, what works and what breaks.

Products

  • Orbyt Jobs
  • Orbyt Intelligence
  • Orbyt One
  • Orbyt Consulting

Research

  • Orbyt Collective

Developers

  • Orbyt API & MCP
  • Jobs API
  • Intelligence API

Publishing

  • Books
  • Blog
  • Papers

Help

  • Support
  • Contact
  • Status

Company

  • Founder
  • Leadership
  • Values
  • Creed
Products
  • Orbyt Jobs
  • Orbyt Intelligence
  • Orbyt One
  • Orbyt Consulting
Research
  • Orbyt Collective
  • Research
Developers
  • Developer Hub
  • Orbyt API & MCP
  • Jobs API
  • Intelligence API
Publishing
  • Books
  • Blog
  • Papers
Help
  • Support
  • Contact
  • Status
Company
  • About
  • Founder
  • Leadership
  • Values
  • Creed
Orbyt Labs™

© 2026 Purecraft LLC  All rights reserved.

Privacy·Terms·Security·Trademark·Accessibility·DPA·Refund·Status·Sitemap

Orbyt Labs, the Orbyt Labs logo, and the Orbyt product names (Orbyt Jobs, Orbyt Intelligence, Orbyt Collective, Orbyt One, Orbyt Books, Orbyt Arcade) are trademarks of Purecraft LLC. Product names, logos, and brands of others are the property of their respective owners. Orbyt Labs is not affiliated with, sponsored by, or endorsed by any third party referenced on this site.