Field note - version 1.0.0
A Model Upgrade on a Frozen Dependency Review
On 17 September 2026, the reviewer accepted 0 of 3 current model reports and 0 of 3 candidate reports. Candidate wins under the pre registered rule: no.
Measured 09-17-2026. Sample: 3 runs per model on the same frozen task.
The data.
- Current model reports accepted by the reviewer0 of 3
- Candidate reports accepted by the reviewer0 of 3
- Current model reports rejected mechanically3 of 3
- Candidate reports rejected mechanically3 of 3
| Metric | Value | How it is counted |
|---|---|---|
| Runs per model | 3 | docs/research/model-upgrade-trial/runs/RESULT.json, trials per model |
| Who performed the substantive reads | agent | docs/research/model-upgrade-trial/runs/RESULT.json, reviewer (absent means founder) |
| Review time clause | not measured: an agent read the reports (Deviations 4) | docs/research/model-upgrade-trial/runs/RESULT.json, review_time_clause |
| Current model | claude-sonnet-4-5-20250929 | docs/research/model-upgrade-trial/runs/RESULT.json, trials.model |
| Current reports accepted by the reviewer | 0 | docs/research/model-upgrade-trial/runs/RESULT.json, trials.human_verdict |
| Current mechanical rejections | 3 | docs/research/model-upgrade-trial/runs/RESULT.json, trials.mechanical_verdict |
| Current unsupported claims found by the reviewer | 12 | docs/research/model-upgrade-trial/runs/RESULT.json, trials.unsupported_claims |
| Current median elapsed minutes | 2.6412833333333334 | docs/research/model-upgrade-trial/runs/RESULT.json, median trials.elapsed_ms / 60000 |
| Current median output tokens | 6513 | docs/research/model-upgrade-trial/runs/<id>/RUN.json, median usage.output_tokens |
| Candidate model | claude-sonnet-5 | docs/research/model-upgrade-trial/runs/RESULT.json, trials.model |
| Candidate reports accepted by the reviewer | 0 | docs/research/model-upgrade-trial/runs/RESULT.json, trials.human_verdict |
| Candidate mechanical rejections | 3 | docs/research/model-upgrade-trial/runs/RESULT.json, trials.mechanical_verdict |
| Candidate unsupported claims found by the reviewer | 14 | docs/research/model-upgrade-trial/runs/RESULT.json, trials.unsupported_claims |
| Candidate median elapsed minutes | 5.590933333333333 | docs/research/model-upgrade-trial/runs/RESULT.json, median trials.elapsed_ms / 60000 |
| Candidate median output tokens | 37456 | docs/research/model-upgrade-trial/runs/<id>/RUN.json, median usage.output_tokens |
| Files changed outside the report lane | 0 | docs/research/model-upgrade-trial/runs/RESULT.json, trials.lane_violations |
| Candidate wins under the pre registered rule | no | docs/research/model-upgrade-trial/runs/RESULT.json, candidate_wins |
| Equal review totals leave the model unchanged | no | docs/research/model-upgrade-trial/runs/RESULT.json, equal_changes_nothing |
| Missing applicable findings | 0 | docs/research/model-upgrade-trial/runs/RESULT.json, trials.missing_findings |
How it was measured.
Locked reviewer verdicts and mechanical verdicts come from the unblinded result. Output tokens come from each run receipt.
The reviewer was the agent. The review time clause was not measured: an agent read the reports (Deviations 4). Equal review totals leave the model unchanged: no.
The registered rule requires every run in each arm to pass and the candidate to use less total review time. Substantive acceptance alone does not establish a win.
The reviewer found 12 unsupported claims across the current model reports and 14 across the candidate reports. An unsupported claim is a sentence the frozen packet does not back.
What this does not show.
- A single seat, a single task and 3 runs per arm. This does not establish performance across other seats or companies.
- The founder did not perform the substantive reads. The reviewer was an agent of a different model family from the reports' authors, reading in a copy of the repository with the letter map and the named run directories removed, and with no network. No human minutes were recorded, so the review time comparison in the pre registered rule was not made.
- The mechanical label detector was calibrated on a single rendering. It rejected 3 current model reports and 3 candidate reports. The protocol records the calibration failures and a second calibration proven after unblind. The evaluator stayed unchanged for this trial.
Revisions.
This URL is permanent. When the data is refreshed the version bumps and a row lands here, so a citation made today still resolves to the finding it cited.
| Version | Date | Change |
|---|---|---|
| 1.0.0 | 09-17-2026 | First publication. |
More from Research.
The other measurements from the same repository, and the two papers they sit beside.
Cite this.
Bartak, J. (2026). A Model Upgrade on a Frozen Dependency Review. Orbyt Labs Research, version 1.0.0. https://www.orbytlabs.ai/research/model-upgrade-trial
CC BY 4.0. Reuse it with attribution.