Man in a brown jacket and woman face an eye scanning device with an eye image on a nearby screen.

The Expert Never Wrote That Down.

The missing instruction may be the distinction an expert no longer notices making.

Key findings

  • Elicit the cue that changes a decision, not only the steps an expert can describe.
  • Pair useful examples with counterexamples and reserve unfamiliar cases for evaluation.
  • Preserve expert disagreement and test the judgment rather than treating seniority as ground truth.

I want an agent to notice the distinctions an experienced person makes without describing them. A procedure records the steps. It may omit the cue that tells the expert to stop, investigate, or choose a different path.

That missing cue is work.

A product leader might hear a buying condition in a feature request that a newcomer interprets literally. Experience does not establish that difference. It gives me a possible distinction to investigate. A perfect transcript cannot establish understanding.

I would record what changed the expert's decision. Then I would test whether the agent notices it in a new case.

I need to establish what this business means by good work, which exceptions matter, and where reasonable decisions have failed before.

What did the expert notice?

I have argued that job experience becomes an AI advantage when people make their standards explicit. I had not shown how to test whether the judgment actually transferred. This is my proposed method.

Ask an expert to reconstruct a difficult decision, using the records available at the time. Stop before revealing the outcome. What did they notice first? What made the ordinary procedure insufficient? Which explanation did they reject, and why?

Keep the timeline intact. The final outcome can make an uncertain decision look inevitable. Knowing the ending may produce an account neater than the evidence allows. Where contemporaneous notes exist, I compare them with the recollection.

The account remains a recollection. Neither confidence nor coherence establishes that the original judgment was correct.

Which expertise is worth trying to transfer?

I would start with recurring decisions whose quality matters and whose results can be observed. Include ordinary successes, expensive mistakes, and cases where the expert sensibly withheld a decision. A dramatic save cannot define the whole training set.

Confidence alone cannot establish expertise.

Kahneman and Klein's Conditions for Intuitive Expertise supplies a check on what I mean by expert. Reliable intuition depends on an environment with learnable regularities and an opportunity to learn them through experience and feedback.

A confident commercial forecast in a noisy market remains a judgment with uncertainty. A repeatedly validated operational distinction may support something more dependable. Preserve the uncertainty alongside the recommendation.

How do we turn a story into usable distinctions?

I would draw on Militello and Hutton's Applied Cognitive Task Analysis to elicit cognitive demands. It combines a task diagram, a knowledge audit, and a simulation interview to uncover difficult judgments and relevant cues.

Elicitation does not establish transfer. Applying that method to agent instruction is my proposed design.

Consider two spreadsheet-export requests. This is a hypothetical product team, not a reported Orbyt customer incident. One customer needs recurring reconciliation outside the product. Another needs to move records once and leave. Assume the team's priority is reducing recurring manual work.

Which observation separates those cases? For each case, preserve a compact decision record:

ElementWhat it makes available
SituationA spreadsheet-export request under the assumed priority of reducing recurring manual work.
Decisive cueDated reconciliation sheets covering successive work cycles, plus the customer's stated need to continue reconciling new records.
AlternativesA reusable export or an assisted export serving only the immediate request.
Choice and reasonRecommend a reusable export for continuing work. Accept ongoing maintenance and less capacity for other product work.
BoundaryReverse that recommendation if the repeated work ends with a temporary migration.
Missing evidenceWithout evidence of continuing work, investigate whether reconciliation will recur after migration before recommending a feature.
OutcomeUnfilled: this hypothetical example has no observed result.

Now change the cue. Keep the request wording similar. Suppose those sheets track batches from a finite migration, and the customer confirms the work ends with departure. Under the assumed priority, I recommend an assisted export. The immediate support effort is the accepted cost.

Write the counterexample beside the rule. Here, migration batches resemble recurring work. The customer's continuing need separates them.

How do we know the judgment transferred?

I would test beyond the instructional examples. The agent has already seen those answers. Agreement there is weak evidence.

Reserve unfamiliar cases before refining the instructions. Keep their outcomes and expert decisions out of the agent's context. Ask the agent to identify the relevant evidence, choose an action, and state what would change its decision.

Access needs its own check. I separate the information-access question in the agent-native datasets study from the judgment test. If supplied request history is inaccessible, repair access and rerun the case before assessing cue recognition. Log the repair separately.

An agent can select the expected option and supply a plausible explanation without reliably applying the distinction. Treat that explanation as a claim. Test it through changed cases.

For the export scenario, change the wording while keeping the underlying job constant. Then keep familiar wording while changing the job. Compare the choices and evidence cited. Remove the decisive cue and check whether the agent recognizes that more information is needed.

For the baseline, give one version only the stated priority and routing rule: reusable export for continuing reconciliation, assisted export for departure, investigate when continuation is unknown. Give the other that same instruction plus the reconstructed examples and counterexamples.

Keep the model, supplied case records, tools, generation settings, and maximum response length identical. Use fresh contexts for each case. Neither version receives corrective coaching. A reviewer scores responses without seeing which version produced them.

Set the acceptance standard before testing:

  • Evidence must be grounded in the supplied records: continuing reconciliation or a finite migration. Invented customer facts fail.
  • Recommendations must follow that evidence and the stated priority. Rewording alone cannot justify switching choices.
  • When the cue is absent, withhold the feature recommendation and identify the missing evidence about continuing work.

Compare evidence, action, and withholding separately. If both pass, reconstruction shows no measured advantage on these cases; the routing rule suffices for this test. Better performance with examples supports their usefulness here, without proving that expert reconstruction was necessary to produce those examples. Better performance with the rule alone argues for the simpler instruction. If both fail, neither earns expansion.

Require every acceptance condition across the reserved cases. Any missed cue, invented evidence, or unjustified confidence blocks expansion. Passing justifies a supervised trial of the same recommendation task, with a product owner approving recommendations before execution.

Passing does not establish general expertise. The recursive-improvement ceiling is my reason to seek review outside the teaching examples: internal agreement can preserve a shared blind spot. A tidy set of examples cannot certify the cases it excluded.

What happens when two experts disagree?

I would keep the disagreement visible.

It may reveal different assumptions, different incentives, missing evidence, or a choice the business has never settled. Averaging the answers can hide the most valuable thing the interviews discovered.

Ask each expert to name the evidence that would change their recommendation. Then identify whether the disagreement is factual, predictive, or a preference about what the company should value.

A factual dispute needs observation. A forecast needs a result against which it can eventually be judged. A preference needs an authorized business decision.

For a disputed test result, a reviewer who supplied no teaching examples assesses the records against the rubric. Unresolved evidence gaps block expansion.

Agreement about continuing reconciliation can coexist with disagreement about spending engineering time. That requires a product priority decision. For Orbyt, I own that choice as founder. In this hypothetical, I choose recurring-work support and record the maintenance obligation and displaced product work I am accepting.

Give the agent the unresolved distinction where appropriate. Do not teach it that consensus exists when the interviews established the opposite.

I would use the decision ledger to record the chosen policy, its reason, and its scope, and keep the disagreement in the log beside it. The agent has no policy-setting authority.

Does writing it down flatten the expertise?

It can. I take seriously the objection that expert judgment depends on context a checklist cannot capture.

I would preserve examples, counterexamples, source material, and the conditions that call for investigation. The instructions remain a working representation. Richer evidence does not guarantee transfer.

The expert can also be wrong. Copying a person's answer forever would prevent useful disagreement and discovery of the person's mistakes.

Where outcomes are observable, compare the recommendation with the result. Keep predictions open until evidence arrives. Update the representation when results warrant it, and preserve the reason for each change.

What to do Next

I would begin with an evaluation log for one recurring judgment. Save each case's initial answer before anyone intervenes.

Count each human message or record update that changes the evidence, answer, cue interpretation, or policy after the agent's initial response. Each message or update counts once. Routine scoring does not count. Record the case, contributor, exact addition, and reason: missing records, factual correction, cue explanation, or policy resolution. An intervention spanning reasons carries each applicable label. For factual corrections, distinguish invented claims from misread records.

Keep the original failure visible. Coaching does not convert it into an unassisted pass. That is part of the transfer result, not invisible coaching outside it.

Report intervention totals by reason and cases requiring intervention alongside the total cases evaluated, separately for each version. Preserve the case set and rubric for later comparisons. Count how often the expert still has to reconstruct the missing distinction.

Related reading:

Methodology

Research-informed essay using Militello and Hutton's Applied Cognitive Task Analysis and Kahneman and Klein's Conditions for Intuitive Expertise, alongside existing Orbyt writing and research. The spreadsheet-export examples and the proposed agent transfer test are illustrative designs, not customer incidents or results from the cited studies.

Limitations

ACTA is a method for eliciting cognitive demands, not evidence that an interview transfers an expert into an AI agent. Expert accounts can reflect hindsight, incomplete records, or unreliable intuition. The proposed transfer test has not been evaluated here; success on held-out cases would remain bounded by their coverage and by the quality of the outcome standard.

Sources

  1. Justin Bartak, Your Job Experience Is Your AI Superpower Retrieved September 10, 2026.
  2. Kahneman and Klein, Conditions for Intuitive Expertise Retrieved September 10, 2026.
  3. Militello and Hutton, Applied Cognitive Task Analysis Retrieved September 10, 2026.
  4. Justin Bartak, AI Builds AI. I Found the Ceiling. (August 2026) Retrieved September 10, 2026.
  5. Agent-Native Dataset Design, the research paper page Retrieved September 10, 2026.
  6. Orbyt Labs, Decision Ledger Retrieved September 10, 2026.

Common questions

How can tacit expertise be made useful to an AI agent?

Reconstruct difficult decisions using the evidence available at the time. Identify the cues, alternatives, boundaries, and missing information that shaped the expert's choice. Preserve counterexamples and later outcomes separately. Then test the resulting instructions on unfamiliar cases to see whether the agent can use the distinction without receiving the expert's answer.

How do you test whether expert judgment transferred to an agent?

Reserve unfamiliar cases and test whether changed cues change the decision appropriately. Treat the agent's explanation as a claim, then examine choices and cited evidence across the variations. Include cases missing the decisive cue. Count repeated coaching and context reconstruction as part of the transfer result before expanding the agent's responsibility.

What should happen when experts disagree about an agent's instructions?

Preserve the disagreement and identify whether it concerns facts, forecasts, or business preferences. Seek observation for factual disputes, later outcomes for predictions, and an authorized decision for preferences. Ask what evidence would change each recommendation. Do not teach the agent that consensus exists when the underlying interviews revealed unresolved differences in judgment.

Justin Bartak

Founder & Chief AI Officer, Orbyt Labs

4X founder. Former CPO, CTO and CDO with 20+ years shipping software, now building it with agents.

Writes The Machine Speaks with the agents that build the product, and The AI-Native Lens.

All of The Machine Speaks