If Your Experts Leave, Who Evaluates the AI?
An organization can automate more of an expert capability without eliminating its dependence on expertise. Someone—or some independently credible mechanism—must still recognize consequential errors, construct difficult tests, interpret novel cases, and decide when the system should stop. Preserving enough unaided human competence behind that escalation path is therefore an AI-resilience question as well as a workforce question, although current evidence does not establish that AI inevitably causes deskilling.
By Ken Ohyama, Founder · Published August 23, 2026 · Reviewed August 23, 2026
- AI resilience
- expertise retention
- human oversight
Automation can remove the work before it removes the dependency
The first stages of AI adoption often feel additive. An experienced person remains close to the work, sees what the system produces, corrects awkward outputs, and knows when an apparently small error carries a larger consequence. The organization receives faster production while retaining the old reference point.
Over time, the arrangement can change quietly. More cases flow through the system. Routine practice becomes review. Review becomes exception handling. Junior people encounter fewer unassisted problems, and the experts who once knew the terrain retire, move, or spend their attention elsewhere. The capability has not disappeared from the process, but the human benchmark against which it was judged may have thinned.
That sequence is not inevitable. AI can support learning, expose patterns, and make expert practices available to more people. The governance question begins when leaders assume that successful automation has also made the old capability unnecessary. Production and evaluation are different forms of work.
The evaluator needs a capability the workflow may no longer exercise
A reliable escalation path depends on someone recognizing that the case deserves escalation. That may require seeing a weak contradiction, questioning a source, noticing an unfamiliar combination of otherwise normal facts, or knowing that a plausible answer would be unacceptable in this local context.
If the system performs the diagnostic work and presents a polished recommendation, the reviewer may receive less practice forming an independent diagnosis. They can become skilled at operating the interface while losing exposure to the raw conditions from which judgment was originally built. The organization still has a human in the loop, but the loop may no longer preserve an independent point of comparison.
Skagway’s term for this risk is benchmark dependency: the system remains governable only while enough credible expertise survives outside its own outputs to say what good, dangerous, or uncertain looks like. This is a practitioner hypothesis, not an established widespread enterprise failure mode.
Capability path
The benchmark can fade while output improves
The risk lies in the sequence, not in any single use of assistance.
- Practice
The expert performs the work
Frequent unaided exposure maintains pattern recognition, diagnosis, and a reference for acceptable performance.
- Assistance
AI shares more of the task
The expert corrects the system and supplies context while productivity or consistency may improve.
- Exposure
Unaided cases become scarce
Review replaces practice, and junior people encounter fewer of the conditions that once built expertise.
- Failure
The model meets a consequential exception
A fluent output is wrong in a way that requires independent domain judgment to recognize.
- Benchmark
Who still knows?
The organization discovers whether a credible evaluator and manual recovery capability remain.
The sequence is a risk to test, not a prediction that every assisted workforce will lose capability.
Medical evidence makes the concern real and keeps it narrow
A 2025 multicentre observational study examined non-AI-assisted colonoscopies at four Polish centres before and after regular exposure to AI-assisted detection. Adenoma detection in the unassisted procedures declined from 28.4 percent before exposure to 22.4 percent afterward. The authors described the result as suggesting a possible negative effect on endoscopist behaviour.
The study does not establish that AI caused the decline. It was observational, covered one clinical activity, compared two time periods, and may be affected by workload or other changes. A 2026 scoping review found the empirical deskilling literature limited while identifying related evidence across several medical specialties. That supports monitoring and further study; it does not justify a claim that enterprise AI broadly erodes expertise.
The useful lesson is methodological. If unaided capability matters to safety or recovery, measure it directly over time. Do not infer its preservation from the fact that AI-assisted output remains strong.
Emerging evidence
In one four-centre observational study, adenoma detection during unassisted colonoscopies declined from 28.4% before regular AI exposure to 22.4% afterward.
Method note: The study compared non-AI-assisted procedures across two three-month periods. It was observational and cannot establish that AI exposure caused the decline or that the result generalizes beyond this clinical setting.
Current AI reliability keeps the human benchmark economically relevant
The International AI Safety Report 2026 describes uneven capability across tasks and contexts, an evaluation gap between controlled tests and real deployment, and persistent problems with long tasks and unexpected obstacles. It also notes that expert human oversight can mitigate some failures while automation bias can make users trust incorrect output more than warranted.
None of that proves humans will always remain the best evaluators. Automated evaluators, process monitors, and independent models may take on more of the work. Yet a second system creates its own validation question: why should the organization trust this evaluator on the cases that matter, and what independent evidence supports that trust?
For now, consequential deployments should be explicit about where the reference standard comes from. Sometimes it will be an observable outcome. Sometimes a rule or external audit can supply it. In ambiguous work, the benchmark may still depend heavily on people who understand the domain well enough to generate and recognize failure.
Decide what competence has to remain before exposure disappears
The organization does not need every person to remain equally practiced at every task. It needs a deliberate account of residual capability. Who can diagnose without the AI? Who can construct an edge case the development data may not contain? Who understands the escalation threshold and has authority to act when the system disagrees?
The answer should follow consequence and recoverability. A low-consequence, reversible task with clear outcomes may not justify expensive human redundancy. A high-consequence capability with delayed feedback, weak observability, or limited recourse needs a stronger independent benchmark and a credible way to restore manual operation.
The amount of unaided practice required to preserve that competence is not known generically. It will differ by domain, rate of change, and the perceptual or analytical demands of the task. A governance plan should state the assumption and create evidence around it rather than hiding it beneath the phrase “subject-matter expert review.”
Continuity question
What competence may need to remain human?
The answer depends on the capability and consequence. These are functions to assign and test, not a universal staffing prescription.
01Failure recognitionKnow when a plausible output is unacceptable.
Retain enough domain understanding to spot contradictions, missing context, and familiar failure patterns without relying on the system to announce them.
02Edge-case generationConstruct cases outside the routine path.
A credible evaluator should be able to vary conditions, combine unusual factors, and test where the system’s decision boundary changes.
03Independent diagnosisForm a view before seeing the recommendation.
Periodic blind or unaided work helps distinguish independent capability from skill at interpreting a system-framed answer.
04Escalation judgmentRecognize when autonomy should contract.
The reviewer needs both the competence to identify consequential uncertainty and the authority to pause, override, or transfer the case.
05Recovery capabilityContinue when the system is unavailable.
For critical operations, resilience may require a tested manual or alternative path rather than a theoretical fallback stored in a procedure.
Retained capability needs its own operating measures
Measure assisted performance and unaided performance separately. Periodic exercises can ask reviewers to diagnose cases before seeing the model’s answer, detect deliberately seeded errors, explain which evidence they would seek next, and recover from a simulated system outage. The tests should include correct model outputs so reviewers are not trained to perform skepticism theatrically.
An exception library should be maintained by more than one expert where possible. New incidents, near misses, overrides, and environmental changes should refresh it. Disagreement belongs in the record: the organization may need to know that qualified people draw different boundaries rather than pretend that one answer is universally correct.
Capability also has an institutional owner. Someone must decide how often the benchmark is tested, what constitutes unacceptable decay, who receives the results, and what action follows when the remaining expert pool becomes too thin.
Preserving expertise can be pro-automation
The strongest reason to retain human competence is not nostalgia for manual work. It is operational resilience. A system can be used more confidently when the organization knows how failures will be recognized, when autonomy will contract, and how the capability will continue if the tool becomes unavailable or the context changes.
The Atlas is relevant where this is genuinely an institutional-continuity problem: a named critical capability is concentrated in a few people, those people may leave, and the organization needs the judgment, cases, and ownership structure to remain usable beyond one transition. That is narrower than broad AI-resilience or model-validation consulting.
Skagway does not validate AI systems. AI assurance, cybersecurity, regulatory compliance, and technical evaluation require qualified specialists. Skagway’s current contribution begins where a critical human capability is at risk of disappearing before the institution understands what must remain behind the escalation path.
Illustrative example
A manufacturer automates first-pass diagnosis of a rare equipment condition. The tool performs well on ordinary faults, and senior diagnosticians are reassigned. Two years later, an unusual combination of sensor drift and maintenance history produces a confident but unsafe recommendation. The continuity question is not whether every technician should have remained manual. It is whether anyone still had the unaided capability, case access, and authority to recognize the combination before action.
When Skagway is a fit
Skagway Succession is a U.S. executive-succession advisory that captures and transfers the tacit judgment of critical leaders. We are a fit when an organization needs a deliberate, evidence-led process for a critical executive, founder, technical expert, or operator. We are not a replacement for legal, tax, executive-search, compensation, fiduciary, or broad leadership-development advice.
Explore The AtlasGlossary
- Benchmark dependency
- Dependence on surviving expertise or independent evidence to determine whether an automated system’s performance remains acceptable.
- Unaided capability
- The ability to perform or evaluate a defined task without relying on the AI system being assessed.
- Deskilling
- A reduction in practiced human capability or opportunity to maintain it; causes and effects must be established within a specific setting.
- AI resilience
- The ability of an organization and its AI-enabled operation to anticipate, withstand, detect, recover from, and adapt to failures or change.
Sources & further reading
- International AI Safety Report 2026 (opens in a new tab) · International AI Safety Report
- Endoscopist deskilling risk after exposure to artificial intelligence in colonoscopy: a multicentre, observational study (opens in a new tab) · The Lancet Gastroenterology & Hepatology
- Artificial intelligence in medicine: a scoping review of the risk of deskilling and loss of expertise among physicians (opens in a new tab) · ESMO Real World Data and Digital Oncology
- Regulation (EU) 2024/1689, Article 14 — Human oversight (opens in a new tab) · EUR-Lex
This guide is founder-led analysis. Sources provide background and are not endorsements of Skagway Succession.
Continue the research
What took decades to learn
should not disappear in a day.
The road ahead should remember how the company came this far.
