The Expert May Matter More—and Less—at the Same Time
AI can reduce the expertise required for parts of execution while still complementing expertise used to frame the problem, supply domain context, evaluate results, and recover from errors. Customer-support evidence shows large gains for less-skilled workers; Anthropic’s Claude Code usage study finds persistent but modest returns to task-specific expertise. These findings are compatible because they concern different layers of work. The balance may change as tools and tasks change.
By Ken Ohyama, Founder · Published August 23, 2026 · Reviewed August 23, 2026
- AI and expertise
- agentic work
- human judgment
Two credible findings appear to pull in opposite directions
In customer support, generative AI raised productivity most for novice and lower-skilled workers. In Anthropic’s study of Claude Code usage, people with greater task-specific domain expertise were more likely to finish successfully and could direct longer chains of work, although the gap between intermediate and expert users was modest.
One reading says expertise is being compressed. Another says expertise earns a return. The contradiction softens when expertise stops being treated as one indivisible thing. Producing an answer, defining the problem, recognizing an error, and deciding whether the result is fit for use are different parts of performance.
AI can substitute for one layer while complementing another. Which layer matters most depends on the task, its representation, the quality of feedback, and the consequence of being wrong.
Working hypothesis
AI can substitute and complement within the same workflow
AI substitutes more for
- Retrieval represented in accessible data
- Routine execution with observable feedback
- Standardized production from known patterns
AI may complement more
- Problem framing and task-specific constraints
- Evaluation against real-world requirements
- Correction and recovery from unusual failures
These are evidence-aligned dimensions, not permanent borders. The distribution depends on the task and may change as agent capability changes.
Customer support shows substitution where patterns are learnable
Brynjolfsson, Li, and Raymond studied 5,179 customer-support agents during a staggered rollout of a generative-AI assistant. Access increased issues resolved per hour by 14 percent on average and by 34 percent among novice and lower-skilled workers, with minimal effects for experienced and highly skilled workers.
In that environment, the system could suggest responses and retrieve relevant documentation during repeated textual interactions. Some performance once accumulated through experience became available inside the workflow. The novice did not need to generate every response or remember every policy unaided.
That is genuine substitution for parts of execution and retrieval. It does not show that every underlying capability transferred or that unusual, high-consequence work would produce the same distribution of gains.
Agentic coding separates planning from execution—for now
Anthropic analyzed roughly 400,000 Claude Code sessions from about 235,000 people between October 2025 and April 2026 using privacy-preserving, model-based classification. In the typical session, the person made about 70 percent of planning decisions—what to do—while Claude made about 80 percent of execution decisions—how to do it.
The division is not a universal law of agentic work. It describes one product, period, and classification framework. Still, it gives concrete shape to a pattern executives often discuss vaguely: implementation can move toward the model while task direction remains substantially human.
The people bringing more task-specific domain expertise tended to achieve higher success and recover more effectively from errors or misunderstandings. Yet the difference between intermediate and expert users was modest. Deep mastery did not produce an unlimited advantage.
Observed division of labor
Method note: Anthropic analyzed approximately 400,000 sessions from about 235,000 people, October 2025–April 2026, using privacy-preserving model-based classifications. This is product-usage evidence, not an economy-wide causal estimate.
Task-specific expertise is different from occupational identity
Anthropic’s classifier inferred expertise from how precisely a person framed directions, what they asked the system to verify, and whether the user corrected the model or the model corrected the user. A senior software engineer could be a novice in an unfamiliar domain; an accountant with little programming experience could be expert in the reconciliation rules the code had to enforce.
That distinction changes the workforce question. The scarce person may not be the one who can type the implementation. It may be the one who can state the true constraint, notice that a month-end edge case was lost, or recognize that passing tests do not answer the operational requirement.
As agents improve, some of those activities may also move toward the model. The study itself proposes tracking whether returns to expertise decline over time. Any claim that domain experts necessarily become more valuable should therefore carry a date and a measurement plan.
Substitution and complementarity can occur inside one task
Consider a leader asking an agent to build a forecast. Retrieval and standardized production may become much cheaper. The agent can assemble data, draft code, run a model, and produce a presentable output. The human contribution may shift toward deciding which business mechanism the forecast represents, whether a discontinuity makes historical data misleading, and what consequence follows from an error.
That shift does not automatically make the human contribution more valuable. It can shrink the total labor required, move authority, or expose the framing itself to automation. Nor does cheaper production guarantee the organization retains enough skill to diagnose a strange failure when normal feedback breaks.
The useful operating question is granular: which layer of expertise is the tool absorbing, which layer is it amplifying, and which layer remains necessary for independent evaluation or recourse?
The answer should remain unsettled because the system is moving
The customer-support study is causal evidence in a bounded workplace setting. Anthropic’s report is large-scale product telemetry analyzed by the company that makes the product, using model-based classifications of sessions and outcomes. It is descriptive rather than economy-wide causal evidence. The two sources deserve different weights and answer different questions.
For continuity planning, the implication is to preserve neither every old skill nor only the glamorous layer called judgment. Keep enough unaided and independent capability to define consequential requirements, test unusual cases, recognize material error, and govern escalation—where the task actually demands it. Let tools remove dependencies they can demonstrably remove.
Skagway will continue treating this as a research question. The balance among execution, framing, evaluation, and recovery may move quickly, and conclusions should move with the evidence rather than with a preferred slogan.
Evidence boundary
Why the two studies can both be true
01Different work
One study examined customer-support conversations with repeated historical examples; the other examined interactive agentic coding and related tasks.
02Different outcomes
The customer-support study measured issues resolved per hour. Anthropic classified task composition, collaboration, and session success using product telemetry.
03Different evidence designs
The staggered rollout supports a causal estimate in one workplace. The Anthropic report is a large descriptive analysis conducted by the product developer.
04A shared implication
Expertise should be decomposed. Assistance can compress some execution differences while task-specific context still affects direction, correction, and success.
Illustrative example
An operating partner asks an agent to model a facility consolidation. The agent retrieves benchmarks, writes the model, and produces scenarios faster than the team could unaided. A domain expert notices that the proposed throughput assumption ignores a product-mix constraint that appears only during two seasonal weeks. Execution expertise became less scarce; task framing and exception recognition still changed the decision. A better future model may absorb that distinction too.
When Skagway is a fit
Skagway Succession is a U.S. executive-succession advisory that captures and transfers the tacit judgment of critical leaders. We are a fit when an organization needs a deliberate, evidence-led process for a critical executive, founder, technical expert, or operator. We are not a replacement for legal, tax, executive-search, compensation, fiduciary, or broad leadership-development advice.
Read our editorial boundaryGlossary
- Substitution
- A technology performs part of work that previously required human labor or skill.
- Complementarity
- A technology increases the usefulness or output of a human capability when they are used together.
- Task-specific expertise
- Knowledge and judgment relevant to the particular problem being solved, distinct from job title or general proficiency.
- Model-based classification
- Use of an AI model to categorize or rate source material according to a defined framework rather than relying entirely on human coding.
Sources & further reading
- Generative AI at Work (opens in a new tab) · Erik Brynjolfsson, Danielle Li, and Lindsey R. Raymond · NBER Working Paper 31161; published in The Quarterly Journal of Economics, 140(2) · National Bureau of Economic Research · 2025
- Agentic coding and persistent returns to expertise (opens in a new tab) · Anthropic Economic Research
This guide is founder-led analysis. Sources provide background and are not endorsements of Skagway Succession.
Continue the research
What took decades to learn
should not disappear in a day.
The road ahead should remember how the company came this far.
