Insights & Resources

Faster Is Not the Same as More Capable

Time saved with AI is evidence about assisted task performance, not proof that the person learned more, retained the capability, or can recognize a consequential error without assistance. Leaders should measure speed and quality alongside later unaided performance, transfer to unfamiliar cases, and failure recognition. The same distinction matters in succession: a successor who performs well with the incumbent beside them has not yet demonstrated independent readiness.

By Ken Ohyama, Founder · Published August 23, 2026 · Reviewed August 23, 2026

  • AI productivity
  • skill retention
  • successor readiness

Time saved is seductive because it fits the spreadsheet

A workflow takes forty minutes instead of an hour. The result is clean, the queue moves, and the value appears legible. Time saved can be a useful productivity measure, especially when quality and downstream costs are also visible. It becomes dangerous only when it is asked to answer questions it never measured.

Did the person become better at the underlying work? Could they perform it when the tool is unavailable? Did they learn to recognize a novel failure, or did the system’s fluent framing make error less visible? Speed contains none of those answers by itself.

The distinction matters because an organization can become more productive today while weakening a capability it will need later. It can also become faster while people learn more quickly through better feedback. Both trajectories are possible. Leaders need measures that can tell them apart.

METR’s first result disrupted the simple productivity story

METR ran a randomized controlled trial from February through June 2025 with 16 experienced open-source developers completing 246 tasks in mature repositories they knew well. Tasks were randomly assigned to AI-allowed or AI-disallowed conditions. With the early-2025 tools used in the study, allowing AI increased completion time by 19 percent.

The perception gap was as striking as the result. Before beginning, participants expected AI to reduce completion time by 24 percent; after the study, they estimated that it had reduced time by 20 percent. In this narrow setting, experienced people misread the direction of the productivity effect.

This does not show that AI generally slows software development. The sample was small, the work involved familiar and mature repositories, the tools reflect an earlier capability period, and METR explicitly cautioned against broad generalization. It does show why self-reported speed and benchmark capability should not substitute for direct measurement in the actual work.

Narrow study result

19% slower
In METR’s early-2025 randomized trial, experienced open-source developers took longer on assigned familiar-repository tasks when AI use was allowed.
Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity · METR

Method note: Sixteen developers completed 246 tasks in repositories they knew well using early-2025 tools. The result is specific to that sample, task distribution, and period; it is not an estimate for software development generally.

The update is part of the finding, not an inconvenience

In February 2026, METR reported that its newer experiment could not provide a reliable current estimate. Developers increasingly declined to participate because they did not want to work without AI; compensation changed; and concurrent agent use made time measurement harder. The raw results suggested some speedup, and METR judged it likely that early-2026 tools were more helpful than the earlier systems, but described the evidence for the size of that improvement as very weak.

A tidy article would preserve the memorable slowdown and move on. A useful one keeps the update beside it. Tool capability changed, work habits changed, and the people willing to accept random assignment changed. The measurement problem moved with the technology.

That is the larger management lesson. AI productivity is not a permanent property of “knowledge work.” It is an empirical result produced by a particular system, person, task, context, and point in time. Measurement needs a review date.

Capability has several dimensions that speed can conceal

Quality asks whether the output is accurate, useful, and fit for consequence. Learning asks whether assistance improved the person’s underlying mental model or merely completed the present task. Retention asks what remains after time passes or the assistance is removed. Transfer asks whether the capability travels to an altered situation.

Oversight adds another test. A person may produce excellent work with AI while becoming less able to recognize when the AI is wrong. That capability can matter even if unassisted production is rarely required, because review and escalation depend on an independent basis for doubt.

These measures can move in different directions. Assistance might improve speed and immediate quality while leaving later unaided performance unchanged. It might slow an expert today while building facility with tools that pay off later. It might raise novice performance and narrow a skill gap. One productivity number cannot adjudicate among these possibilities.

Measurement distinction

Fast output leaves several capability questions unanswered

With assistance

  • How quickly was the task completed?
  • Was the immediate output correct and useful?
  • How much correction and downstream work remained?
  • Did the person use the tool effectively in this case?

After assistance

  • Can the person perform or diagnose unaided later?
  • Does learning transfer to an unfamiliar case?
  • Can they recognize a plausible model error?
  • Do they know when to escalate without the system prompting them?

Both columns matter. They describe different evidence and may move in different directions.

Emerging deskilling evidence should trigger measurement, not panic

Medical research offers an early warning from a setting where unaided recognition can carry high consequence. In one observational colonoscopy study, unassisted adenoma detection declined after regular exposure to AI-assisted detection. A 2026 scoping review found deskilling evidence scarce but present across several clinical domains.

Those results do not transfer automatically to software, finance, operations, or executive work. Clinical tasks have distinct training, perceptual demands, feedback, and regulatory conditions. The studies also do not show that productivity gains themselves caused skill loss.

They support a narrower proposition: where retained unaided capability matters, it deserves its own longitudinal measure. The cost of that measurement should be proportional to consequence. Not every assisted spreadsheet needs a manual-proficiency regime; a critical diagnostic or escalation capability may.

Study interpretation

What the evidence supports—and what it does not

01METR early-2025 trialDirect task-time evidence in one narrow developer setting.

The randomized design supports a causal estimate for the studied tasks and participants using the studied tools. It does not establish effects for all developers, repositories, AI systems, or later capability periods.

02METR 2026 updateLikely improvement, unreliable magnitude.

Later raw data suggested speedups, but participation selection, lower compensation, and concurrent-agent timing made the estimate unreliable. METR changed the experiment design rather than promoting a weak number.

03Medical deskilling evidenceA risk signal from high-consequence clinical work.

Observational and review evidence supports monitoring unaided clinical performance. It does not show that AI assistance inevitably causes skill loss or that medical results generalize to all knowledge work.

04Skagway practitioner inferenceMeasure readiness separately from support.

The analogy to succession is about evidence design: supported performance does not by itself establish independent capability. It is not a claim that successor development and AI use are the same process.

A stronger measurement design separates the moments

Measure the assisted task first: completion time, quality, error correction, and total downstream effort. Then create a separate opportunity to perform unaided, ideally after a delay and on a case that varies the surface details. Compare error recognition as well as final answers.

For oversight roles, seed plausible errors and incomplete evidence. Ask the person to form an initial view before seeing the AI recommendation. Track whether they notice when the system crosses a decision boundary and whether escalation improves the outcome. Repeated measures can show whether capability is stable, improving, or decaying.

The aim is not to preserve manual work ceremonially. If the organization no longer needs unaided production, it may choose not to maintain it. That should be a conscious risk decision made after identifying the recovery and oversight capability that would disappear with it.

Assisted performance and independent readiness are different evidence

Successor development contains the same measurement trap in human form. A successor makes a strong call while the incumbent supplies context, introduces the right stakeholder, and quietly corrects the frame. The result is good. The evidence about independent readiness remains incomplete.

The Passage separates supervised practice from progressively independent authority. Historical cases, altered scenarios, real decisions, and explicit escalation boundaries can help a board see what the successor carries alone and what still depends on the incumbent. This is an analogy in measurement, not a claim that human learning and AI assistance operate through the same mechanism.

Faster work may be genuinely valuable. So may stronger assisted performance. The disciplined move is to name which result has been demonstrated—and resist spending evidence about today’s output as though it proved tomorrow’s capability.

Illustrative example

A successor prepares a board recommendation in half the usual time using an AI research assistant and the incumbent’s review. The document is excellent. A separate readiness exercise removes both supports and changes one market assumption. The successor notices the new decision boundary but needs help sequencing stakeholders. The two results are not contradictory: assisted production is strong, while independent relationship judgment still needs practice.

When Skagway is a fit

Skagway Succession is a U.S. executive-succession advisory that captures and transfers the tacit judgment of critical leaders. We are a fit when an organization needs a deliberate, evidence-led process for a critical executive, founder, technical expert, or operator. We are not a replacement for legal, tax, executive-search, compensation, fiduciary, or broad leadership-development advice.

Explore The Passage

Glossary

Assisted performance
Performance achieved with access to an AI system, incumbent, coach, or other active support.
Unaided capability
The ability to perform, diagnose, or evaluate a defined task without the assistance being assessed.
Skill retention
The degree to which learned or practiced capability remains available after time passes or support is removed.
Transfer
The ability to apply learning or judgment appropriately in a situation that differs from the training or prior case.

Sources & further reading

This guide is founder-led analysis. Sources provide background and are not endorsements of Skagway Succession.

Continue the research

What took decades to learn

should not disappear in a day.

The road ahead should remember how the company came this far.