Insights & Resources

The Most Important AI Skill May Be Knowing When Not to Use It

Effective AI use requires more than prompt skill or broad adoption. Leaders need a current, task-level understanding of where assistance improves performance, where it degrades judgment, and which signal should cause a person to stop, verify, or escalate. Because that boundary can move and appear between similar-looking tasks, capability calibration must become an organizational practice rather than one expert’s private instinct.

By Ken Ohyama, Founder · Published August 30, 2026 · Reviewed August 30, 2026

  • AI capability
  • human judgment
  • AI governance

At a glance

Key takeaways

  • AI assistance can improve and degrade performance within the same professional workflow.
  • In a 758-consultant experiment, AI users completed more work faster inside the tested frontier but were less likely to solve one outside-frontier task correctly.
  • The useful boundary is task- and system-specific, changes over time, and may not match human perceptions of difficulty.
  • Companies need retained human benchmarks, exception cases, escalation routes, and owners responsible for reviewing where AI belongs.

Two assignments can look equally suitable for AI

A team uses the same model to draft a market view and diagnose a difficult business problem. The first answer is faster, broader, and more useful than the work the team usually produces alone. The second is fluent, organized, and wrong in a way that takes an experienced person several minutes to notice.

From the outside, both assignments involved reading information, finding a pattern, and writing a recommendation. The model did not fail at the obviously harder task. It crossed a boundary the workflow never named.

That is why “use AI for easy work” offers little protection. Human difficulty and model capability do not line up neatly. The scarce skill may be recognizing which side of the boundary the present case occupies before confidence hardens around the answer.

The frontier experiment produced gains and a reversal

Dell’Acqua and colleagues ran a preregistered laboratory-in-the-field experiment with 758 Boston Consulting Group consultants. Participants were assigned to work without AI, with GPT-4, or with GPT-4 plus a prompt-engineering overview. The researchers designed realistic consulting tasks both inside and outside the system’s tested capability frontier.[Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality]

Across 18 inside-frontier tasks, AI users completed 12.2 percent more tasks and worked 25.1 percent faster on average, while also producing higher-quality solutions. On a complex managerial problem selected outside the frontier, participants using AI were 19 percent less likely to produce a correct solution than participants without it.[Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality]

The experiment used GPT-4 as available in 2023, a particular set of consulting tasks, and a highly selected professional population. Its percentages are not current universal estimates for knowledge work. The enduring finding is structural: one system can help and harm across nearby parts of the same workflow.

One experiment, opposite directions

AI assistance crossed a boundary inside one professional workflow

Inside the tested frontier

  • 12.2% more tasks completed
  • 25.1% faster completion on average
  • Higher-quality solutions across the tested tasks
  • Eighteen realistic consulting tasks

Outside the tested frontier

  • 19% lower likelihood of a correct solution
  • A complex managerial problem with plausible misleading output
  • AI use degraded human performance on the selected task
  • The boundary was difficult to infer from apparent task difficulty

The figures describe GPT-4, the study population, and tasks tested in the experiment. They demonstrate jaggedness, not the permanent location of today’s frontier.

Jaggedness makes broad policies brittle

A company may divide work into categories such as research, analysis, drafting, and decision support. The frontier can cut through those labels. AI may summarize one source set accurately and lose a decisive qualification in another. It may generate useful alternatives and then anchor the team to a false premise during evaluation.

A permanent allowlist cannot fully solve a moving capability problem. Models change, integrations change, source data change, and users learn new ways to combine them. A task that once required human ownership may later become reliably assistive. A workflow that performed well in testing may deteriorate when the context, stakes, or data source changes.

Policy still matters. It should establish consequence, authority, privacy, legal, security, and escalation boundaries. Operational calibration sits beneath that policy: the continuing work of discovering where the tool’s answer can be trusted, challenged, or refused in this company’s actual decisions.

The warning sign may be an exception, not a low score

Some failures announce themselves through missing citations, impossible numbers, or disagreement with a known record. The dangerous ones fit the visible pattern. The output uses the right language and most of the right evidence while missing the one exception that changes the action.

Experts may recognize that boundary through a small cue: a customer type the historical data barely represents, a regulation that changes recourse, a measurement generated under different operating conditions, or a relationship obligation absent from the prompt. Their value lies partly in knowing when the recurring answer has stopped applying.

Critical Decision Method can recover incidents where a familiar path broke. Cognitive Task Analysis can identify the mental demands and cues around verification. Naturalistic Decision Making keeps those cases tied to consequence, time pressure, and the realities of who can act. None of these methods validates a model; they can help a technical evaluation include the human decision terrain the workflow otherwise omits.

Build an exception library from real work

Start with decisions where AI is already used or being considered. Collect representative cases and the cases experienced people find uncomfortable: ambiguous evidence, conflicting goals, unusual stakeholders, silent failures, and situations where the cost of a plausible mistake is high.

For each case, preserve the source materials, expected boundary, expert rationale, acceptable actions, and escalation route. Test the full workflow rather than the model in isolation. Did the person notice the problem? Did the interface invite verification? Could the reviewer access independent evidence? Did escalation reach someone qualified to decide?

Before You Automate Expert Work, Find the Exceptions explains how this kind of case library can begin with actual incidents rather than a catalogue of imagined edge cases.

A living boundary record

How a company learns where AI belongs

  1. Decision

    Name the real work and consequence

    Define the decision, source context, affected parties, reversibility, and authority rather than testing a vague task category.

  2. Cases

    Use the middle and the edges

    Test representative work alongside exceptions, conflicting goals, missing evidence, and plausible failure conditions.

  3. Benchmark

    Preserve an independent reading

    Identify qualified people, records, or criteria that can challenge the system without relying on its own answer.

  4. Escalation

    Make stopping possible

    Specify the cue, path, context, and authority required to pause, override, or route the decision.

  5. Review

    Redraw the boundary as conditions move

    Revisit the cases after model, data, workflow, consequence, or human capability changes.

Calibration becomes an organizational memory: what was tested, what failed, what changed, and who can still disagree.

The human benchmark can disappear while the model improves

If the system handles routine cases, people may receive fewer opportunities to build the pattern recognition needed at the edge. If the expert retires, the company may retain the model and lose the person who knew which cases deserved doubt. A capable system can therefore coexist with a weaker independent basis for evaluating it.

The organization should decide which unaided or independently grounded capabilities still matter. Maintain them through selected manual work, blind review, scenario exercises, post-incident reconstruction, or another form suited to the domain. Human involvement earns its place when it provides a credible second reading after the first route becomes unreliable.

If Your Experts Leave, Who Evaluates the AI? develops the succession side of this problem, while What Happens to Apprenticeship When AI Does the Junior Work? examines how the future evaluator may lose formative cases.

Someone must own the moving boundary

Technical teams should own appropriate model evaluation, security, data, monitoring, and assurance work. Business leaders should own the decisions being delegated, the consequence of error, and the operating conditions under which escalation must remain real. Legal and regulatory judgments belong with qualified specialists.

The remaining governance question is organizational: where does the company now depend on one model, dataset, vendor, workflow owner, or evaluator to know whether the answer still fits? A capability map should include the second route, not only the primary system.

The Atlas can examine concentrated decisions and second routes across people and organizational systems. It does not replace technical AI assurance. It helps leadership see where a consequential capability has gathered and what still happens when the familiar answer fails.

Illustrative example

An AI assistant reliably drafts routine supplier-risk summaries. A rare case involves a supplier whose ownership changed after the underlying data was collected and whose component has no qualified alternative. The fluent summary recommends the ordinary mitigation. An experienced operator recognizes the ownership date as a boundary cue, stops the workflow, and routes the decision. The company adds the case to its evaluation set and preserves why the ordinary answer failed.

When Skagway is a fit

Skagway Succession is a U.S. executive-succession advisory that captures and transfers the tacit judgment of critical leaders. We are a fit when an organization needs a deliberate, evidence-led process for a critical executive, founder, technical expert, or operator. We are not a replacement for legal, tax, executive-search, compensation, fiduciary, or broad leadership-development advice.

See The Atlas

Glossary

Jagged technological frontier
An uneven capability boundary where AI assistance can improve some tasks and degrade others that appear similarly difficult to people.
Capability calibration
The continuing practice of aligning reliance on a system with evidence about where it performs reliably and where it does not.
Independent benchmark
A person, case set, record, or evaluation criterion capable of testing an AI-supported answer without depending on that answer for its own validity.
Exception library
A maintained set of real or carefully constructed cases where the ordinary path may fail, used for learning and evaluation.

Sources & further reading

This guide is founder-led analysis. Sources provide background and are not endorsements of Skagway Succession.

Continue the research

What took decades to learn

should not disappear in a day.

The road ahead should remember how the company came this far.