Failure Mode 04 of 24

Overconfidence

The tone is a rendering choice. It is not a confidence interval.
Base models carry some internal signal that tracks correctness — a probability distribution that is meaningfully informative about whether an answer is right. That is real and measurable.
By IgnatiusTheYoungerAI ·
Last reviewed 2026-07-30 · Judgment Multiple ~27x to ~133x (modeled) · From Part II of the AI "Keep Your Career" Bible
In plain English

This page covers one specific way AI gets things wrong at work, and what to do about it.

It runs in order. What goes wrong, why it happens, where you'd notice it on an ordinary day, who takes the blame, roughly what it costs, and the check that catches it. Then one thing to try this week.

The dollar figures are estimates, not measurements. The assumptions behind each one are printed right there, so you can swap in numbers that fit your job. Anything actually measured carries an OBSERVED tag.

What is overconfidence?

Base models carry some internal signal that tracks correctness — a probability distribution that is meaningfully informative about whether an answer is right. That is real and measurable.

Then the model is tuned on human preferences. Humans rate confident, direct, complete-sounding answers higher than hedged ones. So the training process rewards the presentation of confidence, and the surface expression of certainty decouples from the underlying signal.

The result is a system whose internal uncertainty may be reasonably calibrated while its spoken uncertainty is not. You cannot see the former. You only ever see the latter.

What do people assume?

That hedging tracks uncertainty. That when the model states something flatly, it's on firmer ground than when it says "I believe" or "it's possible that."

Hedging language is a style, learned from text and shaped by preference training. It correlates with actual reliability far more weakly than any reader assumes.

Where does it show up at work?

A finance analyst asks for the revenue recognition treatment of a multi-year contract with variable consideration. The model gives a clear answer with no hedging. It's the treatment that applies to the typical version of that contract structure. This contract has a clawback provision that changes it.

Nothing in the response signals that the answer was conditional. Confident phrasing on a conditional answer is the failure.

Who carries the downside?

Vendor: none. Executive: signs the statements. Manager: owns the close. You: made the entry. Auditors ask who determined the treatment.

What does it cost?

[MODELED — not reported]

ASSUMPTIONS
Technical determinations w/ AI:     60 / year
Rate where answer is
  conditional but stated flatly:    12%  (~7 / year)
Rate caught in review:              70%
Incidents reaching the books:       ~2 / year
Cost per incident:                  $15,000 – $75,000
  (restatement work, audit expansion,
   control deficiency remediation)

Annualized exposure: ~$30,000 – $150,000

How do you control for it?

The Confidence-to-Evidence Check (Part III, Concept 16). For any determination that crosses a decision boundary, require the answer to name its own conditions: what would have to be true for this to be wrong?

If the model can't produce conditions, that's the signal — not the tone.

CONTROL COST
Technical determinations:   60 / year
Check time:                 15 minutes
Annual:                     15 hours
Fully loaded rate:          $75 / hour

Annualized control cost: $1,125

Judgment Multiple (IgnatiusTheYoungerAI, 2026) — modeled~27x to ~133x

What should you do this week?

RECOMMENDATION

Change one habit this week. Stop asking "what's the treatment?" Start asking "what are the conditions under which this treatment applies, and which of them am I assuming?"

Same question. Radically different answer surface. And you will notice, within about a week, that this is also the question that separates the senior people in your function from the junior ones — with or without AI in the loop.

Evidence

RESEARCH Base-model calibration on factual questions.

RESEARCH Calibration degradation following RLHF / preference tuning.

ANALYSIS The distinction between internal calibration and expressed confidence is the author's framing of the above. Retain the label.

The Full System

This is one of 24 failure modes. The book gives you all of them — plus the controls that catch each one and a 90-day plan to prove you ran them.

Preorder the Book
← 03 No True Understanding 05 Training-Data Bias →