# Polished Output From Bad Data: Training-Data Quality, Mode 15 | IgnatiusTheYoungerAI

> Training-data quality is the ceiling imposed by learning from web-scale text containing errors, outdated material and content optimized for search rather than correctness.

- canonical: https://ignatiustheyoungerai.com/failure-modes/15
- html: https://ignatiustheyoungerai.com/failure-modes/15
- machine-view: https://ignatiustheyoungerai.com/failure-modes/15?view=machine
- markdown: https://ignatiustheyoungerai.com/failure-modes/15.md
- site-index: https://ignatiustheyoungerai.com/llms.txt
- evidence-labels: RESEARCH · REPORTED · ANALYSIS · MODELED · RECOMMENDATION · OBSERVED — carry the label with the claim; MODELED is arithmetic from stated assumptions, never a measured outcome
- publisher: IgnatiusTheYoungerAI (Nathan Battin) · hello@ignatiustheyoungerai.com

---

Failure Mode 15 of 24

# Training-Data Quality

> It learned from the internet, including the parts of the internet that were wrong.

Web-scale training data contains errors, outdated material, marketing copy, forum speculation, and content optimized for search rather than correctness. The model learns statistical regularities across all of it.

By [IgnatiusTheYoungerAI](/about) ·

Last reviewed 2026-07-30 · Judgment Multiple ~5x to ~67x (modeled) · From Part II of the AI "Keep Your Career" Bible

In plain English

**This page covers one specific way AI gets things wrong at work** , and what to do about it.

It runs in order. What goes wrong, why it happens, where you'd notice it on an ordinary day, who takes the blame, roughly what it costs, and the check that catches it. Then one thing to try this week.

The dollar figures are estimates, not measurements. The assumptions behind each one are printed right there, so you can swap in numbers that fit your job. Anything actually measured carries an OBSERVED  tag.

## What is training-data quality?

Web-scale training data contains errors, outdated material, marketing copy, forum speculation, and content optimized for search rather than correctness. The model learns statistical regularities across all of it.

Where a misconception is *widely repeated* , it is well-represented in training, sometimes better-represented than the correction, which is typically stated once in a technical source rather than a thousand times in blog posts. **Frequency in the corpus is not correlated with truth, and in some domains it is inversely correlated.**

Two aggravating factors: technical domains where the correct answer is niche and the popular answer is wrong; and any domain with commercial incentive to publish, where volume tracks marketing spend.

## What do people assume about training-data quality?

That training corpora were curated for accuracy, that some filtering process removed the false, the outdated, and the deliberately misleading.

Filtering targets toxicity, duplication, and low quality by proxy measures. It does not verify factual accuracy at corpus scale, because nothing can.

## Where does training-data quality show up at work?

A benefits analyst asks about the tax treatment of a specific compensation arrangement. The model returns the widely-repeated internet answer, which describes the general case correctly and the analyst's actual case incorrectly.

The general case is what's written about. The exception is in the regulation.

## Who carries the downside?

**Vendor:**  none. **Executive:**  none. **Manager:**  owns the guidance. **You:**  gave the answer. In regulated domains, "the AI said so" is not a defense that has ever worked.

## What does training-data quality cost?

**[MODELED — not reported]**

```
ASSUMPTIONS
Technical questions answered
  from model knowledge:             200 / year
Rate where popular answer ≠
  authoritative answer:             7%  (~14 / year)
Rate caught before acting:          65%
Bad guidance delivered:             ~5 / year
Cost per instance:                  $3,000 – $40,000
  (correction, remediation,
   penalty exposure in regulated cases)
```

**Annualized exposure: ~$15,000 – $200,000**

## How do you control for training-data quality?

Route authoritative questions to authoritative sources. The model is a starting point for *finding*  the governing text, never a substitute for reading it. Maintain the Authoritative Source Register and require primary-source confirmation for anything that carries regulatory, contractual, or financial consequence.

```
CONTROL COST
Authoritative questions:    200 / year
Primary-source check:       12 minutes each
Annual:                     40 hours
Fully loaded rate:          $75 / hour
```

**Annualized control cost: $3,000**

Judgment Multiple (IgnatiusTheYoungerAI, 2026) — modeled  ~5x to ~67x

## What should you do this week?

**RECOMMENDATION**

Identify the three questions in your function where **the popular answer and the correct answer differ.**  Every experienced practitioner knows what these are; almost none have written them down.

Write them down. That document is the highest-density expertise artifact you can produce, it makes you the person others check with, and it is directly reusable as a standing context brief for AI-assisted work.

## Evidence

**RESEARCH**  Corpus filtering targeting quality proxies rather than factual verification.

**RESEARCH**  Models reproducing common misconceptions.

**ANALYSIS**  The frequency-versus-truth argument and the commercial-incentive aggravator are the author's.

The Full System

This is one of 24 failure modes. The book gives you all of them, plus the controls that catch each one and a 90-day plan to prove you ran them.

[Get the Book · $29](https://agrb01-w9.myshopify.com/cart/43747019161690:1)

[← 14 Uncertainty Miscalibration](/failure-modes/14)
[16 Lack Of Accountability →](/failure-modes/16)
