Hallucination

A model stating something false with confidence — a predictable consequence of how models are trained and graded.

Producing a valid statement is strictly harder than judging whether one is valid, so a model that makes errors distinguishing true from false must make errors generating text. Standard pretraining also produces calibrated models, and a calibrated model necessarily emits some falsehoods.

It persists after training because almost every benchmark grades in binary, where “I don’t know” scores the same as a wrong answer. We are, in effect, training models to bluff because bluffing scores better.