Calibration is a promise: what 80% should mean
A number you cannot act on is decoration. When a reading says 80%, it should be wrong about one time in five, no more and no less. That is the difference between a score and a guess with a decimal point.
When a reading says 80%, it is right about 80% of the time. That is the whole promise.
Confidence without calibration is theatre
Many systems report a probability that has no relationship to how often they are right. The number looks precise and means nothing, which is worse than no number at all, because it invites action.
Calibration is the property that ties the reported probability to observed frequency. It is measurable, and it is the first thing we measure.
How we hold the line
Readings are checked against held out outcomes and summarised with expected calibration error. When the curve drifts, the mapping from evidence to probability is corrected rather than the headline accuracy chased.
The reasoning a model produces is kept separate from the probability it reports, so a fluent explanation cannot quietly push a number up.
Why it matters for review
Calibration is what lets a team triage. The confident claims can be skimmed, the uncertain ones get a person, and the split is trustworthy because the number behind it means what it says.
Takeaways
- Calibration ties the reported probability to how often it is right.
- We report expected calibration error, not just accuracy.
- Reasoning is kept separate from the probability.