
Estimate calibration is the practice of recording predictions made inside a business — delivery dates, effort, cost, close probability — and scoring them against what actually happened, so the organisation learns the systematic bias in its own forecasts. Companies measure whether work was delivered. They almost never measure whether the estimate that preceded it was any good.
Bias is stable, which makes it correctable
Individual estimation error is remarkably consistent. People are not randomly wrong — they are reliably wrong in a particular direction, by a roughly stable factor, for a particular class of work. An estimator with a consistent 1.6x factor is more useful than one whose error is unbiased but wildly variable — the first can be corrected, the second cannot.
What gets recorded
Four fields at prediction time: the claim (a specific, falsifiable statement), confidence (stated at the time — never reconstructed later), basis (comparable prior work, calculation, or judgement), and resolution date. Then one field at resolution: what actually happened. Confidence must be captured before the outcome and cannot be edited afterwards — hindsight adjustment is the single failure mode that destroys the whole exercise.
Where AI helps
Manual calibration ledgers are almost all abandoned within a quarter. The automation does three things: extraction (identifies falsifiable predictions in mail, proposals, standups, tickets — no form-filling required), resolution (closes predictions automatically when outcomes become visible), and surfacing at the right moment (calibration data appears while a new estimate is being made, not in a quarterly report nobody reads).
Making it safe
It measures estimates, not effort or output. It corrects rather than judges and leadership goes first, if executive forecasts are exempt, nobody below will trust the exercise.

