methodProbability and statistics

Calibration and reliability

Test whether stated probabilities agree with observed frequencies and remain stable across relevant groups and time.

Evidence status

Cites 1 source, none of which has been read

The source is named, but it has been retrieved and read as part of building this page. Nothing here has been matched to a passage, so the citations show where a reader might look rather than what was checked.

Rely on this page for

Orientation: how the topic is organised, which terms matter, and where to start reading.

Do not rely on it for

A claim you intend to act on or repeat. Follow the cited material yourself first.

Working definition

A probabilistic forecaster is calibrated when events assigned probability p occur at approximately frequency p over an appropriate reference class. Reliability diagrams and calibration error summarize agreement, but calibration must be evaluated with discrimination, sample size, dependence, and subgroup stability.

Notation

P(Y=1 | p̂=p) ≈ pcalibration error = observed frequency − forecast probability

Assumptions

  • Forecasts are locked before outcomes.
  • Outcome definitions and horizons are stable.
  • Reference classes are large enough to estimate frequencies.

Invariants

  • A constant base-rate forecast can be calibrated but uninformative.
  • Calibration depends on the evaluated population.
  • Retrospective relabeling invalidates the test.

Reproducible procedure

  • Bin or smooth locked forecasts without viewing outcomes during design.
  • Compare predicted and observed frequencies with uncertainty.
  • Assess discrimination, sharpness, and subgroup drift.

Error and boundary controls

  • Small bins create noisy estimates.
  • Adaptive binning can bias summaries.
  • Non-stationarity can make historical calibration stale.

What this does not establish

Calibration alone does not show useful skill over a baseline, causation, or transportability to a new decision context.

Explicit applications

2 cross-domain bridges

Semiconductor manufacturingmeasurement

Yield learning and process control

Compare measured defect and yield behavior with stable process limits and calibrated metrology.

Inputs

  • lot measurements
  • control limits
  • tool and recipe identifiers

Outputs

  • control signals
  • calibration status
  • subgroup diagnostics

Transformation: Estimate reliability and monitor departures from the qualified process distribution.

Limit: Control limits detect distributional change; they do not identify the physical root cause.

Open connected system →
Empirical validationempirical test

Probability calibration audit

Test whether events forecast at a stated probability occur at that frequency prospectively.

Inputs

  • locked forecasts
  • binary outcomes
  • forecast strata

Outputs

  • calibration curve
  • calibration error
  • subgroup stability

Transformation: Estimate reliability curves and uncertainty alongside discrimination.

Limit: A base-rate forecaster may be calibrated without adding useful discrimination.

Open connected system →

Authoritative references

  1. [1]NIST/SEMATECH e-Handbook of Statistical Methods · National Institute of Standards and Technology

    Establishes: Methods for uncertainty analysis, calibration, time-series modeling, process monitoring, experimental design, reliability, and statistical comparison.

    Boundary: Statistical procedures quantify evidence under a design and model; they do not repair biased sampling, outcome leakage, post-hoc hypotheses, or unmeasured confounding.

Direct answer

  • A probabilistic forecaster is calibrated when events assigned probability p occur at approximately frequency p over an appropriate reference class. Reliability diagrams and calibration error summarize agreement, but calibration must be evaluated with discrimination, sample size, dependence, and subgroup stability.

Mechanism and method

  • Bin or smooth locked forecasts without viewing outcomes during design.
  • Compare predicted and observed frequencies with uncertainty.
  • Assess discrimination, sharpness, and subgroup drift.

What is measured

  • A constant base-rate forecast can be calibrated but uninformative.
  • Calibration depends on the evaluated population.
  • Retrospective relabeling invalidates the test.

Limitations

  • Small bins create noisy estimates.
  • Adaptive binning can bias summaries.
  • Non-stationarity can make historical calibration stale.
  • Forecasts are locked before outcomes.
  • Outcome definitions and horizons are stable.
  • Reference classes are large enough to estimate frequencies.

What this does not establish

  • Calibration alone does not show useful skill over a baseline, causation, or transportability to a new decision context.

Bridge: Yield learning and process control

  • Compare measured defect and yield behavior with stable process limits and calibrated metrology.
  • Input: lot measurements
  • Input: control limits
  • Input: tool and recipe identifiers
  • Output: control signals
  • Output: calibration status
  • Output: subgroup diagnostics
  • Limit: Control limits detect distributional change; they do not identify the physical root cause.

Bridge: Probability calibration audit

  • Test whether events forecast at a stated probability occur at that frequency prospectively.
  • Input: locked forecasts
  • Input: binary outcomes
  • Input: forecast strata
  • Output: calibration curve
  • Output: calibration error
  • Output: subgroup stability
  • Limit: A base-rate forecaster may be calibrated without adding useful discrimination.

Related records

Related mathematical concepts