methodProbability and statistics

Proper scoring rules and multiplicity

Score probabilistic forecasts honestly while controlling the many-comparisons problem created by large rule libraries.

Evidence status

Cites 1 source, none of which has been read

The source is named, but it has been retrieved and read as part of building this page. Nothing here has been matched to a passage, so the citations show where a reader might look rather than what was checked.

Rely on this page for

Orientation: how the topic is organised, which terms matter, and where to start reading.

Do not rely on it for

A claim you intend to act on or repeat. Follow the cited material yourself first.

Working definition

A proper scoring rule gives an expected optimum when the forecaster reports its true probability distribution. Forecast evaluation must also account for multiplicity: searching many planets, windows, outcomes, and subgroups inflates false discoveries unless the analysis plan, correction, or held-out evaluation is fixed in advance.

Notation

Brier = (p − y)²log score = −log p(y)

Assumptions

  • Forecast probabilities and outcomes are valid.
  • The score and baseline are selected before outcomes.
  • The family of tested hypotheses is declared.

Invariants

  • Proper scores reward honest probabilities in expectation.
  • Lower Brier and log loss are better under their standard definitions.
  • Adding undisclosed tests changes the error budget.

Reproducible procedure

  • Lock forecast, outcome, horizon, score, and baseline.
  • Compute paired score differences.
  • Report uncertainty, multiplicity controls, and all registered analyses.

Error and boundary controls

  • Rare outcomes require large samples.
  • Log loss is sensitive to overconfident errors.
  • Repeated observations may violate independence.

What this does not establish

A score measures performance on a specified task; it does not prove a universal mechanism or justify claims beyond the registered population and horizon.

Explicit applications

1 cross-domain bridges

Empirical validationempirical test

Prospective forecast scoring

Compare locked probabilistic forecasts with outcomes using a preselected proper score and paired baselines.

Inputs

  • locked probabilities
  • resolved outcomes
  • baseline forecasts

Outputs

  • Brier or log score
  • baseline delta
  • confidence or credible interval

Transformation: Compute paired score differences with multiplicity and dependence controls.

Limit: Results apply only to the registered task, population, and horizon.

Open connected system →

Authoritative references

  1. [1]NIST/SEMATECH e-Handbook of Statistical Methods · National Institute of Standards and Technology

    Establishes: Methods for uncertainty analysis, calibration, time-series modeling, process monitoring, experimental design, reliability, and statistical comparison.

    Boundary: Statistical procedures quantify evidence under a design and model; they do not repair biased sampling, outcome leakage, post-hoc hypotheses, or unmeasured confounding.

Direct answer

  • A proper scoring rule gives an expected optimum when the forecaster reports its true probability distribution. Forecast evaluation must also account for multiplicity: searching many planets, windows, outcomes, and subgroups inflates false discoveries unless the analysis plan, correction, or held-out evaluation is fixed in advance.

Mechanism and method

  • Lock forecast, outcome, horizon, score, and baseline.
  • Compute paired score differences.
  • Report uncertainty, multiplicity controls, and all registered analyses.

What is measured

  • Proper scores reward honest probabilities in expectation.
  • Lower Brier and log loss are better under their standard definitions.
  • Adding undisclosed tests changes the error budget.

Limitations

  • Rare outcomes require large samples.
  • Log loss is sensitive to overconfident errors.
  • Repeated observations may violate independence.
  • Forecast probabilities and outcomes are valid.
  • The score and baseline are selected before outcomes.
  • The family of tested hypotheses is declared.

What this does not establish

  • A score measures performance on a specified task; it does not prove a universal mechanism or justify claims beyond the registered population and horizon.

Bridge: Prospective forecast scoring

  • Compare locked probabilistic forecasts with outcomes using a preselected proper score and paired baselines.
  • Input: locked probabilities
  • Input: resolved outcomes
  • Input: baseline forecasts
  • Output: Brier or log score
  • Output: baseline delta
  • Output: confidence or credible interval
  • Limit: Results apply only to the registered task, population, and horizon.

Related records

Related mathematical concepts