Mynd Healthcare

Home / Health Data & AI

Calibration and Decision Thresholds

Calibration concerns the reliability of predicted probabilities, while decision thresholds determine when an output leads to a particular action.

#When probabilities are reliable

Calibration describes how closely predicted probabilities agree with observed outcome frequencies across groups of predictions. If a model assigns similar risks to many people, the event frequency in that group should be reasonably close to those risks. This is a property assessed across cases, not proof about one person's future.

A model can rank people well while consistently overstating or understating risk. That means good discrimination does not guarantee good calibration. Calibration also depends on the outcome definition and time period. A probability of an event within one month cannot be interpreted as the probability of that event within a year.

#Choosing when to act

A decision threshold is a cut-off used to connect a score or probability with an action, such as further review. Lowering a positive threshold generally identifies more cases but also produces more false positives. Raising it generally reduces false positives while increasing the chance of missing cases.

There is no universally correct threshold. The choice depends on the action, the consequences of errors, available resources and the evidence supporting the care pathway. A threshold used to prompt a low-burden review may not be suitable for an invasive procedure. Operational convenience alone does not establish an acceptable clinical trade-off.

#Checking reliability over time

Calibration can be examined by comparing predicted and observed risks across ranges of predictions. These comparisons need enough outcomes and appropriate follow-up. Small samples can make estimates unstable, and good average calibration can conceal problems in particular groups or at the highest and lowest predicted risks.

A model may need reassessment when event rates or care practices change. Recalibration adjusts the mapping from model outputs to probabilities, but it cannot fix every underlying problem. Changes to probabilities or thresholds should be evaluated because they can alter who receives attention, how workload is distributed and which errors occur.

#Common misunderstandings

A well-calibrated prediction is not a guarantee about what will happen to one person. If a model assigns a 20% risk to many people, good calibration means that roughly 20% of that group experience the outcome within the stated time period. It does not identify which individuals will be affected.

Calibration is also different from the ability to separate people at higher and lower risk. A model can rank people reasonably well while consistently giving probabilities that are too high or too low.

A decision threshold is not a boundary between “safe” and “unsafe”. It is a chosen point for considering an action, based on the expected benefits, harms and practical consequences. Different actions may justify different thresholds.

Finally, strong results in one setting do not establish reliability everywhere. Changes in the population, clinical practice or how information is recorded can affect predictions and warrant further checks.

#Questions worth asking a clinician

  • If this tool predicts a 20% risk, do about 20 in 100 similar patients actually develop the condition?
  • Has this tool’s calibration been checked in patients with health conditions and backgrounds like mine?
  • What risk threshold would lead you to recommend testing or treatment, and why was that threshold chosen?
  • How does the threshold balance missed diagnoses against unnecessary tests or treatments?
  • Could my preferences and tolerance for risk change the threshold you use for my care?