Mynd Healthcare

Home / Health Data & AI

Labels and Ground Truth

Reference labels provide targets for training and evaluation, but their meaning and reliability depend on how they are created.

#What a label represents

In many AI projects, a label is the answer a model is trained to predict or compared against during evaluation. Labels might describe a finding on an image, a recorded diagnosis or an event during follow-up. They define the task, so an unclear label can make the whole project unclear.

The term ground truth is often used for these reference answers. It can suggest more certainty than the evidence supports. Some outcomes are directly measured, while others depend on interpretation, incomplete records or a chosen definition. A reference standard is best understood as the basis for comparison, not automatic proof.

#How reference answers are created

Labels may come from clinical records, laboratory results, expert review or patient reports. Reviewers may work independently and use written criteria. Disagreements can be discussed or referred to another reviewer. These processes can improve consistency, but agreement does not guarantee correctness if everyone shares the same mistaken assumption.

Timing matters too. A diagnosis confirmed later may be unavailable when a prediction would actually be made. An outcome may appear absent because follow-up ended or care continued elsewhere. Labels should therefore specify what counts as an event, the relevant time period and how uncertain or unavailable outcomes are handled.

#Understanding label limitations

Convenient labels may be proxies for the question of real interest. A billing code does not necessarily capture disease severity. Recorded treatment may reflect access to care as well as clinical need. A model trained on a proxy can reproduce those influences instead of measuring the intended health condition.

Useful reporting explains label sources, reviewer procedures and known uncertainty. Checking a sample against a stronger reference can reveal problems, although that reference may also have limits. When labels are imperfect, performance results describe agreement with those labels and should not be presented as certainty about clinical reality.

#Common misunderstandings

“Ground truth” can sound like a final, unquestionable answer. In health data, it usually means the reference chosen for a particular task. That reference might come from a clinician’s assessment, a laboratory result, a medical record, or several reviewers working together. Each source answers different questions and can have limitations.

A label is not always the same as a confirmed diagnosis. For example, a diagnosis code may reflect what was recorded during a visit rather than an independently verified condition. Likewise, the absence of a recorded diagnosis does not necessarily mean a condition was absent.

Agreement between reviewers can support confidence, but it does not prove that a label is correct. Reviewers may share assumptions or lack the same information. A model’s agreement with reference labels therefore measures performance against that reference—not automatically clinical usefulness, safety, or accuracy in every patient group. Those require additional evaluation in relevant care settings.

#Questions worth asking a clinician

  • How were the reference labels created, and what clinical outcome or judgment was the model trained and evaluated to match?
  • How were uncertain cases, missing evidence, and disagreements between clinicians handled when assigning reference labels?
  • If labels used billing codes or treatment decisions as proxies, how did you check for effects of access, documentation, and care practices?
  • How did you check reference labels for errors, and how might those errors affect reported model performance?
  • Beyond agreement with reference labels, what independent clinical evidence supports the correctness of the model’s predictions?