Mynd Healthcare

Home / Scientific Methods

External and Prospective Evaluation

External evaluation tests transfer beyond development data, while prospective evaluation plans observation forward in time.

#Understand the two separate dimensions

External evaluation assesses a method using data outside its development process. The difference may involve another location, population, organization or time period. Describe that difference precisely. A random split of one development dataset is usually an internal evaluation, even when the held-out portion was not used for training.

Prospective evaluation starts with a plan and observes relevant data or outcomes forward in time. It can occur in the same setting where a method was developed. An external evaluation can use previously collected records. External and prospective therefore describe different features, and neither term guarantees the other.

#Protect the evaluation and define independence

Specify the method, version and decision thresholds before examining evaluation results where possible. Document any overlap in participants, sites or data sources with development. Repeatedly checking results and adjusting the method turns evaluation data into development information, so a later untouched evaluation may be needed for an unbiased performance estimate.

Data independence and investigator independence are also different. A development team may evaluate genuinely separate data, while a separate team may still use overlapping records. Explain who controlled the data, made analysis decisions and interpreted findings. Clear reporting is more useful than an unsupported claim that a study was independent.

#Evaluate the intended conditions of use

Prospective observation can reveal missing inputs, workflow delays and changes in data quality that historical records do not capture well. However, observing a model without showing its output to clinicians does not establish the consequences of using it. Study design should distinguish performance assessment from evaluation of effects on care.

Compare performance across relevant settings and groups, with uncertainty clearly reported. If local adjustment is needed, separate evidence before and after adjustment and explain how new settings will be assessed. Successful evaluation in one additional setting strengthens evidence, but it does not establish that performance will remain stable everywhere or indefinitely.

#Common misunderstandings

External and prospective evaluation answer different questions; neither label alone establishes that a clinical tool is ready for routine care. An external study tests performance beyond the development data, but the new setting may still closely resemble the original one. It does not automatically show that results will transfer to every hospital, population or workflow.

Prospective evaluation follows a plan forward in time, but it is not necessarily external, independent or a test of patient benefit. A tool can be assessed prospectively at the same organisation that developed it. Researchers may also record its predictions without letting them influence care.

Strong average performance can hide important differences between patient groups or clinical settings. Accuracy alone does not establish whether using the tool improves decisions, avoids harm or makes care more accessible. These questions require evidence matched to the intended role, including how clinicians respond to outputs and what happens when outputs are wrong.

#Questions worth asking a clinician

  • Was the model evaluated on external data, prospectively collected data, or both, and what makes each label appropriate?
  • How did evaluation data differ from development data in patients, hospitals, time periods, equipment, or clinical workflows?
  • Were evaluation data independent of development data, and were the investigators independent of the team that developed the model?
  • For prospective evaluation, were the model, performance measures, and analysis plan fixed before patient observation began?
  • Did the prospective study only measure predictive performance, or did it test whether using the model improved clinical decisions or patient outcomes?