Explainability and Its Limits
Explanations can help people inspect model behaviour, but they do not establish that an output is correct, causal or clinically useful.
#Different kinds of explanation
Explainability concerns ways to make a model's behaviour more understandable. Some models use relatively simple relationships that can be inspected directly. Other methods add an explanation after a prediction is produced. These may highlight influential inputs, mark regions of an image or show how changing an input affects an output.
An explanation may describe the model overall or focus on one result. These are different tasks. A factor that matters across a population may not explain a particular case. Explanations also need an audience and purpose: a developer investigating errors may need different information from a clinician reviewing an alert.
#Why a plausible story is not proof
An explanation can look convincing without faithfully representing the model's actual reasoning. Some explanation methods approximate complex behaviour, and different methods may produce different accounts of the same output. Highlighted image regions or ranked variables can therefore be useful clues, but should not automatically be treated as definitive evidence.
Importance is not the same as causation. A model may use a variable because it is associated with an outcome, even when changing that variable would not change the outcome. Inputs that are closely related can complicate attribution further. A suggested change may also be impossible or inappropriate in real life.
#Using explanations responsibly
Explanations can support error investigation, reveal unexpected dependencies and help users ask better questions. Their usefulness should be tested rather than assumed. Relevant questions include whether they reflect model behaviour, remain reasonably stable and help the intended audience identify mistakes instead of merely increasing confidence.
An understandable output can still be wrong, and a difficult-to-explain model can still perform well on a defined task. Neither observation settles whether use is appropriate. Evaluation also needs reliable performance evidence, information about limitations and a clear care process. Explanations complement these safeguards; they do not replace them.
#Common misunderstandings
A clear explanation is not the same as a correct result. A model can give a confident, readable account of an answer that is wrong. The explanation itself may also be incomplete or inaccurate.
Highlighting an input does not show that it caused the outcome. For example, a feature linked with illness in training data might reflect patterns in testing or record keeping rather than a biological relationship. Its apparent importance may change when other inputs change.
Another misunderstanding is that explanations reveal everything the model considered. Some methods provide only an approximation of its behaviour; generated descriptions may not faithfully represent how an output was produced.
Finally, explainability is not a substitute for clinical evaluation. An understandable tool can still perform poorly, miss important cases or work less well in some groups. Evidence about accuracy, safety and effects on care must be assessed separately from how convincing its explanations sound.
#Questions worth asking a clinician
- Does this explanation show how the model actually reached its answer, or just offer a plausible story?
- What independent evidence supports the model’s output for someone with my medical history?
- Does a highlighted factor actually cause the predicted outcome, or is it merely associated with it?
- Could small changes in my information produce a different answer or explanation, and how would you check?
- Has using this model improved patient outcomes compared with usual care, rather than simply making its recommendations easier to understand?