Home / Evidence & Research Literacy
When Studies Disagree
Apparently conflicting studies may differ in their questions, methods or precision, so disagreement requires investigation rather than a simple vote count.
#Check whether the questions match
Studies can appear to disagree while examining different things. Participants may differ in age, illness severity, previous treatment or other relevant characteristics. An intervention may have different effects across these groups, although an apparent subgroup difference needs careful testing rather than an explanation invented after seeing the results.
Details of the intervention and comparison also matter. Dose, duration, support, adherence and usual care can vary between settings. A study measuring short-term symptoms is not answering the same question as one measuring long-term hospital admissions, even when both investigate an intervention with the same name.
#Compare estimates, not labels
One paper may report a statistically significant result while another does not, yet their effect estimates may be similar. Different sample sizes and amounts of variation can produce different levels of precision. The labels 'positive' and 'negative' can therefore exaggerate a disagreement that the numbers do not support.
Compare effect sizes and their uncertainty, using compatible measures and time periods. Overlapping confidence intervals do not by themselves prove that effects are identical, and testing whether effects differ requires an appropriate analysis. Chance remains a possible explanation, especially when studies are small or examine many outcomes.
#Examine methods and the wider pattern
Differences in recruitment, measurement, missing data and analysis can produce different results. Observational studies may also account for confounding in different ways. A larger or newer paper is not automatically more reliable; the relevance and quality of its design need to be assessed directly.
#Avoid choosing a convenient answer
A systematic review can help organise these differences and assess whether combining results is appropriate. Counting how many studies favour each conclusion is usually misleading because studies differ in size, precision and risk of bias. Several weak studies need not outweigh one strong, relevant comparison.
Sometimes disagreement remains unresolved. The honest conclusion may be that effects vary by setting or that the evidence is not yet precise enough. Further research is most useful when it addresses the specific reasons for uncertainty rather than merely repeating a poorly defined question.
#Common misunderstandings
Two studies can reach different conclusions without either being dishonest or useless. They may examine different populations, compare different treatments, or measure outcomes at different times. Even similar studies can produce different estimates through chance, especially when the estimates are imprecise.
A common mistake is to treat “statistically significant” and “not statistically significant” as opposite findings. Those labels alone do not establish that the results differ meaningfully. The estimated effects and their uncertainty may be quite similar. Likewise, finding no clear evidence of benefit does not necessarily establish that there is no benefit.
Counting studies is not a reliable way to settle disagreement. Several small studies with similar weaknesses do not automatically outweigh a more rigorous study. Nor does the newest study necessarily replace everything published before it. The aim is to understand what explains the differences and what the evidence, considered together, can reasonably support.
#Questions worth asking a clinician
- Did these studies examine comparable populations, interventions, comparison groups and outcomes, or were they answering different questions?
- Are the estimated effects and confidence intervals similar despite one study reporting statistical significance and another not?
- Which studies have lower risks of bias, and how should that affect our interpretation rather than counting positive and negative findings?
- Could differences in treatment dose, follow-up time or outcome measurement explain why the studies appear to disagree?
- What disagreement remains after accounting for differences in methods and precision, and what further research would help resolve it?