Evaluating Evidence
How to assess the quality and reliability of evidence used in arguments — from scientific studies to statistics, testimony, and anecdotes.
For the complete documentation index, see llms.txt.Skip to main content
Not all evidence is created equal. In science and medicine, evidence is ranked by how reliably it establishes causal relationships. At the top are systematic reviews and meta-analyses that combine results from multiple studies. Next come randomized controlled trials. Below those are observational studies, then case reports, and at the bottom, expert opinion and anecdote.
This hierarchy applies beyond medicine. In policy debates, a well-designed study of a program's effects across multiple cities is stronger evidence than a single success story. A statistic from a government statistical agency is more reliable than a number cited by an advocacy group. A pattern that holds across different countries and time periods is more convincing than a single data point.
Who funded it? Industry-funded studies are not automatically wrong, but they consistently show a funding effect — studies funded by drug companies are more likely to find positive results for the funder's drug. Always check the disclosure statement.
How large was the sample? Small studies (under 100 participants) are more likely to produce extreme results by chance. Be skeptical of dramatic findings from small samples.
Was it peer-reviewed? Peer review is imperfect but provides a baseline quality check. Preprints (not yet peer-reviewed) should be treated as preliminary.
Has it been replicated? A single study, no matter how well-designed, is a starting point. The replication crisis across psychology, medicine, and social science has shown that many published findings do not hold up when other researchers repeat the experiment.
Suppose you see the headline: 'Coffee prevents cancer, new research finds.' Run the checklist before believing it.
Step 1 — What kind of study is it? The article describes an observational cohort that followed coffee drinkers and non-drinkers. Observational studies can show association but cannot, on their own, establish that coffee causes lower cancer rates.
Step 2 — How big and how long? A cohort of 500,000 people tracked for 10 years is far more credible than 200 people surveyed once.
Step 3 — Who funded it, and was it replicated? If a coffee trade association funded a single unreplicated study, treat the headline as a hypothesis, not a fact.
Step 4 — What is the alternative explanation? Coffee drinkers may differ from non-drinkers in income, exercise, or smoking — confounders that could drive the result. A responsible study statistically adjusts for these; a headline rarely mentions them.
The table below maps each tier of the hierarchy to how much weight it should carry:
| Evidence type | Reliability | Why |
|---|---|---|
| Systematic review / meta-analysis | Highest | Pools many studies, averages out chance and single-study bias |
| Randomized controlled trial | High | Randomization isolates cause from confounders |
| Cohort / observational study | Moderate | Shows association; vulnerable to confounding and reverse causation |
| Case report / case series | Low | Vivid but has no comparison group |
| Expert opinion / anecdote | Lowest | No systematic data behind it |
Applied to the coffee headline, a single moderate-reliability cohort simply cannot support the word 'prevents.' The honest verdict is: interesting association, causation unproven.