Evidence-based medicine: PICO questions and levels of evidence
What evidence-based medicine is, how to write a PICO question with examples, how levels of evidence work (Oxford CEBM and GRADE) and how to use AI without losing rigor.
Evidence-based medicine (EBM) combines the best available research with clinical expertise and patient values to make care decisions. In practice it follows five steps: ask a focused PICO question, acquire the evidence, appraise it by level of evidence, apply it to the individual patient and assess the result. AI can speed up searching, not judgment.
Editorial update: October 10, 2026. Published by Arkangel AI. This content is educational and does not replace clinical judgment.
What is evidence-based medicine?
The most cited definition comes from David Sackett and colleagues: "the conscientious, explicit, and judicious use of current best evidence in making decisions about the care of individual patients" (Sackett et al., BMJ 1996). The same paper spells out what EBM is not: it is not "cookbook" medicine, and it is not restricted to randomized trials. External evidence informs clinical expertise and patient preferences; it never replaces them.
That three-part model still holds:
- Best available evidence, with its level of certainty.
- Clinical expertise, to judge whether the evidence applies to this patient.
- Patient values and circumstances, including cost, access and preferences.
The 5 steps of evidence-based practice (the 5 A's)
- Ask: turn a clinical uncertainty into an answerable question (PICO).
- Acquire: find the best evidence in appropriate sources.
- Appraise: judge validity, effect size and applicability.
- Apply: integrate the evidence with context and shared decision-making.
- Assess: measure the outcome and adjust.
Step 1 determines the quality of everything that follows. A vague question produces a vague search.
What is a PICO question?
PICO is a format for structuring clinical questions described by Richardson and colleagues in 1995 (ACP Journal Club). Each letter is a component:
| Component | Meaning | Guiding question |
|---|---|---|
| P | Patient, population or problem | Who does this apply to? |
| I | Intervention or exposure | What are you considering doing? |
| C | Comparison | Compared with what? |
| O | Outcome | Which result matters? |
Common variants: PICOT adds time, PICOS adds study design and PECO uses "exposure" for etiology or harm questions.
PICO question example
Starting doubt: "Are the newer anticoagulants better than warfarin?"
PICO breakdown:
- P: adults with non-valvular atrial fibrillation.
- I: direct oral anticoagulants.
- C: warfarin.
- O: stroke, systemic embolism and major bleeding.
Final wording: "In adults with non-valvular atrial fibrillation, do direct oral anticoagulants, compared with warfarin, reduce stroke and systemic embolism without increasing major bleeding?"
The second version already suggests search terms, the ideal study design and the outcomes to look for in each paper.
Which study design answers which question
| Question type | Best design |
|---|---|
| Therapy or prevention | Randomized controlled trial; systematic review of trials |
| Diagnosis | Cross-sectional study against a reference standard |
| Prognosis | Cohort study |
| Etiology or harm | Cohort or case-control study |
What are the levels of evidence?
Levels of evidence rank studies by their risk of bias. The familiar picture is the evidence pyramid: systematic reviews at the top, then randomized trials, cohort studies, case-control studies, case series and expert opinion. It is a useful starting point but oversimplifies. Two systems dominate today.
Oxford CEBM levels of evidence (2011)
The Oxford Centre for Evidence-Based Medicine assigns levels by question type. For treatment benefits, for example:
- Level 1: systematic review of randomized trials.
- Level 2: randomized trial, or observational study with a dramatic effect.
- Level 3: non-randomized controlled cohort or follow-up study.
- Level 4: case series, case-control or historically controlled studies.
- Level 5: mechanism-based reasoning.
The full table is published by the CEBM at the University of Oxford.
GRADE: certainty of evidence and strength of recommendations
GRADE rates the certainty of evidence for each outcome as high, moderate, low or very low (Guyatt et al., BMJ 2008). Randomized trials start at high and observational studies at low. Certainty is then rated down for risk of bias, inconsistency, indirectness, imprecision or publication bias, and can be rated up for a large effect, a dose-response gradient or plausible confounding that would reduce the observed effect.
GRADE separately rates the strength of a recommendation (strong or conditional), weighing benefits and harms, values, cost and feasibility. It is used by the WHO, Cochrane and many guideline bodies, including several U.S. specialty societies; the GRADE Working Group maintains the method.
Level of evidence is not the same as grade of recommendation
A strong recommendation can rest on moderate-certainty evidence when benefits are clear and harms minimal. Conversely, high-certainty evidence can lead to a conditional recommendation when the benefit is small or the cost high. When you read a guideline, check both.
Where to find the evidence
- Clinical practice guidelines: U.S. specialty societies, the USPSTF for prevention and national bodies elsewhere. See AI for clinical guidelines.
- Systematic reviews: the Cochrane Library and Epistemonikos.
- Primary studies: PubMed with study-type filters.
- Synthesized resources: point-of-care references and AI search engines that cite their sources.
How AI helps with evidence-based medicine
AI is useful for steps 1 to 3 when it is used with method:
- From case to question: it can help rephrase a clinical doubt in PICO format.
- Search: a medical AI search engine can scan hundreds of abstracts and return a cited synthesis in seconds.
- First-pass appraisal: it can surface the design, sample size and outcomes of each cited study.
It also has limits. General-purpose language models can fabricate references, and an automated summary does not replace a risk-of-bias assessment. Always ask for the source, open the study and check its level of evidence before applying it.
Arkangel AI answers clinical questions with cited studies and guidelines so every statement can be checked. In a peer-reviewed randomized trial with 83 medical students, those using Arkangel AI took ~55% less time per case and scored higher on answer validity (Intelligence-Based Medicine, 2026); the study was run by Arkangel's own team. See the workflow in evidence-based medicine AI or read our guide to AI in healthcare.
Frequently asked questions
What is a PICO question?
A PICO question turns a clinical doubt into a question you can search and answer. It defines four elements: patient or problem (P), intervention (I), comparison (C) and outcome (O). It helps you pick search terms, the right study design and the outcomes worth appraising, which makes searches faster and more reproducible.
What are the levels of evidence?
Levels of evidence rank studies by their risk of bias. On the Oxford CEBM scale for treatment, level 1 is a systematic review of randomized trials and level 5 is mechanism-based reasoning. GRADE instead rates the certainty of evidence for each outcome as high, moderate, low or very low.
What is the difference between level of evidence and strength of recommendation?
The level or certainty of evidence describes how confident we are in a result. The strength of a recommendation says how firmly an action is advised, also weighing benefits, harms, cost and patient values. That is why a strong recommendation can rest on moderate evidence, and high-certainty evidence can still yield a conditional one.
Can AI replace a systematic literature search?
No. An AI search engine speeds up exploratory searching and summarizes studies with citations, but a systematic review requires a protocol, reproducible searches across several databases, dual screening and formal risk-of-bias assessment. Use AI to get oriented quickly, and always verify the original sources before making a clinical decision.
Sources
- Sackett DL et al. Evidence based medicine: what it is and what it isn't. BMJ 1996.
- Richardson WS et al. The well-built clinical question. ACP J Club 1995.
- Guyatt GH et al. GRADE: an emerging consensus. BMJ 2008.
- Oxford Centre for Evidence-Based Medicine. OCEBM Levels of Evidence.
- Arkangel AI. Randomized trial with 83 students, Intelligence-Based Medicine 2026 (authors are Arkangel staff).