What Is a Systematic Review? A Complete Guide
A systematic review is a type of literature review that uses a predefined, documented, and reproducible method to identify, evaluate, and synthesize all available research evidence relevant to a specific research question. Unlike a general literature review, it follows a strict protocol — usually registered before the review begins — that sets out the search strategy, the databases to be searched, and the criteria for including or excluding studies. Because the method is documented in advance, another researcher could in principle repeat the same steps and arrive at a similar set of included studies.
A systematic review is a research method that identifies, evaluates, and summarizes all available evidence on a specific question using predefined, documented methods. It follows a systematic search across multiple databases, screens and appraises the studies found against set eligibility criteria, and synthesizes their findings — either narratively or, when the studies are similar enough, through statistical meta-analysis.
Systematic Review vs. Literature Review
A general literature review is typically narrative: the author selects sources based on their own reading and judgment, and organizes them into a discussion of what is known and unresolved in a field. There is no fixed protocol, and the exact set of sources included is shaped by the author's expertise and choices.
A systematic review, by contrast, follows a documented protocol that is usually finalized before the review begins. The protocol specifies the databases to be searched, the exact search terms, and the criteria a study must meet to be included or excluded. Two systematic reviews addressing the same question, run by different teams following the same protocol, should identify a similar (though rarely identical) set of studies — a level of reproducibility a narrative literature review doesn't aim for.
Systematic Review vs. Meta-Analysis
A systematic review and a meta-analysis are related but not the same thing. A systematic review is the overall process — the protocol, the search, the screening, the quality assessment, and the synthesis of what the included studies found.
A meta-analysis is a statistical technique that can be applied within a systematic review when the included studies are similar enough, in their design and reported outcomes, to be combined mathematically into a single pooled estimate of effect. Not every systematic review includes a meta-analysis — if the included studies use different outcome measures, populations, or designs, statistical pooling may not be appropriate, and the review reports its findings narratively instead.
PRISMA 2020: The Reporting Standard
PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) is a reporting standard, not a research method in itself — it doesn't tell you how to conduct a systematic review, but it specifies what needs to be reported once you have. The current version, PRISMA 2020, updated the original 2009 statement with a revised, expanded checklist and an updated flow diagram, reflecting how systematic review methodology had developed since 2009 — including more detailed guidance on reporting the search strategy, risk-of-bias assessment, and synthesis methods.
The PRISMA 2020 flow diagram is one of the most recognizable outputs of the standard: it tracks how many records were identified through database searching and other sources, how many were removed before screening (for example, duplicates), how many were screened, how many were excluded and why, how many full-text reports were assessed for eligibility, and finally how many studies were included in the review and, if applicable, in any quantitative synthesis. Reporting these numbers at every stage is what allows a reader to judge how thorough and unbiased the search and selection process was.
Most major journals in medicine and many in the social sciences now require or strongly recommend PRISMA 2020 compliance for systematic review submissions, and many journals ask authors to submit a completed PRISMA checklist alongside the manuscript.
Worked Example: A Hypothetical PRISMA Flow
The numbers below are an illustrative example only — not data from a real Tezyar project or a published study — meant to show how the PRISMA flow diagram's stages relate to each other in a typical review.
Records identified through database searching: 1,842 Duplicates removed: 412 Records screened (title and abstract): 1,430 Records excluded at this stage: 1,190 Full-text reports sought for retrieval: 240 Reports not retrieved: 8 Full-text reports assessed for eligibility: 232 Full-text reports excluded, with reasons (for example, wrong population, wrong outcome, or wrong study design): 187 Studies included in the review: 45
Reporting a specific reason for every exclusion at every stage is exactly what lets a reader judge how thorough and unbiased the search and selection process was.
Defining the Research Question
Before any searching begins, a systematic review starts with a clearly defined research question — a vague question produces an unfocused search and inconsistent screening decisions. In quantitative reviews of interventions, the PICO framework is the most common way to structure this: Population (who is being studied), Intervention (what is being done to or by them), Comparison (what it's being compared against, if anything), and Outcome (what is being measured). Some reviews extend this to PICOS, adding Study design as a fifth element to specify which types of studies will be considered.
For questions about diagnostic accuracy, prevalence, or qualitative evidence, other frameworks fit better — for example, a PIRD-style structure (Population, Index test, Reference test, Diagnosis of interest) for diagnostic test accuracy reviews, or a PICo structure (Population, Interest, Context) for some qualitative synthesis questions. Whichever framework fits the question, the goal is the same: a research question specific enough that a reader could predict what kind of studies would answer it, and specific enough to build a search strategy around.
Worked Example: Applying the PICO Framework
The example below is a hypothetical illustration of how PICO turns a broad interest into a focused research question — it is not a real study Tezyar has conducted or reviewed.
Population: adults hospitalized with type 2 diabetes. Intervention: a structured nurse-led discharge education program. Comparator: standard discharge instructions without structured education. Outcome: 30-day hospital readmission rate.
Example research question: in adults hospitalized with type 2 diabetes, does a structured nurse-led discharge education program, compared with standard discharge instructions, reduce the 30-day hospital readmission rate?
Registering the Protocol
A systematic review protocol is a pre-specified document describing the research question, planned search strategy, eligibility criteria, and analysis plan before the review actually begins. Registering it — most commonly through PROSPERO, an international registry for systematic review protocols — creates a public, time-stamped record of what the review team intended to do, which lets readers check whether the final published review changed its methods partway through (for example, quietly narrowing eligibility criteria after seeing which studies turned up).
Registration isn't universally mandatory, but many journals, funders, and ethics committees now expect it, particularly for reviews of health interventions. Some protocols are also published as standalone journal articles before the review itself is completed, which subjects the planned methodology to peer review before any results exist that could bias that assessment.
Designing the Search Strategy
A systematic search aims to be exhaustive and reproducible, not just convenient. That usually means searching multiple databases relevant to the field — for example, MEDLINE/PubMed, Embase, and CENTRAL for health topics, or subject-specific databases in other fields — rather than relying on a single source, since database coverage varies meaningfully.
Search terms are typically built around the concepts in the research question (for example, the P and I in PICO) and combined using Boolean operators (AND, OR, NOT), often mixing controlled vocabulary (like MeSH terms in PubMed) with free-text keywords to capture articles indexed differently. Many reviews also search beyond databases — checking the reference lists of included studies, searching trial registries, and searching grey literature such as conference abstracts or dissertations — to reduce the risk of missing relevant unpublished or hard-to-find studies.
The full search strategy for each database, including the exact search string used and the date it was run, is expected to be reported in the final review (typically as a supplementary file), so another researcher could rerun the same search and get a comparable set of results.
Setting Eligibility Criteria
Eligibility criteria define which studies from the search results actually qualify for inclusion, and they follow directly from the research question — usually specifying acceptable populations, interventions or exposures, comparators, outcomes, study designs, and sometimes constraints like publication language or date range. These criteria are meant to be fixed in the protocol before screening begins, precisely so decisions about which studies to include aren't influenced by what the results of those studies show.
In practice, criteria occasionally need adjusting once screening reveals the literature doesn't match initial assumptions — for example, finding almost no randomized trials on a topic and deciding to also include observational studies. When this happens, transparent reviews report the change and the reasoning behind it, rather than silently applying revised criteria as if they were the original plan.
Screening Studies for Inclusion
Screening is usually done in two stages. First, titles and abstracts from the search results are reviewed against the eligibility criteria to exclude clearly irrelevant records — a high-volume stage with a relatively fast per-record decision. Second, the full text of any records that pass the first stage is retrieved and assessed in detail, since abstracts alone often don't contain enough information to confirm eligibility.
Methodological guidelines generally recommend that screening at both stages be done independently by at least two reviewers, with disagreements resolved through discussion or a third reviewer, since a single screener's judgment calls are a documented source of error and bias in systematic reviews. The number of records excluded at each stage, along with the main reasons for full-text exclusions, is exactly the information the PRISMA flow diagram is designed to report.
Extracting Data from Included Studies
Once studies are selected, relevant data is extracted from each one into a structured, standardized format — typically a data extraction form or spreadsheet designed before extraction begins, so every study is captured consistently. Extracted data usually includes study characteristics (design, setting, sample size), participant characteristics, details of the intervention or exposure, outcome measures and how they were defined, and the actual results reported.
As with screening, extraction is ideally performed independently by two reviewers and compared, since transcription errors or differing interpretations of what a study reported are easy to introduce when only one person extracts the data. Many reviews pilot their extraction form on a handful of included studies first, refining it before extracting data from the full set.
Assessing Risk of Bias
Not all included studies are equally trustworthy, and a systematic review is expected to assess this explicitly rather than treating every study as equally reliable. The appropriate tool depends on study design: Cochrane's RoB 2 tool is used for randomized controlled trials, ROBINS-I is used for non-randomized studies of interventions, and discipline-specific tools — such as the Newcastle-Ottawa Scale for cohort and case-control studies, or JBI's critical appraisal checklists elsewhere — each examine different domains of potential bias relevant to that study design (for example, randomization, blinding, missing data, or selective reporting).
The results of this assessment matter beyond just describing study quality: they inform how much weight a study's findings are given in the synthesis, whether it's included in a meta-analysis at all, and how confidently the review's overall conclusions can be stated — a review built on high-risk-of-bias studies warrants more cautious conclusions than one built on low-risk ones, regardless of how many studies were included.
Synthesizing the Evidence
Once data is extracted and quality-assessed, it needs to be brought together into an answer to the original research question. When included studies are similar enough in population, intervention, and outcome measurement, this can take the form of a meta-analysis — combining results statistically into a pooled estimate. When they aren't similar enough, the synthesis is narrative instead: describing patterns, agreements, and contradictions across studies in prose, sometimes supported by structured summary tables, without calculating a single pooled number. This approach is formalized in guidance such as Synthesis Without Meta-analysis (SWiM), which sets out how to report a narrative synthesis with the same rigor and transparency expected of a meta-analysis.
Either way, the synthesis should account for the risk-of-bias assessment (for example, by examining results separately for low- versus high-risk studies) and should be explicit about the certainty of the evidence — not just what the studies found, but how much confidence that finding deserves given the quality and consistency of the underlying studies.
When Is a Meta-Analysis Appropriate?
A meta-analysis is only appropriate when the included studies are similar enough — clinically, methodologically, and statistically — that pooling their results produces a meaningful number rather than an average of genuinely different things. Clinical similarity means the studies address comparable populations, interventions, and outcomes; methodological similarity means comparable study designs and risk-of-bias profiles; statistical similarity (heterogeneity) is often assessed formally using measures like the I² statistic, which estimates what proportion of the variation across study results reflects real differences between studies rather than chance.
When heterogeneity is high and can't be explained (for example, through subgroup analysis), many methodologists recommend against pooling the studies into a single estimate at all, since the result can be misleading — an average that doesn't represent any of the actual studies well. In that situation, a narrative synthesis, or a meta-analysis restricted to a more homogeneous subset of studies, is usually the more defensible choice. Where pooling is appropriate, reviewers also choose between fixed-effect and random-effects models depending on whether they're willing to assume all studies estimate one true, identical effect, or whether the true effect is expected to vary across studies.
Common Mistakes in Systematic Reviews
A handful of methodological problems account for most of the criticism systematic reviews receive during peer review. Conducting the review without a registered protocol — or registering one but quietly deviating from it without disclosure — undermines the reproducibility the whole method is meant to provide. Screening or data extraction performed by a single reviewer, without independent duplication, removes a key safeguard against individual error and bias.
Other frequent issues include an incomplete or poorly documented search strategy that can't be reproduced or verified, searching only one database, excluding grey literature or non-English studies without justification (which can introduce publication or language bias), pooling statistically or clinically heterogeneous studies into a meta-analysis without acknowledging the heterogeneity, and omitting a risk-of-bias assessment altogether or conducting one but not letting it inform the conclusions. Finally, some reviews fail to address publication bias — the tendency for studies with positive or striking results to be published more readily than null results — which can skew a review's conclusions if left unexamined, for example through a funnel plot or statistical test where a meta-analysis is performed.
When to Choose a Systematic Review
A systematic review is well suited to questions where an exhaustive, transparent, and reproducible account of the existing evidence matters — for example, questions about the effectiveness of a treatment or intervention, where decisions may be made based on the review's conclusions. It requires a meaningful investment of time, usually run by more than one reviewer, and is not the right tool for every research question.
If you're exploring an emerging topic, want to identify patterns across a broad or loosely related body of work, or don't need the level of methodological rigor a systematic review demands, a narrative literature review is often a more practical and appropriate choice.
Frequently asked questions
No. A meta-analysis is only appropriate when the included studies are similar enough in design, population, and outcome measures to be statistically combined. Many systematic reviews report their findings narratively instead, without a meta-analysis.
PROSPERO is an international registry for systematic review protocols. Registering your protocol before you begin screening isn't mandatory everywhere, but many journals and funders expect it, and it helps demonstrate that your methods were defined in advance rather than adjusted after seeing the results.
PRISMA 2020 revised and expanded the reporting checklist and flow diagram from the original 2009 version, adding more detailed guidance on reporting the search strategy, risk-of-bias assessment, and synthesis methods, to reflect how systematic review methodology had developed since 2009. Reviews published after PRISMA 2020's release are generally expected to follow the newer version rather than the 2009 one.
It depends on the study design being assessed — Cochrane's RoB 2 for randomized controlled trials, ROBINS-I for non-randomized intervention studies, and tools like the Newcastle-Ottawa Scale or JBI's critical appraisal checklists for other observational designs. Using a tool designed for a different study type than the one you're assessing produces an unreliable result.
A meta-analysis statistically pools results from multiple studies into a single combined estimate, which requires the studies to be similar enough to make that pooling meaningful. A narrative synthesis instead describes patterns, agreements, and disagreements across studies in prose, without calculating a pooled number, and is used when studies are too different to responsibly pool.
It varies widely depending on the topic, the number of databases searched, and team size, but a well-resourced systematic review commonly takes several months from protocol registration to a completed manuscript, given the multiple screening and quality-assessment stages involved.
It's possible, but most methodological guidelines recommend at least two reviewers working independently for the screening and quality-assessment stages, since this reduces the risk of errors or bias in which studies get included.
References
- PRISMA Statement
The official PRISMA 2020 reporting guidelines, checklist, and flow diagram for systematic reviews and meta-analyses.
- Cochrane Handbook for Systematic Reviews of Interventions
A comprehensive, widely used methodological handbook for conducting systematic reviews, maintained by Cochrane.
- PROSPERO
The international registry for systematic review protocols, maintained by the University of York's Centre for Reviews and Dissemination.
- JBI
The Joanna Briggs Institute, publisher of internationally recognized systematic review methodology and critical appraisal tools.
- PubMed
A free search engine for biomedical literature, maintained by the National Library of Medicine (NCBI) — a primary database for systematic review searches.
