How to Do a Thematic Analysis: A Step-by-Step Guide
Thematic analysis is one of the most widely used methods for analyzing qualitative data — interview transcripts, open-ended survey responses, focus group discussions, and similar text-based data. It's flexible enough to fit many research questions and theoretical approaches, but that flexibility can also make it hard to know where to start. This guide walks through the standard six-phase process for conducting a thematic analysis, from first read-through to final write-up.
Thematic analysis follows six phases: familiarizing yourself with the data, generating initial codes, searching for themes among those codes, reviewing the themes against the full dataset, defining and naming each theme, and writing up the analysis. Codes are short labels applied to meaningful segments of text; themes are broader patterns you build by grouping related codes together. The process can be inductive (themes emerge from the data) or deductive (themes are guided by an existing framework or research question), and can be done manually or with qualitative data analysis software such as NVivo.
What Thematic Analysis Is (and When It Fits)
Thematic analysis is a method for identifying, analyzing, and reporting patterns — themes — within qualitative data. It doesn't require a specific theoretical framework the way some qualitative approaches do (such as grounded theory or discourse analysis), which is part of why it's widely used across disciplines including psychology, education, health research, and the social sciences. It fits well when your research question asks what patterns of meaning appear across a dataset — for example, how participants describe an experience, what themes recur across interview transcripts, or how a set of open-ended survey responses cluster around common ideas.
The Six-Phase Process at a Glance
The most widely used framework for thematic analysis, developed by Braun and Clarke, breaks the process into six phases: familiarizing yourself with the data, generating initial codes, searching for themes, reviewing themes, defining and naming themes, and producing the final write-up. The phases are presented in order, but the process is rarely strictly linear in practice — researchers commonly move back and forth between phases, for example returning to re-code the data after an early round of theme development suggests a pattern was missed.
Phase 1: Familiarizing Yourself With the Data
Before any coding starts, the first phase is simply reading and re-reading the data — transcripts, responses, or notes — closely enough to become genuinely familiar with its content and to start noticing recurring ideas. If you conducted the interviews or collected the data yourself, this phase also includes transcribing recordings accurately, since transcription itself is already a first pass of close engagement with the data. Taking informal notes on initial impressions during this phase, without trying to finalize anything yet, makes the next phase faster.
Phase 2: Generating Initial Codes
Coding means systematically working through the dataset and applying short, descriptive labels to segments of text that seem meaningful or relevant to the research question. A single passage can carry more than one code if it touches on multiple ideas. At this stage, the goal is to code as comprehensively as possible rather than to filter for importance — a code that seems minor on a first pass sometimes turns out to be part of a larger pattern once more of the dataset has been coded. Most projects end up with a working codebook of somewhere between roughly 30 and 80 active codes, though this varies widely by dataset size and research question.
Inductive vs. Deductive Coding
Inductive coding lets codes emerge directly from the data itself, without trying to fit the data into a pre-existing framework — it's a good fit for exploratory research questions where you don't want to presuppose what the important patterns will be. Deductive coding starts from an existing framework, theory, or set of research questions, and codes the data against those predefined categories. Many projects use a hybrid approach: starting with a small set of deductive codes tied directly to the research question, while remaining open to adding inductive codes for patterns that didn't fit the original framework. Deciding which approach fits before coding begins keeps the process more consistent.
Phase 3: Searching for Themes
Once the dataset is coded, the next phase is stepping back from individual codes to look for broader patterns among them — a theme is a cluster of related codes that together represent something meaningful about the research question, not just a single repeated code. This phase often involves grouping codes visually (on paper, in a spreadsheet, or using software's code-grouping features), and sorting out which candidate themes are strong enough to stand on their own, which should be combined, and which are better treated as sub-themes within a broader one.
Phase 4: Reviewing Themes
Candidate themes from Phase 3 need to be checked against two things: the coded extracts that support each theme, and the dataset as a whole. A theme should hold together coherently around the extracts coded into it, and themes should be distinct enough from each other that they aren't really the same idea under two names. This is also the point to check for coherence at the level of the whole dataset — do the themes, taken together, tell an accurate and complete story of what's in the data, or are there patterns the current theme set doesn't account for.
Phase 5: Defining and Naming Themes
For each theme that survives the review phase, this step involves writing a clear definition of what the theme captures, what its boundaries are, and how it relates to the research question — plus choosing a concise, descriptive name. A useful test is whether someone unfamiliar with the full dataset could read the theme's name and definition and understand, in a sentence or two, what it represents. Themes that are hard to define crisply at this stage sometimes turn out to actually be two themes that got merged too early.
Phase 6: Writing Up Your Analysis
The final phase turns the analysis into a coherent narrative: presenting each theme with a clear description, supported by direct quotes or extracts from the data, and connecting the themes back to the research question and existing literature. Selected extracts should be genuinely representative of the theme, not just the most quotable or extreme example in the dataset. The write-up should also make the analytical process reasonably transparent — briefly explaining how codes were developed and how themes were derived from them — since this is part of what makes a qualitative analysis credible to reviewers.
Manual Coding vs. Qualitative Software
Thematic analysis can be done manually — using a word processor, a spreadsheet, or annotated printouts — or with dedicated qualitative data analysis software such as NVivo, which supports highlighting and labeling text, organizing codes into hierarchies, and visualizing how codes overlap across the dataset. Software becomes more useful as a dataset grows larger or more complex, since manually tracking which passages carry which codes gets harder to manage by hand past a certain point. For smaller datasets, a well-organized manual approach can be entirely adequate — the software doesn't determine the quality of the analysis, the coding and theme-development decisions do.
Common Mistakes in Thematic Analysis
The most common mistakes are treating a topic summary (for example, "comments about cost") as a theme rather than identifying what the data actually says about that topic; stopping at Phase 2 and presenting a list of codes as if it were a finished analysis; letting themes overlap so heavily that they're not really distinct from each other; and choosing quotes for the write-up because they're vivid rather than because they're representative of the theme as a whole. Deciding upfront whether the approach is inductive or deductive, rather than drifting between the two without acknowledging it, also helps keep the analysis consistent.
When to Get Support With Qualitative Coding
For a small, well-scoped dataset and a fairly direct research question, working through all six phases yourself is often manageable. For larger datasets, more complex or overlapping research questions, multi-coder projects that need to establish coding consistency across researchers, or when you're not yet confident distinguishing a genuine theme from a topic summary, getting input from someone experienced in qualitative analysis — ideally while the codebook is still being developed, not after the analysis is finished — is generally the more efficient path.
Frequently asked questions
A code is a short label applied to a specific, meaningful segment of text. A theme is a broader pattern built by grouping related codes together that, taken as a group, represent something meaningful about the research question.
No — thematic analysis can be done manually with a word processor or spreadsheet, especially for smaller datasets. Software like NVivo becomes more useful as a dataset grows larger or more complex, but it doesn't determine the quality of the analysis on its own.
Inductive coding lets codes emerge from the data itself without a predefined framework, which suits exploratory questions. Deductive coding codes the data against an existing framework or theory. Many projects combine both approaches.
There's no fixed number — it depends on the dataset and research question. What matters more is that each theme is distinct, coherent, and genuinely supported by the coded data, rather than hitting a specific count.
Yes. Thematic analysis is used with sample sizes ranging from a handful of in-depth interviews to hundreds of survey responses — what matters is whether the data can meaningfully answer the research question, not a minimum sample size.
References
- University of Arizona Libraries — Qualitative Coding & Analysis
A university library guide to qualitative coding and analysis tools and approaches.
- George Washington University Libraries — Qualitative Data: Best Practices, Analysis, and Tools
A research guide covering qualitative data methods, including coding and thematic analysis.
- National University Library — Thematic Data Analysis in Qualitative Design
A university library guide to conducting thematic analysis in qualitative research design.
