The findings chapter of many qualitative theses looks like this: “Theme 1: Financial difficulties”, “Theme 2: Family support”, “Theme 3: Time pressure”, each followed by three or four quotes. The reader finishes and still does not know what the study discovered. Those are not themes. They are drawers for sorting data.
A genuine theme is a claim about the data, something that could be right or wrong and is supported by evidence. “Mature students treat study as a loan against family time, so they measure success by not inconveniencing anyone” is a theme. “Family support” is only a heading.
Codes, categories and themes
- Code: a short label attached to a segment of data, often a few words. For example, “studying while children sleep”.
- Category: a group of related codes. For example, “scheduling around family”.
- Theme: a meaningful pattern across the dataset that answers the research question, usually expressible as a full sentence.
The move from category to theme is interpretation, which is where the researcher contributes. Stopping at categories means stopping at description.
A worked coding example
Excerpt (master’s student, 38, two children): “I can only study after ten at night once the kids are asleep. Some nights I open the laptop and just fall asleep on it. My husband is supportive, but I don’t want to ask him, because going back to study is my thing, I chose it.”
| Data segment | Initial code | Category |
|---|---|---|
| only study after ten, once the kids are asleep | studying in leftover time | study pushed to the margins |
| open the laptop and fall asleep | exhaustion while studying | bodily cost |
| husband supportive but I don’t want to ask | declining available help | carrying it alone |
| going back to study is my thing, I chose it | study as personal choice | carrying it alone |
If the pattern “support exists but goes unused because study is framed as a private choice” recurs across participants, you are close to a theme with real weight — one that also cuts against the familiar assumption that lack of support is the main barrier.
Six working phases
- Familiarisation: read every transcript at least once before coding, making margin notes.
- Initial coding: go line by line, keeping codes close to participants’ words. Do not fear having many; 150–300 initial codes across fifteen interviews is normal.
- Searching for categories: cluster codes and sketch relationships.
- Reviewing: return to the raw data and check whether each candidate theme is supported and whether any cases contradict it.
- Defining and naming: write two or three sentences defining each theme and what falls outside it.
- Writing: writing is part of analysis; many themes only take shape when you try to explain them in prose.
Inductive or deductive
Inductive coding lets codes emerge from the data and suits exploratory questions. Deductive coding starts from a theory-based codebook and suits testing an existing framework. In practice most studies combine both, using a framework for the broad containers while remaining open to new codes. State which approach you took; reviewers ask.
Codebooks and memos
A codebook records for each code its name, definition, when to use it, when not to, and an example. It is essential with multiple coders and useful even alone, because the “pressure” code you applied in week one may drift from the one you apply in week six.
Analytic memos are short pieces of writing about ideas, doubts and links between codes. Date every memo. When you write your methods chapter, they form your audit trail.
Software helps but does not analyse
NVivo, ATLAS.ti and MAXQDA help manage codes, retrieve segments and count frequencies. For small projects, the open-source Taguette or even a spreadsheet will do. Software does not generate themes; it just helps you find data quickly. Avoid presenting code counts as strong evidence — frequency in qualitative data mostly reflects who talked the most.
Inter-coder agreement: needed or not?
It depends on your approach. In codebook-driven analysis, reporting an agreement statistic such as Cohen’s kappa on a subset of data is standard. In interpretive thematic analysis, many researchers argue that agreement statistics are a poor fit; team discussion and documenting how disagreements were resolved serve better. Choose what fits your methodology and explain why.
Presenting findings: quotes serve the argument
- Open each theme with the claim, then bring in the data.
- Use quotes as evidence, not as a substitute for your voice. A useful rule of thumb: your analytic commentary should outweigh the quoted material.
- Tag each quote with a participant code and brief context (for example, P07, female, 38).
- Include cases that run against the theme and explain them.
- Avoid vague quantifiers like “most” or “some” without explanation.
Next step: take your current list of themes and rewrite each one as a complete sentence with a verb. Any theme that only works as a noun phrase is still a category — go back to the data and ask what in it would surprise a reader.
Câu hỏi thường gặp
What is the difference between a code and a theme?
A code is a short label for a segment of data, while a theme is a meaningful pattern across the dataset that answers your research question. Good themes can usually be stated as a complete sentence making a claim.
Do I need NVivo for thematic analysis?
No. NVivo and similar tools help manage codes and retrieve quotes, but small projects can use Taguette or a spreadsheet. The quality of analysis comes from your interpretation, not the software.
What is the difference between inductive and deductive coding?
Inductive coding lets codes emerge from the data, while deductive coding starts from a theory-based codebook. Many studies combine the two and should explain the approach in the methods section.
Do I need a second coder in qualitative research?
It depends on your approach. Codebook-based analysis usually reports agreement statistics, whereas interpretive approaches often rely on team discussion and documented resolution of disagreements instead.
How many quotes should I include in qualitative findings?
Enough to evidence each claim, but your own analysis should take up more space than the quotes. Label each quote with a participant code and brief context.