FESK.COMYour global study desk
Email us
Research Methods

Quantitative Content Analysis: Codebooks, Text Sampling and Intercoder Reliability

Counting how often newspapers mention vocational education sounds easy until two coders read the same article and disagree. Here is how to design a content analysis whose numbers survive scrutiny.

Quantitative Content Analysis: Codebooks, Text Sampling and Intercoder Reliability

A team wants to know whether online news covers vocational education positively or negatively. Two members code 300 articles each. When they compare 40 overlapping articles, they agree on “tone” for only just over half. One treats a story about students dropping out of vocational courses as negative; the other codes it neutral because it simply reports figures. Neither is wrong, but the codebook never told them how to decide.

Quantitative content analysis turns texts, images or video into countable data through a systematic coding procedure, so that others can repeat it and get similar results. Unlike qualitative thematic analysis, the emphasis falls on consistent coding and on generalising to a larger body of material.

When to use quantitative content analysis

  • Your question is about frequency, trend or comparison: what share of articles, which sources get quoted more, how coverage changed over five years.
  • You have a large corpus: news, social media posts, textbooks, policy documents, job adverts, research abstracts.
  • Your variables can be defined clearly enough for two independent people to apply them the same way.

If you care about subtle meaning, metaphor or how a discourse is built, qualitative analysis fits better. The two combine well: read a small sample qualitatively to build categories, then code a large sample quantitatively.

Units of analysis and units of coding

The sampling unit is what you select (issues, days, accounts). The coding unit is what receives a code (whole article, paragraph, sentence, image). They can differ. Coding tone for a whole article is a very different decision from coding each paragraph and aggregating. Choose the smallest unit that answers your question; the bigger the unit, the more judgement coders need and the harder reliability becomes.

Sampling texts

News has weekly rhythms: education sections may run early in the week and entertainment at weekends. Simple random sampling of days can be skewed. A constructed week draws one random Monday, one random Tuesday and so on from different weeks in the period, balancing that rhythm. For social media, document exactly how you collected posts (keywords, dates, tool), because feed algorithms shape what you see.

Keep a collection log: date collected, search string, number of hits and exclusion rules (duplicates, photo-only items, advertorials). It plays the role of the participant flow diagram in survey research.

Writing the codebook: one page per variable

The codebook is the heart of the method. Each variable should have:

  1. A name and the question coders must answer.
  2. Values that are mutually exclusive and exhaustive, including “other” and “cannot determine”.
  3. A definition of each value with one typical example and one borderline example.
  4. Decision rules for hard cases: “An article that only reports figures, with no evaluation by the journalist or quoted sources, is coded neutral even if the figures are negative.”
Variable typeExamplesCoding difficulty
ManifestIs a student quoted? Number of named sourcesLow; agreement comes easily
LatentTone; responsibility frame (individual or systemic)High; needs rules and careful training

Training coders and piloting

A standard cycle: coders read the codebook; the team codes 10 to 20 texts together and discusses every disagreement; the codebook is revised; coders independently code a fresh pilot sample; reliability is calculated; repeat until thresholds are met. Log every revision. Never use texts from training rounds to compute the final reliability figures.

Intercoder reliability

Simple percentage agreement is inflated when one value dominates: if 90 per cent of articles quote no students, two coders guessing “no” every time will agree most of the time. You need coefficients that correct for chance agreement:

  • Cohen’s kappa: two coders, nominal variables.
  • Krippendorff’s alpha: any number of coders, handles missing data and different levels of measurement; widely used in communication research.

Commonly cited thresholds are alpha of 0.80 or above for firm conclusions and 0.667 or above for tentative ones, though journals may set their own. Report reliability for each variable, never a single average. The reliability sample is often around 10 to 20 per cent of the corpus, drawn at random and large enough to contain the rare values. Variables that miss the threshold after several revisions should be dropped or have values merged, and you should say so.

Main coding and data management

  • Enter codes through a form with drop-down lists rather than free typing, so typos do not become new values.
  • Store the text ID, coder ID and coding date on every row.
  • Recheck reliability midway if coding runs for weeks, because coders drift from the rules.
  • Archive the original texts or stable links, since online articles get edited or removed.

Reporting

Your methods section should cover the corpus and sampling procedure, the coding unit, number of coders and their training, reliability coefficients per variable with the size of the reliability sample, and how disagreements were settled during main coding. Put the full codebook in an appendix or an open repository. If you use automated classification, report its agreement with human coders on a test sample, just as you would for two humans.

Before you code the first article: pick the hardest variable in your codebook, write down five borderline texts you suspect a colleague would code differently, and write a decision rule for each. Those rules will save you several pilot rounds.

Câu hỏi thường gặp

What is quantitative content analysis?

It is a method for systematically coding texts, images or video into countable data using a fixed codebook, so you can analyse frequencies, trends and comparisons. The procedure must be clear enough for others to repeat it and get similar results.

How is content analysis different from thematic analysis?

Quantitative content analysis applies predefined categories and prioritises consistency between coders so that material can be counted and compared. Qualitative thematic analysis builds themes from the data and focuses on meaning rather than frequency.

What is an acceptable intercoder reliability score?

A commonly cited benchmark is Krippendorff’s alpha of 0.80 for firm conclusions and 0.667 for tentative ones. Report it for each variable and check any specific requirements of your target journal.

Should I use Cohen’s kappa or Krippendorff’s alpha?

Cohen’s kappa suits two coders and nominal variables. Krippendorff’s alpha is more flexible: it works with multiple coders, different levels of measurement and missing codes.

How many texts should be double-coded for reliability?

Often around 10 to 20 per cent of the corpus, selected at random and containing enough of the rarer values. Do not reuse texts from the training rounds for the formal reliability check.

Need specific advice for your case?

We will contact you within 24 hours.

Request consultation now

Related articles

🧭
Bạn đang ở chặng nào của đường học vị?
Nhập chỗ bạn đang đứng và đích bạn nhắm — công cụ trả về số năm, chi phí và việc phải làm từng chặng.
Show my pathway →
Miễn phí, không cần tài khoản. Xem tất cả công cụ

Need advice? Talk to us

Leave your details and our team will contact you within 24 hours. The first consultation is completely free.

or
info@fesk.com