FESK.COMYour global study desk
Email us
AI in Research & Learning

Using Language Models to Classify Text in Research: Validate Them Like a Scale

Having AI label 20,000 student comments in an afternoon is entirely feasible. But once those labels become variables, the model is a measurement instrument and needs validating like any other.

Using Language Models to Classify Text in Research: Validate Them Like a Scale

An education research team has 20,000 open-ended comments from students about online courses. Coding them by hand would take months. They write a prompt asking a generative AI assistant to assign each comment to one of five categories: content, instructor, technical platform, workload, other. By the end of an afternoon they have a new variable for their regression models. The reviewer asks a single question: “How do the authors know the machine labels measure what they intend to measure?”

It is the right question. When a language model’s output becomes a variable in your analysis, the model is acting as a measurement instrument, like a questionnaire or a human coder. And instruments need validating.

What language models are being used to measure

  • Topics in comments, social media posts and policy documents.
  • Sentiment, attitude or stance.
  • Features of student essays (is there a clear claim, is evidence cited?).
  • Structured information extracted from long texts (dates, amounts, organisation names).

In each case the validation questions are the same: are the labels stable, are they correct, and are they systematically worse for some groups?

Step 1: build a human-coded gold standard

Without a gold standard you cannot validate anything. The minimum process:

  1. Write a codebook with definitions, clear examples and borderline examples for each label, exactly as you would to train human coders.
  2. Draw a random sample, say 500 texts; for rare labels, add a purposive sample so you have enough examples.
  3. Have two people code independently, calculate their agreement and discuss to settle final labels.
  4. Split the gold standard: one part for developing prompts, one part held out for final evaluation only.

Human–human agreement is your reference ceiling. If two experts agree only moderately, do not expect the machine to do much better, and revisit your label definitions. A common mistake is reporting accuracy on the same texts used to refine the prompt: the prompt has been tailored to them, so the figure will flatter the model. The held-out set exists to prevent exactly that.

Step 2: choose appropriate metrics

MetricTells youCaveat
Overall accuracyShare of texts labelled correctlyMisleading with imbalanced labels; if one label covers 80%, guessing gets 80%
Per-label precision and recallWhich labels are over-assigned and which are missedMost important for rare labels
Macro-averaged F1A balanced summary across labelsReport alongside per-label results
Cohen’s kappa or Krippendorff’s alphaChance-corrected agreement, comparable to human–human agreementReport next to human–human figures
Confusion matrixWhich labels get mistaken for whichGuides revisions to definitions and prompts

Step 3: test sensitivity to the prompt

A good instrument does not change its readings when you phrase the question slightly differently. Try:

  • Two or three wordings of the same prompt, with the label order shuffled.
  • With and without examples in the prompt.
  • Repeated runs of the same prompt on the same data, even with randomness settings turned down.

If label proportions across the whole dataset shift noticeably between variants, conclusions built on those proportions are fragile. Reporting this variation is a point in your favour with reviewers.

Step 4: check for systematic bias

High overall accuracy can hide a problem: the machine may measure worse for a particular group. Think of comments written in dialect, without diacritics, mixing languages, or very short ones. Compute metrics separately for groups that matter to your research question. If you plan to compare students from two regions and the model labels one region’s writing style less accurately, the difference you find may simply be measurement error.

Step 5: make it reproducible

  • Record the exact model name and version, run date and parameters.
  • Save the full prompt and all raw outputs, not just final labels.
  • Providers update or retire models; rerunning a year later may give different results. Your saved outputs are your data, so archive them like raw data.
  • Consider open-weight models run locally if long-term reproducibility matters, or if sensitive data must not leave your institution.

Ethics and data

Student comments, posts and case files often contain personal information. Check that your ethics approval permits processing by external services, anonymise before sending, and prefer tools that commit not to train on your data.

What to report

Your methods should include the model and version, the full prompt (in an appendix), how the gold standard was built, human–human agreement, per-label performance on the held-out set, prompt sensitivity results and bias by group. Add a frank sentence on limitations: how measurement error from the model might affect downstream estimates. Some teams also rerun their main analysis on the human-labelled gold standard to see whether conclusions hold, which is a very persuasive check.

If you plan to use a language model to label your own data, start with a codebook and 200 texts you have coded yourself, before writing the first line of a prompt. Without that step you will never know what the numbers in your results table are actually measuring.

Câu hỏi thường gặp

Can I use a large language model to label research data?

Yes, if you validate it like any measurement instrument: compare it with a human-coded gold standard, report per-label metrics, test prompt sensitivity and group bias, and save all outputs for reproducibility.

How many human-coded texts do I need to validate LLM labels?

There is no fixed number; a few hundred randomly sampled texts is a common starting point, plus purposive samples for rare labels. What matters is keeping a held-out portion for final evaluation.

Which metrics should I use to evaluate AI text classification?

Report per-label precision, recall and F1, plus Cohen’s kappa or Krippendorff’s alpha to compare with human agreement. Overall accuracy is misleading when labels are imbalanced.

How can I keep LLM-based research reproducible when models change?

Record the exact model, version, date and parameters, and archive the prompt and all raw outputs as primary data. Locally run open-weight models are an option when long-term reproducibility is critical.

Can AI text classification be biased against certain groups?

Yes. Models may perform worse on dialect, missing diacritics, mixed languages or very short texts. Compute metrics separately for the groups relevant to your research question.

Need specific advice for your case?

We will contact you within 24 hours.

Request consultation now

Related articles

🧭
Bạn đang ở chặng nào của đường học vị?
Nhập chỗ bạn đang đứng và đích bạn nhắm — công cụ trả về số năm, chi phí và việc phải làm từng chặng.
Show my pathway →
Miễn phí, không cần tài khoản. Xem tất cả công cụ

Need advice? Talk to us

Leave your details and our team will contact you within 24 hours. The first consultation is completely free.

or
info@fesk.com