FESK.COMYour global study desk
Email us
AI in Research & Learning

AI in Systematic Reviews: Screening 4,000 Abstracts Without Losing Key Studies

Screening thousands of titles and abstracts is the most laborious stage of a systematic review. AI can shorten it, but only safely if you measure recall, set a stopping rule and keep humans at the point of decision.

AI in Systematic Reviews: Screening 4,000 Abstracts Without Losing Key Studies

A team is running a systematic review of interventions for test anxiety in school students. Their searches across four databases return 4,180 records after deduplication. The standard is for two reviewers to screen every title and abstract independently. At roughly 30 seconds a record, that is more than 30 hours per person for the first stage alone. No wonder they want AI to help. The right question is not “can AI screen?” but “how many relevant studies will it miss, and how will we know?”

In a systematic review, wrongly excluding a relevant study is far worse than wrongly including an irrelevant one. A false inclusion is caught at full-text stage; a false exclusion disappears for good.

Two ways AI takes part in screening

ApproachHow it worksStrengthsRisks
Prioritisation by active learningYou screen progressively; the model learns from your decisions and pushes likely relevant records to the topHumans still read and decide; most relevant studies surface earlyNeeds a sound stopping rule, or you either read everything or stop too soon
Screening with a large language modelYou give an AI assistant the eligibility criteria and each abstract and receive include or exclude decisionsFast, no training data neededMisreads criteria, output shifts with prompt wording, hard to reproduce

The first approach has been built into review management tools for years and has a substantial evaluation literature. The second is newer and is currently best treated as a support, typically as a second screener alongside one human, not as a replacement.

Recall is the number you must measure

Recall is the proportion of truly relevant records that the AI keeps. To measure it on your own data:

  1. Have two people screen a random sample manually, say 400 records, and reconcile their decisions.
  2. Have the AI screen the same 400 without seeing the human decisions.
  3. Suppose the humans kept 40 records and the AI kept 37 of those: recall is 37/40 = 92.5%.
  4. Study the three records the AI excluded. Were the criteria vague, the abstracts uninformative, or did the AI misunderstand?
  5. Revise the criteria or prompt and test again on a fresh sample, not the same one.

There is no single recall threshold accepted everywhere, but systematic reviews usually aim very high. If the recall you measure is below what you could defend to a peer reviewer, do not use the AI as an independent screener.

Stopping rules for active learning

Once the tool has pushed relevant records to the top, you eventually read hundreds in a row without finding anything to include. When do you stop? Practical stopping rules:

  • Stop after a predefined number of consecutive irrelevant records, set in your protocol in advance, not decided when you are tired.
  • Combine this with a check: screen an additional random sample from the unread records; if any relevant ones appear, continue.
  • Cross-check against a list of “must-find” studies you know are in scope. If any are still unfound, you are not done.

Write criteria that humans and machines read the same way

Vague criteria cause disagreement between two human screeners and even larger errors with AI. Write them using PICO (or your field’s equivalent) with explicit definitions:

  • Who counts as a “student”? Primary, secondary, university?
  • Must “test anxiety” be measured with a validated instrument, or does a simple self-report count?
  • Which designs qualify: randomised trials only, or quasi-experiments too?
  • When the abstract lacks information: include for full-text review. State this rule explicitly in your prompt.

A realistic plan for a two-person team

For the 4,180 records in the example, one plan that balances savings and safety:

  1. Both reviewers independently screen 400 random records, reconcile them and use that set to measure the AI’s recall.
  2. Reviewer one screens the remaining records in the order suggested by an active learning tool until the stopping rule is met.
  3. The AI acts as second screener on the same records; every record where the human and the AI disagree goes to reviewer two.
  4. Reviewer two also checks a random sample of records that both the human and the AI excluded.

Reviewer two’s workload drops sharply, yet every record is still read by at least one person and every disagreement is looked at again.

After screening: data extraction and risk of bias

An AI assistant can help extract details such as sample size, participant characteristics and intervention length into a table. But errors here flow straight into your meta-analysis. The minimum rule: every figure extracted by AI is checked against the original paper by a human, just as in traditional double extraction. Risk-of-bias assessment is expert judgement; AI can flag points to look at but should not deliver the verdict.

Reporting under PRISMA

PRISMA 2020 asks you to describe any automation tools used in study selection. Your methods should state:

  • The tool, version and role (prioritisation, second screener or extraction).
  • The prompt or configuration, provided in full in supplementary material.
  • Recall results from your manual validation sample.
  • Your stopping rule and how many records humans actually screened.
  • How disagreements between humans and the AI were resolved.

If you are drafting a review protocol, write your AI plan and stopping rule into it now, before registration. Deciding before you see the data is what lets readers trust your results, exactly as with every other analytical decision.

Câu hỏi thường gặp

Can I use AI to screen studies for a systematic review?

Yes, most safely as a prioritisation tool or as a second screener alongside a human. Measure recall on a manually screened sample and report the process in line with PRISMA.

What is recall in systematic review screening?

Recall is the share of truly relevant records that a screening method retains. In systematic reviews it matters more than precision because wrongly excluded studies are lost for good.

When should I stop screening with an active learning tool?

Follow a stopping rule set in your protocol, such as a number of consecutive irrelevant records, combined with a random check of unread records and a list of must-find studies.

Can AI extract data for a meta-analysis?

It can assist, but every value it extracts must be checked against the original paper by a person, because extraction errors feed directly into pooled results.

How should AI use be reported in a systematic review?

Report the tool, version, role, prompt or configuration, recall validation results, stopping rule and how human–AI disagreements were resolved, as PRISMA 2020 requires.

Need specific advice for your case?

We will contact you within 24 hours.

Request consultation now

Related articles

🧭
Bạn đang ở chặng nào của đường học vị?
Nhập chỗ bạn đang đứng và đích bạn nhắm — công cụ trả về số năm, chi phí và việc phải làm từng chặng.
Show my pathway →
Miễn phí, không cần tài khoản. Xem tất cả công cụ

Need advice? Talk to us

Leave your details and our team will contact you within 24 hours. The first consultation is completely free.

or
info@fesk.com