A research team surveys 600 first-year students about their sense of belonging, planning to measure again at the end of years two and four. At wave two only 390 respond. Worse, 60 responses cannot be linked to wave one because students typed their ID incorrectly or used a different email. By wave three the team realises that those still taking part are mostly high achievers, precisely the students least in need of support.
A longitudinal study measures the same people repeatedly over time. It is the only way to see individual change and the order in which things happen. But every design decision has to assume you will lose people along the way.
What longitudinal data give you that cross-sections cannot
| Question | Cross-section | Repeated cross-section | Panel |
|---|---|---|---|
| What is the current rate? | Yes | Yes | Yes |
| Has the overall rate changed over the years? | No | Yes | Yes |
| Who changed, and by how much? | No | No | Yes |
| What came first within each person? | No | No | Yes |
If your question sits only in the first two rows, do not run a panel. It costs several times more and the attrition risk is not worth it.
Number and spacing of waves
Spacing must match how quickly the phenomenon changes. Exam stress moves week by week; children’s reading moves term by term; graduate earnings move year by year. Measure too rarely and you miss peaks and troughs; measure too often and participants tire and drop out.
- Two waves give a difference score and cannot separate real change from measurement error.
- Three waves are the minimum to describe a linear trajectory.
- Four or more let you test curved trajectories or breakpoints.
Anchor waves to meaningful moments, such as before and after a placement or the start and end of the academic year, rather than spacing them evenly for convenience.
Identifiers: decide before wave one
Unlinkable responses are the cheapest error to prevent and the most expensive to fix. Common options:
- Researcher-assigned codes sent through personalised links. The most accurate, but the name-to-code key must be stored securely and separately from the data.
- Self-generated codes built from a rule, for example the first two letters of your mother’s first name, your day of birth and the last two digits of your phone number. No identities stored, but near-matches need manual checking.
- Avoid anything that changes easily: email address, class group, room number.
Document the linking rule in your data management plan and test it on dummy data before wave one opens.
Keeping participants
Attrition cannot be eliminated, but it can be reduced considerably. Practical measures:
- Collect two or three contact routes at the outset, a backup contact person where appropriate, and explicit consent to be recontacted.
- Send a thank-you and a plain-language summary of early findings between waves so people feel part of the project.
- Keep later questionnaires as short as, or shorter than, the first. Adding questions later is the quickest way to lose people.
- Use a reminder sequence: advance notice, launch, reminder after three days, final reminder.
- Do not drop people who miss a wave. Someone absent at wave two may return at wave three.
- If you offer incentives, increase them slightly at later waves rather than holding them flat or cutting them.
Consistent measurement across waves
To compare, you must measure the same thing in the same way. Do not reword items, reorder scales or switch from paper to online mid-study if you can avoid it. If a change is unavoidable, have a small group answer both versions to estimate the difference.
With psychological scales there is a subtler issue: do people understand the items the same way at each point? A final-year student may interpret “belonging” quite differently from a first-year. Testing longitudinal measurement invariance helps answer this before you compare means over time.
Analysing data when the sample shrinks
Always start by comparing stayers with leavers on their wave-one data. If leavers had markedly lower belonging scores, any conclusion based only on stayers will be too optimistic.
- Complete-case analysis is the simplest option and usually the most biased.
- Mixed-effects or growth models use every available observation for each person, including those who missed a wave, under a missing-at-random assumption.
- Multiple imputation using variables that predict dropout.
- Sensitivity analysis: try different assumptions about those who left and see whether conclusions change.
Report retention at each wave, the characteristics of leavers and your handling method, together with a participant flow diagram across waves.
Checklist before wave one
- The question genuinely requires individual-level change.
- Number and timing of waves match the pace of the phenomenon.
- Identifier rule tested on dummy data.
- Two or three contact routes and consent to recontact.
- A fixed measurement battery for all waves.
- Wave-one sample size inflated for expected attrition.
- A missing-data analysis plan written in advance.
Work the sample size backwards from the final wave: if your analysis needs 250 people at wave three and you expect to retain 80 per cent at each wave, wave one needs about 250 / (0.8 × 0.8) ≈ 391 participants. Do that sum before you apply for funding, not after wave two.
Câu hỏi thường gặp
What is the difference between longitudinal and cross-sectional studies?
A cross-sectional study measures each person once at a single point in time. A longitudinal study measures the same people repeatedly, which reveals individual change and the order in which events occur.
How many waves does a longitudinal study need?
Two waves only give a difference. Three waves are the minimum for describing a trajectory, and four or more are needed to test non-linear change.
What is attrition in a longitudinal study?
Attrition is the loss of participants between waves. It reduces sample size and, more seriously, can leave a remaining sample that differs systematically from the original one.
How do you link survey responses across waves?
Use a stable identifier, either researcher-assigned through personalised links or self-generated by participants from a fixed rule. Avoid email addresses or class groups because they change.
Should I only analyse participants who completed every wave?
Usually not, because stayers rarely resemble leavers. Mixed-effects models or multiple imputation use more of the data, and sensitivity analyses test how much your missing-data assumptions matter.