A communications PhD student collects 40,000 comments from parenting groups to study anxiety about exams. In her thesis she quotes twenty “typical” comments verbatim, with names removed. A week after the thesis appears in the university repository, a parent pastes her own comment into a search engine and finds it in the thesis, next to an analysis of “pressure-applying parents.” The name was gone, but one search was enough to trace the author.
Data from social media, forums and websites is a goldmine for social science, education and public health. It also raises ethical questions that procedures designed for surveys and interviews do not fully answer.
“Public” is a spectrum, not a switch
Someone posting a public comment does not necessarily imagine it becoming research data. Expectations of privacy vary with context:
| Type of content | Expectation of privacy | Approach |
|---|---|---|
| Organisational accounts, public figures speaking officially | Low | Can be quoted and named, like published statements |
| Individuals’ public posts on everyday topics | Moderate | Aggregate analysis is fine; verbatim quotes need care |
| Public groups on sensitive topics (illness, mental health, sexuality) | High | Ethics review needed; limit verbatim quotation |
| Closed or approval-only groups | Very high | Consent from moderators and usually members |
| Private messages | Complete | Do not use without explicit consent |
Joining a closed group as an ordinary member and quietly harvesting its content is the clearest violation, even if you technically “have access.”
Do you need ethics approval?
Rules vary. Some institutions treat analysis of public, non-identifiable data as exempt; others require an application for any research using data about people. The safe move is to ask your ethics committee before collecting, describing the source, collection method, storage and how you will quote. Even if you are exempt, you will then have documentation to show a journal.
The Association of Internet Researchers (AoIR) ethics guidelines are a useful reference for your application, because they were written for exactly these grey areas.
Platform terms and automated collection
- Read the platform’s terms of service and any researcher policy. Some prohibit automated collection outside official programming interfaces.
- Prefer official research access routes where they exist, or datasets that have been shared legitimately.
- Collect the minimum: only fields your question needs. If you do not need profile photos, friend lists or locations, do not collect them.
- Record the collection date, search terms, scope and tools so your methods section is transparent.
Verbatim quotes: the most overlooked issue
Removing usernames does not anonymise a quote, because search engines can trace the exact wording. Options, from safest:
- Do not quote verbatim; describe and summarise.
- Controlled paraphrase: create composite examples that keep the meaning but change the wording, and state in your methods that quotes were altered to protect authors. Test with a search engine that they cannot be traced.
- Ask the author’s permission when the exact wording is essential, for instance when phrasing itself is what you analyse.
For sensitive content, options 1 or 2 are almost always right. Screenshots of individuals’ posts in a paper are almost never appropriate.
Storing and sharing what you collect
Raw social media datasets contain usernames, profile links and sometimes sensitive details. Separate identifiers into their own file, encrypt storage and limit access. When sharing under open data expectations, the usual approach is to share post IDs (so others can retrieve them under the platform’s terms) plus your analysis code, not the raw content. If users have since deleted posts, respect that decision in anything you publish.
When researchers do more than observe
Posting test content to see how people react, creating a fake account to chat with users, or messaging people to recruit them are interventions, not observation. Such designs almost always need full ethics review, and any design that conceals the researcher’s identity needs a compelling justification and a plan for debriefing participants afterwards.
When minors or vulnerable groups may be involved
User ages are often unknowable. On platforms popular with young people, or for topics such as self-harm, bullying or mental health, assume vulnerable people are in your data. Increase anonymisation, avoid verbatim quotes and plan in advance what you will do if you encounter content suggesting someone is in danger.
A checklist before you collect
- Where does the source sit on the privacy spectrum?
- Do the platform’s terms allow this collection method?
- Has the ethics committee reviewed it or confirmed an exemption?
- Which fields will I collect, and which can I drop?
- How will I quote so that nobody can be traced?
- Where will raw data be stored, who can access it and when will it be deleted?
A good next step: write half a page describing your source, collection method and quoting approach against this checklist, and send it to your ethics committee or supervisor before running any collection tool. That half page will later become the backbone of your methods section.
Câu hỏi thường gặp
Can I use public social media posts for research without consent?
Aggregate analysis is often acceptable, but it depends on how sensitive the content is, the platform’s terms and your ethics committee’s rules. Closed groups, sensitive topics and verbatim quotes need particular care or consent.
Do I need ethics approval to analyse Twitter or Facebook data?
It depends on your institution; many require an application or a formal exemption. Ask before collecting so you have documentation if a journal requests it.
Is removing usernames enough to anonymise social media quotes?
No, because search engines can trace exact wording back to the original post. Summarise, use controlled paraphrase or ask permission if you must quote verbatim.
Can I collect data from a private group I belong to?
Not without consent from the moderators and usually the members. Membership gives you access to read, not permission to use content as research data.
How should social media data be shared for open science?
The usual practice is to share post identifiers and analysis code so others can retrieve the data under platform terms, rather than sharing raw content containing user information.