FESK.COMYour global study desk
Email us
AI in Research & Learning

What Is Safe to Paste Into an AI Assistant? Personal Data, Manuscripts, Student Work

Interview transcripts, student grades, a manuscript under review: the things researchers most often paste into AI assistants are exactly what must not leave their hands. A traffic-light rule and how to de-identify first.

What Is Safe to Paste Into an AI Assistant? Personal Data, Manuscripts, Student Work

A master's student pastes 20 teacher interview transcripts into an AI assistant to help code themes. The transcripts name the school and the principal, and include a teacher's account of a conflict with her manager. The consent form the teachers signed promised that data would be “accessible only to the research team”. With one paste, that promise was broken, though nobody meant any harm.

The safety question is not whether the AI tool is good, but where your data goes once it leaves your computer, and whether you have the right to send it there.

Where your data goes

When you send content to an online AI service, depending on its terms, that content may be:

  • Stored on the provider's servers for some period.
  • Used to train or improve models, unless you opt out or use a plan with different commitments.
  • Reviewed by staff in some circumstances, for example to check for policy violations.
  • Processed in a different country from where it was collected.

Free consumer versions and institution-licensed versions often have very different terms on retention and training. Many universities now have agreements with providers, and those are usually the safer option for academic work.

The traffic-light rule

LevelContentWhat to do
Red: do not pasteIdentifiable research participant data; student records, grades and submissions; health information; manuscripts or grant proposals you are reviewing; confidential institutional documents; passwords and access keysKeep out of external tools unless explicitly approved and the tool is licensed for that data
Amber: prepare firstYour own unpublished manuscripts; coded data that still contains specific details; work emailsDe-identify, strip details, prefer institution-licensed tools
Green: generally finePublished text; general methodological questions; fake or simulated data; passages with nothing sensitiveUse normally, still checking outputs

Why manuscripts under review are red

Accepting a review means promising confidentiality. The ideas, data and results are unpublished and belong to the authors. Uploading a manuscript to an AI tool for a “quick summary” or to “draft comments” hands a confidential document to a third party. Many publishers and funders have said explicitly that this is not permitted. The same applies to grant applications you assess.

De-identify before you use AI

If you need AI help with qualitative data, prepare it first:

  1. Replace names of people, institutions and places with codes (T01, School A, Region X).
  2. Remove indirect identifiers: unique job titles, distinctive events, specific dates.
  3. Share only the passages you need, never the whole dataset.
  4. Reread and ask: would someone who knows this participant recognise them?

Even with de-identified data, if the consent form said nothing about external tools, check with your ethics committee. For new projects, state in the consent form and ethics application if you plan to process data with AI tools.

Leaks people overlook

  • Shared conversation links: others can open them, and they may contain data you have forgotten about.
  • Extensions and built-in assistants: browser extensions or email assistants may read whatever you are viewing, not just what you actively send.
  • Photos of documents: a snapshot of a grade sheet to “have AI read the numbers” includes every student's name.
  • Recordings and transcription: online speech-to-text services receive the full audio, including identifiable voices.
  • Synced chat history on shared or lab computers.

The legal backdrop

Many countries have data protection laws requiring a lawful basis and appropriate consent for processing or transferring personal data, especially across borders. The GDPR in Europe applies whenever you work with partners there, and other jurisdictions have comparable rules. You do not need to be a lawyer, but you do need to know one thing: putting personal data into an external service is a form of processing, and responsibility lies with whoever put it there. When in doubt, ask your institution's data protection or legal office.

Safer options

  • Institution-licensed tools that commit not to train on your data.
  • Models running locally or on institutional servers, so data never leaves your infrastructure. Quality may be lower, but they suit sensitive data.
  • Turning off history and training options where available, though this is not a solution for red-level data.
  • Fake data to obtain code or procedures, then applying them to real data on your own machine.

If you have already pasted sensitive data

  1. Delete the conversation and, where the service allows, submit a data deletion request.
  2. Record what you pasted, when and into which tool.
  3. Tell your supervisor or project lead, and your institution's data protection office if red-level data was involved. Many procedures have reporting deadlines for incidents.
  4. Review your workflow so it does not happen again.

Your next step: find out whether your institution offers a licensed AI tool and what its data terms say, then keep the traffic-light table beside your screen. Every time you are about to paste, a one-second glance is enough.

Câu hỏi thường gặp

Can I paste interview transcripts into an AI assistant?

Not if they are identifiable and the consent form does not allow it. De-identify carefully, use an institution-licensed tool, and check with your ethics committee if unsure.

Can peer reviewers use AI to summarise a manuscript?

No. Manuscripts under review are confidential, and uploading them to external AI tools breaches that confidentiality and many publishers' policies.

How do free and institutional AI versions differ on data?

Institution-licensed versions usually commit not to train on your data and have stricter retention terms. Read the specific terms of the tool you use.

Is turning off chat history enough to protect data?

It reduces risk but is not enough for sensitive data such as participant information or student records. That data should stay within approved infrastructure.

What should I do if I pasted personal data into an AI tool?

Delete the conversation, request data deletion if the service supports it, document what happened and inform your supervisor and your institution's data protection office.

Need specific advice for your case?

We will contact you within 24 hours.

Request consultation now

Related articles

🧭
Bạn đang ở chặng nào của đường học vị?
Nhập chỗ bạn đang đứng và đích bạn nhắm — công cụ trả về số năm, chi phí và việc phải làm từng chặng.
Show my pathway →
Miễn phí, không cần tài khoản. Xem tất cả công cụ

Need advice? Talk to us

Leave your details and our team will contact you within 24 hours. The first consultation is completely free.

or
info@fesk.com