Listening changes what you learn.
A participant gives a short answer, pauses, then says, “Actually, that is not quite what I meant.” Whether the interviewer waits, asks a follow-up, or moves to the next question can change what comes next. A transcript captures the words, but may not show everything that helped the person decide what to share.
AI can draft questions, transcribe speech, translate passages, and suggest ways to group what people say. Researchers interviewed by Schroeder and colleagues described useful possibilities alongside concerns about privacy, bias, access, and the handling of participant data.[1] The practical question is which of these tasks would help your study and how you would check the result.
For each proposed use, write down what material the tool receives, what it produces, and who reviews it. Ask what the participant needs to know about that use. Saving time is valuable when it helps you listen and examine the material more closely. It does not remove responsibility for protecting people’s information or representing them fairly.
Leave room for a participant to pause, disagree, and change what you thought you understood.
Explain AI use before asking people to participate.
“This study uses AI” leaves too much unexplained. Using a tool to draft an invitation is different from sending a recording to a transcription service or analyzing personal accounts. Tell participants what will be recorded, where it will go, who can access it, how long it will be kept, and whether any provider may use it for model improvement. Offer separate choices where the study can support them, and explain what declining a use means for participation.
Check those arrangements before recruitment. Trace a test recording through storage, transcription, analysis, export, and deletion using invented material. Confirm that the tools can support the withdrawal process you intend to offer, and explain any limits. Separate identifying details from research material as early as the study permits.
Put the agreed uses and restrictions in the study plan, with a reviewer for anything requiring approval. Some researchers in Schroeder and colleagues’ study avoided sharing participant data with proprietary language models because of confidentiality concerns.[1] A tool that is acceptable for preparing an invitation may still be unsuitable for the interview material itself.
Use invented participants to test assumptions, not supply testimony.
Asking AI to act as a participant can produce plausible answers immediately. In a study involving nineteen qualitative researchers, generated accounts initially included familiar themes. As the conversations continued, researchers found that the accounts lacked concrete detail and depth. The simulated participants also could not consent, make their own choices, or challenge how they were represented.[3]
A generated character reflects assumptions about someone’s experience. It cannot tell you what that person actually lived through. Suppose a study assumes people avoid a service because its form is confusing. A real participant may instead describe a previous encounter that made them afraid to contact the agency. Recruiting people gives an unexpected account like that a chance to change the study.
AI can help review recruitment wording, prepare translations for checking, or identify assumptions in screening questions. A simple recruitment table can show whose experience is still missing. Do not infer sensitive traits or decide who is credible from their wording. If a group is missing, recruit further or report the gap; do not fill it with a generated quotation.
A plausible generated answer is not evidence of someone’s experience.
Keep attention on the person during the conversation.
Automated interviews can reach more people than a small team could interview individually. Geiecke and Jaravel’s work examines AI-led interviews, including voice interviews, and reports promising quality assessments across several applications.[2] This supports trying the method under suitable conditions. It does not establish that it works equally well for every topic, person, or sensitive disclosure.
Test how the interviewer responds, not just whether people finish. Does a follow-up reflect the previous answer? Does it leave room to disagree? Check what happens when someone corrects a statement, says they do not know, becomes distressed, stays silent, or withdraws. Review misunderstandings and early departures by group, alongside the substance of the answers.
In a human-led interview, a stream of AI suggestions can draw your attention away from the participant. If you use live assistance, explain it and give it a narrow role, such as timekeeping or reminding you of an omitted topic after a section ends. You decide when to pause, wait, or leave the guide. Stop automation when consent is withdrawn, distress needs attention, a system error needs correction, or the tool prevents you from listening closely.
Check that the transcript contains what was actually said.
An automatic transcript is a draft. It may miss a name, change a technical term, merge overlapping speech, or lose a pause or correction. Speech that moves between languages can be especially difficult to check from text alone. Once an error enters the transcript, it can also enter the analysis and the final quotation.
Errors may affect some speakers more than others. A 2020 study tested five commercial speech-recognition systems using conversations from seventy-three Black speakers and forty-two white speakers. Mean word-error rates were 0.35 and 0.19 respectively, or 35% and 19%.[4] These are findings from that study, not current ratings of every transcription tool. They show why you should check the speakers and recording conditions relevant to your own work, with more review where errors are more likely.
Mean word error rate by speaker group
2020 study · Five commercial speech recognition systems
Keep recordings linked to timestamped transcripts in the approved storage environment. Mark unclear passages, overlapping speech, pauses that matter to the analysis, and your corrections. Check every quotation against the recording before publication. For an important interpretation, replay the surrounding conversation too: the sentence alone may not explain what the person meant.
Check the recording before a transcript error becomes a finding.
Keep the original language available during translation.
A fluent translation can still lose meaning. A polite refusal, a family term, or an everyday metaphor may not have a direct equivalent in the report’s language. Van Nes and colleagues explain how translation adds another act of interpretation when research conducted in one language is published in another.[5] The wording needs to preserve the participant’s meaning as well as read naturally.
Keep the original transcript beside the translation, with matching passage identifiers. Note terms with several plausible meanings and record why a version was chosen. Have a bilingual researcher review passages that support central findings. Where possible, begin analysis in the original language so an early translation does not narrow the interpretation.
AI can suggest alternative translations for that review. Agreement between models is not a vote that settles the meaning. Sometimes it is better to retain an original word and explain it. In the report, say which languages were used in the interviews and analysis, who translated, and how important passages were checked.
Explain why you chose an interpretation.
Researchers often label passages to help develop an analysis; these labels are called codes. A list of codes does not automatically become a set of findings. In reflexive thematic analysis, the researcher develops themes by examining the material and reflecting on how their questions, knowledge, and position shape the interpretation.[7] An AI-generated list can make that work look finished before the choices have been examined.
Read a varied selection yourself before asking AI for suggestions, and use only material approved for that purpose. You might ask it to find passages related to a tentative code, compare two possible interpretations, or locate accounts that challenge a theme. Require passage identifiers and return to the full account before using a suggestion. Choose these steps to fit your analytic method rather than treating them as a universal coding procedure.
In a separate study comparing six open models with human coders, people performed consistently across complex sentences, while models’ performance was better on simpler sentences. Some model labels matched the reference labels yet received lower expert ratings.[6] That difference is a reason to examine the interpretation, not just count label matches. Record useful disagreements and explain why a theme belongs in the report, especially where passages are ambiguous or emotionally complex.
A suggested theme needs evidence and a researcher’s explanation.
Agree the uses and data handling before recruitment. These examples still require a study-specific decision and a responsible reviewer.
Make the final account clear to readers and participants.
Consider whether participants should be offered a transcript correction, a review of interpretations, or another opportunity to respond. The appropriate approach depends on the method and relationship. A review can reveal a misquotation or missing context, but it does not automatically settle the truth: people can disagree, change their account, or decline further involvement.
COREQ, a reporting checklist for interviews and focus groups, asks researchers to describe their team, methods, context, and analysis.[8] Apply the same clarity to significant AI use. State the task, tool and version, material supplied, whether it left the research environment, and the human checks performed. Explain any changes the tool made that affected the findings.
Finish the data handling as well as the report. Follow the agreed retention and deletion arrangements, remove temporary copies, and identify who is responsible for any later reuse. Keep only the records the consent and study plan permit. The reader should be able to follow how participants’ words became findings, and participants should have the protections and choices you promised.

