Interview research workflow
How to Organize Interview Transcripts With AI Without Losing the Speaker
The transcript is a working representation of the interview. Preserve the recording, speaker, timestamp, context, and editorial status behind every extracted quote.
Method and ownership note: ChatGrid publishes this workflow and supports short audio/video and recorded WebM context in bounded use. It does not claim perfect transcription, automatic speaker identification, or support for arbitrary recording length.
Direct answer
The working method
Keep the original recording and raw transcript unchanged, then work from a review copy. Normalize speaker labels and timestamps, segment by topic, and ask AI to propose themes and candidate quotations without rewriting the speaker. Each usable quote should retain the speaker, exact words, timestamp, surrounding context, and verification status. Draft from verified quote cards and thematic notes, never from a polished AI summary alone.
Preserve the raw interview before organizing it
Keep the original recording, a raw machine transcript, and a separate working transcript. Do not correct the only transcript in place. The raw layer lets an editor inspect how names, pauses, interruptions, and uncertain words were originally captured; the working layer can add structure without pretending those changes came from the speaker.
Record interviewee, interviewer, date, location or medium, recording parts, transcript method, and any usage conditions supplied for the project. The Smithsonian Archives' oral-history guide recommends contextual introduction material about why the person was selected and the circumstances of the interview. That context also helps a future writer interpret the exchange.
Normalize speakers and timestamps before themes
Replace unstable labels such as Speaker 1 with confirmed names or neutral identifiers. Add timestamps at regular intervals and at every speaker change used for a quotation. Mark uncertain speaker assignments instead of guessing. If several recordings were joined, preserve the part number so a timestamp remains reproducible.
Correct obvious transcription errors in the working copy while keeping an edit log for material changes. Names, organizations, numbers, dates, and technical terms deserve a recording check. Use a consistent token such as [unclear 00:14:22] when audio cannot be resolved; a plausible completion is more dangerous than visible uncertainty.
Segment by topic without breaking the exchange
Ask AI to propose topical boundaries using timestamp ranges and short neutral labels. Review those boundaries against the conversation. Keep the question that elicited an answer with the answer, and preserve interruptions or follow-ups when they change meaning. A quote about “it” may be unusable once detached from the object named thirty seconds earlier.
Use overlapping segments when one passage belongs to two themes, rather than duplicating and silently editing it. Keep a segment ID, start and end time, speakers, topic, one-sentence description, and relevant project question. The segment becomes a navigation aid, not a rewritten account of the interview.
Generate themes as hypotheses
Give the model a bounded set of segments and request themes with supporting segment IDs, exceptions, and alternative interpretations. A theme without a locator is a suggestion, not a finding. Ask which statements do not fit the proposed pattern; outliers often matter more than the most repeated phrase.
Separate descriptive themes—what topics recur—from interpretive themes—what those recurrences may mean. Keep the interpretation attributed to the writer unless the interviewee explicitly said it. If several interviews are compared, preserve participant identity and context rather than merging everyone into an anonymous composite voice.
Build quote cards and verify every word
A quote card should contain the exact transcript excerpt, speaker, recording and timestamp range, preceding question, short context, theme, edit status, and verification status. Replay the audio before accepting the quote. Check that cuts do not change tense, subject, certainty, or the relationship between question and answer.
If the publication style allows light cleanup, distinguish it from verbatim transcription and follow the interview project's editorial rules. Never let a generated paraphrase acquire quotation marks. Store paraphrases as writer notes with a link to the underlying segment.
Draft from verified segments, then return to the whole interview
Outline with themes and quote-card IDs. Use quotations where the speaker's language adds authority, character, or precision; summarize routine material in the writer's voice with accurate attribution. Keep contradictory or qualifying statements close enough that the reader is not given a false impression of certainty.
Before release, listen to the wider passage around every central quote and scan the full interview for later correction or withdrawal. Confirm that the planned use matches the permissions and editorial commitments governing the project. AI can reduce navigation time, but the writer remains responsible for representation.
Frequently asked questions
Can AI identify speakers in an interview transcript?
It can propose labels, but speaker identity should be confirmed against the recording and interview notes. Mark uncertain assignments explicitly instead of treating diarization as fact.
Should I clean up interview quotes with AI?
Use AI to flag disfluencies or propose edits only under the project's editorial rules. Verify the final wording against the recording, disclose non-verbatim treatment when appropriate, and never put an AI paraphrase in quotation marks.
What metadata should an interview quote keep?
Keep speaker, recording or part, start and end timestamps, preceding question, relevant context, theme, edit status, and verification status. This makes the quote reproducible and reviewable.
Primary sources
- How to Do Oral History
Smithsonian Institution Archives
Primary institutional guide covering interview preparation, transcript context, and post-interview documentation.
- Add subtitles and captions
YouTube Help
Primary documentation illustrating how caption text and timing represent spoken media.