The hardest part of finishing an interview isn't the conversation. It's the hour you lose afterwards staring at a blank page, trying to remember whether the recording opens with consent or whether your subject's title was "associate professor" or "associate director."
Use the template below. Drop your audio in, fill in the gaps, and you have a transcript a colleague (or a citation manager) can actually read.
- A working interview transcript needs three layers: metadata, the conversation itself, and post-interview notes.
- Pick verbatim, cleaned, or summary up front and stick with it for the whole document.
- Cleaned-up is the right default for journalism, qualitative research, and podcasts.
- AI transcription + 20–40 minutes of cleanup is the fastest path for most workflows.
What goes into a usable interview transcript
A working transcript has three layers: metadata (who, when, where, consent), the conversation itself with speaker labels and timestamps, and a short post-interview note. Skip any of those and you'll re-do work later, usually at 11 p.m. when you're trying to cite a quote and can't find the original date.
The template below covers all three. It's plain text, so it pastes cleanly into Google Docs, Word, Notion, or your research notes app.
The template (copy-paste this)
INTERVIEW TRANSCRIPT
Interviewer: [Your name]
Interviewee: [Full name, title/role]
Organization: [If relevant]
Date: [YYYY-MM-DD]
Location: [City, or "Zoom" / "Phone" / "Google Meet"]
Duration: [HH:MM:SS]
Recording file: [Filename or link]
Consent: [Verbal / written, time recorded]
Transcription: [Verbatim / Cleaned / Summary]
Project: [Research project, article, episode]
---
ABSTRACT
[2–3 sentences summarizing the conversation. Write this last.]
KEY QUOTES
- "..." (00:00:00)
- "..." (00:00:00)
---
TRANSCRIPT
[00:00:00] Interviewer: ...
[00:00:00] [Interviewee initials]: ...
[00:00:00] Interviewer: ...
[00:00:00] [Interviewee initials]: ...
---
POST-INTERVIEW NOTES
- [Follow-up questions, observations on tone, things to verify]
- [Permissions: can you quote on the record? attribute by name?]
A filled-in example
INTERVIEW TRANSCRIPT
Interviewer: Maya Chen
Interviewee: Dr. Lena Ortiz, Associate Professor of Sociology
Organization: University of Texas at Austin
Date: 2026-05-14
Location: Zoom
Duration: 00:42:18
Recording file: ortiz-interview-2026-05-14.m4a
Consent: Verbal, recorded at 00:00:08
Transcription: Cleaned (verbatim available on request)
Project: Housing instability dissertation, Ch. 3
---
ABSTRACT
Dr. Ortiz discusses fieldwork methods for studying renter displacement
in Austin. She argues against using surveys as a primary instrument and
favors longitudinal interviews with the same households over two years.
KEY QUOTES
- "Surveys ask people to summarize a year of their life in 12 questions. That's not data, that's pressure." (00:11:42)
- "We followed 38 households for 22 months. By the end, eleven had moved twice." (00:24:09)
---
TRANSCRIPT
[00:00:08] Chen: Thanks for making time. Before we start, are you okay with me recording this and quoting you by name?
[00:00:15] LO: Yes, on the record is fine.
[00:00:21] Chen: Tell me about the project you wrapped up last spring.
[00:00:28] LO: We were looking at displacement in East Austin...
---
POST-INTERVIEW NOTES
- Follow up: ask for the 22-month dataset codebook.
- Quote permissions: on the record, attributable.
- Sounded tired toward the end; consider a second short call rather than a 90-minute follow-up.
How to fill in each section
A few notes on the fields people most often get wrong.
- Consent. Record the moment of consent inside the audio itself, and note the timestamp. If you ever need to defend an attribution, that timestamp is what you'll reach for. For sensitive topics, also keep a signed consent form alongside the file.
- Transcription type. Pick one and stick to it. Verbatim keeps every "um" and false start (right for linguistics, accessibility, court). Cleaned drops fillers but preserves wording (right for journalism and most research). Summary paraphrases. Mixing them mid-document makes the transcript impossible to cite. If you're unsure, the verbatim vs intelligent transcription guide lays out the trade-offs.
- Timestamps. Every 30–60 seconds is enough for most uses. Every speaker turn is overkill unless you're doing conversation analysis. Anchor them to the source file's timecode, not the wall clock.
- Speaker labels. Use full names on first appearance, then initials. Two speakers with the same initials? Use a short surname instead. Don't ship the final document with "Speaker 1 / Speaker 2" — that's fine in a draft, but it's brutal to read later.
Verbatim, cleaned-up, or summary — which to use?
Default to cleaned-up for journalism, qualitative research, podcast show notes, and anything you'll share with the interviewee. It reads like the person speaks, not like a court reporter's printout.
Go verbatim when accuracy of every utterance matters: linguistic analysis, legal proceedings, accessibility captions, or when you're studying speech itself. Verbatim is also what you want if there's any chance the transcript will be challenged.
A summary is not really a transcript. It's notes. Don't call it a transcript when you cite it.
How to actually get the transcript
You have three options.
- Type it yourself. A 60-minute clear-audio interview takes a fluent typist 4–6 hours. Most accurate, most painful. Worth it only when the stakes demand a human touch on every sentence.
- Hire a human service. $1.00–$3.00 per audio minute, 24–48 hour turnaround, around 99% accuracy on clean audio. Sensible for legal, medical, or anything you'd defend in court.
- Use AI transcription, then clean. A few cents per minute, turnaround in seconds to a few minutes, 90–95% accuracy on clean audio. You spend 20–40 minutes per interview cleaning it up, but you're editing, not typing.
For most researchers and journalists, the third option wins on time and cost. Transcribe the recording with VTS, paste the output into the TRANSCRIPT block above, and spend your time on the parts a machine can't do: the abstract, the key quotes, and the follow-up plan.
If you have a multi-speaker recording, the Zoom multi-speaker guide covers diarization. And before you record, run through the audio quality checklist. Ten minutes there saves an hour later.
Paste any public link or upload a file and get a clean transcript in minutes. First 3 clips every month are on us — no card required.
Common questions about interview transcripts
How long should an interview transcript be?
As long as the interview. A 30-minute conversation produces roughly 4,500–5,500 transcribed words at typical speaking pace. Don't cut for length. Cut for clarity in your final article or paper, but keep the transcript whole.
Should I send the transcript to the interviewee?
For research interviews, often yes. Many IRB protocols require "member checking" so subjects can correct factual errors. For journalism, almost never. Sending a transcript back invites edits and retractions you don't want.
Do I need to transcribe the off-topic parts?
Transcribe everything inside the recording window. If your subject went off the record at minute 14, you should have stopped the recording then, not transcribed and redacted it later. Mark gaps as [off the record, 00:14:02 — 00:18:30].
How do I cite an interview transcript?
APA: Ortiz, L. (2026, May 14). Personal interview. MLA: Ortiz, Lena. Personal interview. 14 May 2026. Both styles assume the transcript is on file with the author and unpublished. If you publish the transcript, link to where it lives.
What format should I save it in?
Plain text or Markdown for the working copy (small, searchable, version-controllable). PDF for the archival copy you submit with a paper or article. Avoid .docx as the only copy. Proprietary formats age badly.



