Your grandfather recorded six hours of stories on a microcassette in 1998. The library asked for a transcript future researchers can search and quote. Or you're running a community oral history project with forty interviews queued up, each two hours long, and a grant deadline.
Oral history transcription is not the same job as transcribing a podcast or a research interview. The recording is usually long, often old, and the transcript has to outlive the audio. That changes how you handle it from the first minute.
What makes oral history transcription different
The Oral History Association treats the transcript as part of the archive, not a convenience copy. That changes three things compared to a normal interview:
- The transcript will be cited. Page numbers, timestamps, and speaker labels need to be unambiguous.
- The audio quality is often bad. Old tapes, narrators in their nineties, kitchen-table recordings. You work with what you have.
- The narrator has a stake. They or their family usually get to review and correct the transcript before it goes public.
If you're transcribing a one-hour Zoom session for a paper, our guide on transcribing an interview for a research paper is the closer fit. This post is about the longer, archival version.
What format should an oral history transcript use?
Most US oral history programs follow a Baylor- or Columbia-style template. The pieces in order:
- A title page with narrator name, interviewer name, location, date(s), and a short abstract.
- A list of session lengths and the audio file names.
- The body, with timestamps every five minutes, speaker labels at every turn, and the narrator's first name (or initials) as the label.
- A subject index or keyword list at the end.
For the body, the timestamp format that holds up in both PDF and Word is [HH:MM:SS] on its own line, followed by the next speaker turn. Page numbers from there on become citation-stable.
Step by step: transcribing an oral history interview
Take notes on speaker shifts, accents, place names, and proper nouns. This index pass saves hours of going back to look things up.
Upload the audio and let a tool produce a base transcript with timestamps and speaker turns. For long files, pick one that handles multi-hour audio without splitting it. You can transcribe long-form recordings here and export with timestamps and speaker labels.
Most tools guess "Speaker 1" and "Speaker 2." Find the first clean turn of each voice, label them by name, then do a find-and-replace.
Headphones, play at 0.75x, edit as you go. Slower audio catches misheard words and the [unintelligible] markers that need a second listen.
Keep or drop um/uh, false starts, and laughter the same way throughout. Record the choice in your style sheet.
Two or three paragraphs describing what's in the interview. This is what researchers read first.
Give them a window, usually 30 to 60 days. Mark their corrections as part of the archival record; do not silently overwrite.
Export to PDF for the public copy and keep a Word file as the editable source. Pair both with the audio file in your archive.
How do you handle older or low-quality recordings?
Microcassette, dictabelt, mini-DV: if it's pre-2010, audio cleanup matters more than transcription tooling. A pass through Audacity to bring the gain up and roll off rumble below 80 Hz can cut your [inaudible] count in half. For more on the audio side, see best practices for audio quality before transcribing.
Then transcribe a five-minute test clip with one or two tools before committing the whole collection. AI accuracy drops sharply on hiss, overlap, and strong regional accents. If you're getting under 85% on the test, budget more editing time per hour.
Paste any public link or upload a file and get a clean transcript in minutes. First 3 clips every month are on us — no card required.
Verbatim or clean? What oral historians actually do
Most US oral history programs use true verbatim with light cleanup: keep meaningful repetitions, false starts that show how the narrator thinks, and significant laughter or pauses. Drop crutch sounds that don't carry meaning. The Baylor style guide gives concrete examples.
This is different from journalism, which leans toward heavy cleanup, and different from court reporting, which captures every utterance. If you're new to the distinction, verbatim vs intelligent transcription covers the tradeoffs.
Whatever you pick, write it in a one-page style sheet and apply it consistently across the whole collection. Inconsistency is worse than either choice.
How long does it take to transcribe an oral history interview?
A safe rule of thumb is 4 to 8 hours of work per recorded hour for a polished, verbatim, narrator-reviewed transcript. The AI pass takes minutes. The cleanup, speaker corrections, and unfamiliar-name research are what actually consume the time.
For a 90-minute interview, that's a full working day. Plan for it. If your project has fifty interviews, that's a year of part-time work. Line up funding before you record.
Common pitfalls and how to avoid them
A few things that trip up new oral history projects:
- Naming files badly. Use
LastnameFirstname_Session1_2026-05-22.wav, notinterview_final.mp3. The transcript inherits the file name. - Trusting AI speaker diarization on the first pass. It mislabels frequently, especially when one voice is much quieter. The speaker diarization explainer covers why.
- Losing the original. Always keep the raw audio and the raw AI transcript untouched. Edit on a copy.
- Forgetting timestamps. Once a transcript is finalized without them, adding them back means another full listen-through. Bake them in from step 2.
- Single-pass editing. A second pass a week later catches errors the first pass introduced: names, dates, place spellings.
For longer interviews, the editing workflow itself starts to dominate; working with long-form transcripts gets into that.
Sources
- Oral History Association, Principles and Best Practices
- Baylor University Institute for Oral History, transcription style and program resources
- Library of Congress, Veterans History Project



