You're about to start an interview. The conversation will run sixty minutes. There are no second takes. The single thing that decides whether the transcript is usable is something most people only think about once the recording is already corrupted: the audio.

This checklist runs in four phases — gear, room, settings, opening. Walk it before every interview. It takes ten minutes the first time and three minutes once it's habit. Following it gets you a recording that any modern speech-to-text engine can transcribe cleanly the first time.

Phase 1: Gear (the night before)

Don't troubleshoot equipment ninety seconds before a guest joins. Verify the chain end-to-end the day before.

Phase 2: Room and signal (15 minutes before)

The room shapes the transcript more than the mic does. Speech-to-text engines all degrade on reverb, room rumble, and HVAC hiss, and no plugin un-makes a bad room.

Phase 3: Settings — what to actually record at

Bad defaults silently degrade transcript accuracy. Set these once on the device and forget them.

Phase 4: The opening 30 seconds

The first half-minute decides everything. Don't skip these.

Common mistakes that wreck the transcript

Most ruined recordings come from the same handful of errors. The pattern matters more than the gear.

After the interview: the 60-second wrap

Don't close the laptop yet.

If you want a deeper read on what the engine does to those files once it has them, our pre-transcription audio quality guide digs into the room-treatment side. For the downstream workflow, how to transcribe an interview for a research paper carries it from raw audio to a cited quote.

Try it now — it's free
Transcribe your video with Ask Giya

Paste any public link or upload a file and get a clean transcript in minutes. First 3 clips every month are on us — no card required.

Start transcribing No subscription · 8¢/min after free clips

Common questions

Does any of this matter if I'm paying for a premium AI tool?

Yes, more than people expect. Modern speech-to-text engines degrade hard on reverb, overlap, and low signal-to-noise — the Whisper paper documents this directly. A clean recording at 48 kHz/16-bit pulls noticeably better word accuracy from the same model than a phone-mic recording in a noisy room. The engine doesn't care what you paid; it cares what's in the file.

What if the interview is remote and I can't control the guest's room?

Control what you can. Send the guest a one-paragraph prep email the day before: wired headphones (not AirPods), a quiet room with the door closed, sit close to the laptop. Then record both sides on separate tracks so your clean side isn't muddied by their noisy side. The fallback is your problem to fix in post; their side is theirs to record well in the first place.

Do I need to do all of this if it's a casual chat I'll just listen to?

No. The checklist is for recordings you intend to transcribe or quote from. If you'll never read it back as text, skip phases 3 and 4.

Sources