Good audio files: how to get much better transcripts
June 30, 2026 · 7 min read
Format, microphone, levels, and room acoustics drive recognition accuracy. Practical tips for lectures, interviews, and calls.
No transcription tool — AI or human — extracts quality from noisy, clipped, or reverberant recordings. Getting capture right upfront saves hours of correction later.
Recommended file formats
Sprechverlauf supports MP3, WAV, M4A, OGG, FLAC, and WebM. For best recognition:
- WAV or FLAC — best quality, larger files
- M4A (AAC) — good balance, common on phones
- MP3 — at least 128 kbps, better 192 kbps for speech
- Avoid: very low bitrates, repeatedly re-encoded files, long phone-quality mono
Recording checklist
- Mic close to the speaker (15–30 cm), not near laptop fan
- Quiet room — curtains and carpet reduce echo
- Check levels: neither whisper-quiet nor clipping
- Interviews: one mic per person beats one room mic
- Groups: less simultaneous talking
- Add names and jargon in the special-words field (Sprechverlauf step 2)
Before upload: 2 minutes of prep
Listen to the loudest section. Trim long silence at the start. Light noise reduction in Audacity can help — use sparingly. For lectures, sit near the front. For remote calls, use native recording or a dedicated mic, not speaker playback capture.