AI transcription limits: file size, length, and accuracy
August 7, 2026 · 9 min read
File size is only one limit. Recording length, bitrate, audio quality, speakers, language, processing time, and privacy all shape an AI transcript.
AI transcription can turn hours of speech into searchable text, but every service has limits. Some are hard technical boundaries, such as a maximum upload size. Others are practical: a file may upload successfully yet still need careful review because speakers overlap, the recording is noisy, or the vocabulary is highly specialized. Knowing the difference prevents failed uploads and unrealistic accuracy expectations.
What are the limits of AI transcription?
AI transcription is limited by file size, recording duration, processing runtime, audio quality, language support, speaker separation, and the amount of human review required. At Sprechverlauf, the current hard limits are 100 MiB per uploaded file, 30 minutes per file, 90 minutes of audio per day, 10 hours of audio per calendar month, and up to five transcription starts per hour.
| Limit | What it controls | Practical response |
|---|---|---|
| File size | Whether the upload is accepted | Compress or split files above 100 MiB |
| Recording length | How much audio fits inside the file limit | Use a speech-friendly bitrate |
| Processing runtime | Whether the model finishes in time | Split unusually long recordings |
| Audio quality | Recognition accuracy | Reduce noise and keep microphones close |
| Speaker overlap | Speaker labels and wording | Record separate tracks or review exchanges |
| Language and jargon | Names and specialist terms | Set the language and provide important words |
File-size limit: 100 MiB per upload
Sprechverlauf currently accepts uploaded files up to 100 MiB. This is a binary limit: a larger file must be compressed or divided before upload. Supported formats include MP3, WAV, M4A, OGG, FLAC, and WebM, but the format alone does not determine the size. Bitrate, channel count, sample rate, and compression make the difference.
How much recording fits into 100 MiB?
There is no single conversion from megabytes to minutes. A compact speech MP3 can contain hours, while an uncompressed WAV reaches the same size much sooner. The estimates below assume a constant bitrate and exclude a small amount of container metadata, so real files may be slightly shorter.
| Encoding | Approximate audio within 100 MiB | Typical use |
|---|---|---|
| MP3 at 64 kbps | About 3 h 38 min | Speech where small size matters |
| MP3 at 128 kbps | About 1 h 49 min | Good general-purpose speech |
| MP3 at 192 kbps | About 1 h 13 min | Higher-quality speech or music sections |
| Mono WAV, 16-bit/44.1 kHz | About 20 min | Uncompressed editing master |
For spoken-word material, exporting a copy as mono MP3 or M4A often cuts the upload size dramatically without removing information the speech model needs. Keep the original recording separately; transcribe the smaller working copy.
Is there a maximum recording length?
Sprechverlauf does not impose a separate paid minute balance or daily duration quota. In practice, file size and processing runtime create the useful boundary. A long, low-bitrate recording may fit under 100 MiB, but splitting very long sessions into logical sections makes uploads easier, reduces the cost of retrying a failed section, and produces more manageable transcripts.
- Keep interviews or meetings as one file when they fit comfortably and speaker continuity matters
- Split multi-hour conferences by session, speaker, or agenda item
- Remove long silence and music-only passages before export
- Preserve a few seconds of context at each cut instead of splitting mid-sentence
- Use clear filenames so the transcript sections remain in order
Usage limit: five transcription starts per hour
The free service currently permits up to five transcription starts per hour as technical anti-abuse protection. This is not a bank of paid minutes and it does not become a smaller allowance for long audio. If you have a large batch, prepare and check all files first, then process the most important recordings in order rather than spending starts on duplicate or incomplete uploads.
Accuracy has no honest universal percentage
A single accuracy claim cannot describe every recording. Clear close-mic speech in a supported language is much easier than a distant group discussion with echo, interruptions, and domain-specific names. Even a transcript that looks fluent can contain a wrong number, proper noun, or speaker assignment, so apparent readability is not proof of factual accuracy.
- Background noise can hide consonants and word endings
- Echo and clipping remove details that software cannot reconstruct reliably
- Overlapping speech makes both wording and speaker attribution harder
- Accents, code-switching, and rare languages may increase errors
- Names, abbreviations, product terms, and numbers deserve manual checking
- Music, applause, and long cross-talk can create unusable passages
Speaker recognition is useful, not infallible
Speaker diarization answers who spoke when; it is separate from recognizing the words. Distinct voices with clean turn-taking usually separate well. Similar voices, one-word interruptions, shared microphones, and simultaneous speech can merge or swap labels. Listen to the opening, map generic labels to real names, and verify every attribution that will be quoted or published.
Language and specialist vocabulary limits
Multilingual models cover many languages, but performance is not identical across them. Setting the known recording language prevents unnecessary detection errors. For courses, research interviews, and technical meetings, add important names, abbreviations, and specialist words to the prompt field. This supplies context, but it does not replace proofreading.
Processing and privacy limits
Sprechverlauf sends the recording to Replicate for AI processing. Replicate documents a 30-minute runtime timeout for predictions and removes API prediction inputs, outputs, files, and logs automatically after about one hour by default. Sprechverlauf does not permanently store the audio or transcript content on its own server, so export the result as Word or PDF when you want to keep it.
When human review is still required
AI output is best treated as a fast first draft when mistakes have consequences. Verify quotations against the recording, check figures and names, and use a qualified human workflow for legal, medical, safeguarding, or other high-stakes material. The correct question is not only whether the model can transcribe the file, but whether the resulting evidence is reliable enough for its intended use.
Checklist before starting a long transcription
- Confirm the file is 100 MiB or smaller
- Convert oversized WAV files to a speech-quality MP3 or M4A copy
- Split multi-hour recordings at natural section boundaries
- Select the recording language instead of relying on detection when known
- Add names and specialist terms as context
- Use the correct number of speakers for diarization
- Plan to verify quotations, figures, names, and sensitive claims
- Export the finished transcript before leaving the result page