Sprechverlauf

All articles

AI transcription limits: file size, length, and accuracy

August 7, 2026 · 9 min read

File size is only one limit. Recording length, bitrate, audio quality, speakers, language, processing time, and privacy all shape an AI transcript.

AI transcription can turn hours of speech into searchable text, but every service has limits. Some are hard technical boundaries, such as a maximum upload size. Others are practical: a file may upload successfully yet still need careful review because speakers overlap, the recording is noisy, or the vocabulary is highly specialized. Knowing the difference prevents failed uploads and unrealistic accuracy expectations.

What are the limits of AI transcription?

AI transcription is limited by file size, recording duration, processing runtime, audio quality, language support, speaker separation, and the amount of human review required. At Sprechverlauf, the current hard limits are 100 MiB per uploaded file, 30 minutes per file, 90 minutes of audio per day, 10 hours of audio per calendar month, and up to five transcription starts per hour.

LimitWhat it controlsPractical response
File sizeWhether the upload is acceptedCompress or split files above 100 MiB
Recording lengthHow much audio fits inside the file limitUse a speech-friendly bitrate
Processing runtimeWhether the model finishes in timeSplit unusually long recordings
Audio qualityRecognition accuracyReduce noise and keep microphones close
Speaker overlapSpeaker labels and wordingRecord separate tracks or review exchanges
Language and jargonNames and specialist termsSet the language and provide important words

File-size limit: 100 MiB per upload

Sprechverlauf currently accepts uploaded files up to 100 MiB. This is a binary limit: a larger file must be compressed or divided before upload. Supported formats include MP3, WAV, M4A, OGG, FLAC, and WebM, but the format alone does not determine the size. Bitrate, channel count, sample rate, and compression make the difference.

How much recording fits into 100 MiB?

There is no single conversion from megabytes to minutes. A compact speech MP3 can contain hours, while an uncompressed WAV reaches the same size much sooner. The estimates below assume a constant bitrate and exclude a small amount of container metadata, so real files may be slightly shorter.

EncodingApproximate audio within 100 MiBTypical use
MP3 at 64 kbpsAbout 3 h 38 minSpeech where small size matters
MP3 at 128 kbpsAbout 1 h 49 minGood general-purpose speech
MP3 at 192 kbpsAbout 1 h 13 minHigher-quality speech or music sections
Mono WAV, 16-bit/44.1 kHzAbout 20 minUncompressed editing master

For spoken-word material, exporting a copy as mono MP3 or M4A often cuts the upload size dramatically without removing information the speech model needs. Keep the original recording separately; transcribe the smaller working copy.

Is there a maximum recording length?

Sprechverlauf does not impose a separate paid minute balance or daily duration quota. In practice, file size and processing runtime create the useful boundary. A long, low-bitrate recording may fit under 100 MiB, but splitting very long sessions into logical sections makes uploads easier, reduces the cost of retrying a failed section, and produces more manageable transcripts.

  • Keep interviews or meetings as one file when they fit comfortably and speaker continuity matters
  • Split multi-hour conferences by session, speaker, or agenda item
  • Remove long silence and music-only passages before export
  • Preserve a few seconds of context at each cut instead of splitting mid-sentence
  • Use clear filenames so the transcript sections remain in order

Usage limit: five transcription starts per hour

The free service currently permits up to five transcription starts per hour as technical anti-abuse protection. This is not a bank of paid minutes and it does not become a smaller allowance for long audio. If you have a large batch, prepare and check all files first, then process the most important recordings in order rather than spending starts on duplicate or incomplete uploads.

Accuracy has no honest universal percentage

A single accuracy claim cannot describe every recording. Clear close-mic speech in a supported language is much easier than a distant group discussion with echo, interruptions, and domain-specific names. Even a transcript that looks fluent can contain a wrong number, proper noun, or speaker assignment, so apparent readability is not proof of factual accuracy.

  • Background noise can hide consonants and word endings
  • Echo and clipping remove details that software cannot reconstruct reliably
  • Overlapping speech makes both wording and speaker attribution harder
  • Accents, code-switching, and rare languages may increase errors
  • Names, abbreviations, product terms, and numbers deserve manual checking
  • Music, applause, and long cross-talk can create unusable passages

Speaker recognition is useful, not infallible

Speaker diarization answers who spoke when; it is separate from recognizing the words. Distinct voices with clean turn-taking usually separate well. Similar voices, one-word interruptions, shared microphones, and simultaneous speech can merge or swap labels. Listen to the opening, map generic labels to real names, and verify every attribution that will be quoted or published.

Language and specialist vocabulary limits

Multilingual models cover many languages, but performance is not identical across them. Setting the known recording language prevents unnecessary detection errors. For courses, research interviews, and technical meetings, add important names, abbreviations, and specialist words to the prompt field. This supplies context, but it does not replace proofreading.

Processing and privacy limits

Sprechverlauf sends the recording to Replicate for AI processing. Replicate documents a 30-minute runtime timeout for predictions and removes API prediction inputs, outputs, files, and logs automatically after about one hour by default. Sprechverlauf does not permanently store the audio or transcript content on its own server, so export the result as Word or PDF when you want to keep it.

When human review is still required

AI output is best treated as a fast first draft when mistakes have consequences. Verify quotations against the recording, check figures and names, and use a qualified human workflow for legal, medical, safeguarding, or other high-stakes material. The correct question is not only whether the model can transcribe the file, but whether the resulting evidence is reliable enough for its intended use.

Checklist before starting a long transcription

  • Confirm the file is 100 MiB or smaller
  • Convert oversized WAV files to a speech-quality MP3 or M4A copy
  • Split multi-hour recordings at natural section boundaries
  • Select the recording language instead of relying on detection when known
  • Add names and specialist terms as context
  • Use the correct number of speakers for diarization
  • Plan to verify quotations, figures, names, and sensitive claims
  • Export the finished transcript before leaving the result page

Sources & further reading