Speech-to-text models have diversified rapidly since Whisper’s release. In 2026, practitioners choosing a local STT model face a genuine selection problem: there are fast models and accurate models and streaming models and multilingual models, and the best choice depends on the specific deployment scenario. This guide cuts through the options with a task-first framework — what are you building, what are the hard constraints, and which model family wins given those constraints.
The Four Deployment Scenarios
Most real STT use cases fall into four scenarios, each with different priority orderings. Batch transcription of pre-recorded audio prioritises accuracy and throughput; latency is not a concern since the audio already exists. Real-time streaming transcription prioritises latency — partial results must appear while the speaker is still talking — with accuracy as a secondary constraint. Edge or embedded deployment prioritises model size and CPU performance — the model must run on a device with 4–8GB RAM and no dedicated GPU. Multilingual transcription prioritises language coverage and code-switching handling — the model must handle multiple languages reliably, ideally within the same audio segment. Different models win in each scenario, and conflating them leads to suboptimal choices.
Batch Transcription: Faster-Whisper Wins
For batch transcription — processing audio files after the fact, latency not a concern — Faster-Whisper large-v3 is the default recommendation. It delivers the same word error rates as original Whisper large-v3 (the accuracy ceiling for local open-source models on diverse audio) at 4–8× higher throughput, with lower memory usage via INT8 quantisation. On a single A10G GPU, Faster-Whisper large-v3 processes audio at approximately 20× real-time — a 60-minute meeting transcribes in under 3 minutes. On a modern CPU without GPU, it achieves approximately 3–5× real-time. The model’s multilingual capability (99 languages), robustness to noisy audio, and high accuracy on specialised vocabulary make it the right default for any batch workload. The only scenario where a different model wins on batch is English-only high-volume workloads where Parakeet’s 40–60× speed advantage justifies its English-only limitation.
Real-Time Streaming: Parakeet or Streaming Whisper
For real-time streaming transcription, the requirement is producing partial results quickly enough that they appear responsive — typically, partial words or completed words within 200–400ms of being spoken. Two options compete here. Parakeet-TDT (NVIDIA) achieves sub-50ms latency per audio chunk on GPU, produces word-level timestamps, and is designed for streaming from the ground up. Its CTC architecture allows chunk-by-chunk processing without re-processing previous audio. The limitation is English-only and reduced accuracy on heavily accented or noisy audio compared to Whisper. Streaming Whisper implementations (using whisper-live, faster-whisper streaming mode, or WhisperKit for Apple) work by processing overlapping audio windows and using speculative decoding to produce partial results faster. These achieve lower accuracy than batch Whisper due to the windowing approach but support all 99 languages. For English-language real-time applications on GPU, Parakeet is the better choice. For multilingual streaming or CPU-only streaming, use a streaming Whisper implementation.
Figure 1 — STT model selection by deployment scenario
Edge Deployment: Whisper.cpp
Whisper.cpp is a C++ reimplementation of Whisper optimised for CPU inference, including Apple Silicon acceleration via Metal and x86 optimisation via AVX2. On an Apple M2 Pro, Whisper.cpp medium achieves approximately 5–8× real-time on CPU — fast enough for near-real-time transcription of short clips. The small model achieves 15–20× real-time on M2 Pro. On Raspberry Pi 5, Whisper.cpp tiny achieves approximately 1–2× real-time (processing in roughly the same time as the audio). For embedded deployment where a Python runtime is unavailable, Whisper.cpp’s C++ codebase compiles to a standalone executable. The Android and iOS bindings make it the primary option for on-device mobile transcription. Accuracy is limited to the model size you can run — the tiny model’s 8–12% WER is acceptable for simple voice commands but not for accurate meeting transcription.
Multilingual and Translation: SeamlessM4T
For applications requiring transcription and translation in a single pass — a voice assistant that handles multiple languages and translates responses — Meta’s SeamlessM4T handles approximately 100 languages with integrated speech-to-speech and speech-to-text translation. The large model achieves word error rates competitive with Whisper large-v3 on supported languages. The v2 large model supports 30-second audio segments and handles code-switching better than Whisper, which tends to default to a dominant language. SeamlessM4T’s translation capability is unique among local models: it can transcribe audio in French, German, Hindi, or Japanese and simultaneously translate to English text, in a single model call. For multilingual enterprise deployments where translation is needed alongside transcription, SeamlessM4T avoids the latency and cost of a separate translation step.
Speaker Diarisation: Adding Pyannote
Standard STT models produce a transcript but do not identify who is speaking. Speaker diarisation — attributing each speech segment to a specific speaker — requires an additional model layer. Pyannote Audio is the standard open-source solution: it identifies speaker changes, clusters speech segments by speaker identity, and produces a diarised output that can be merged with a Whisper transcript to produce “Speaker A: …”, “Speaker B: …” formatted output. The combined pipeline (Whisper for transcription, Pyannote for diarisation, alignment to merge them) adds approximately 30–50% to processing time but produces meeting transcripts that are dramatically more readable when multiple speakers are involved. The main limitation is accuracy with more than four speakers and with heavily overlapping speech — Pyannote handles clean turn-taking well but degrades on simultaneous speech.
Vocabulary Adaptation
Whisper models sometimes struggle with domain-specific vocabulary — technical terms, proper nouns, product names — that appear rarely in training data. Two approaches address this. Prompt priming: Whisper accepts an initial prompt text that biases transcription toward certain vocabulary — including key terms in the prompt increases their probability in the output. For a medical transcription use case, priming with common medical terminology improves accuracy on those terms without any fine-tuning. Fine-tuning on domain audio: if you have 10–50 hours of labelled domain audio, fine-tuning Whisper small or medium on that data produces a model specialised for your domain. The Hugging Face Whisper fine-tuning guide provides a straightforward path using the Transformers library and a single GPU. Fine-tuning is most valuable when the domain vocabulary is highly specialised and prompt priming provides insufficient improvement.
The local STT landscape in 2026 offers a genuine choice between models optimised for different constraints. Faster-Whisper remains the default for batch accuracy; Parakeet is the speed leader for English streaming; Whisper.cpp opens up edge and mobile deployment; SeamlessM4T covers multilingual translation. Match the model to the scenario rather than defaulting to one model for everything — the right choice on each of the four deployment scenarios is different, and the wrong choice has real costs in either accuracy, latency, or hardware feasibility.
Evaluating on Your Own Audio
Published benchmark numbers — LibriSpeech clean, CommonVoice, Fleurs — are measured on standardised test sets that may not reflect your actual audio distribution. A model that ranks first on LibriSpeech clean can perform significantly worse than a lower-ranked model on your specific audio type if your recordings differ in accent, recording environment, or vocabulary. Before committing to a model for production, collect 30–60 minutes of representative audio samples from your actual use case (or a realistic simulation of it), manually transcribe a subset, and measure WER on each candidate model. Five minutes of properly labelled test audio for your domain is more informative than any published benchmark number. Include edge cases in your test set: heavy background noise, non-native accents, fast speech, and domain-specific terminology — these are where model rankings diverge most from published benchmarks and where the performance difference matters most for your users.
Cost Comparison: Local vs Cloud STT
Running local STT incurs hardware cost (GPU purchase or rental) but no per-request API fee. Cloud STT (OpenAI Whisper API, Google Speech-to-Text, AssemblyAI) charges per minute of audio — typically $0.006–$0.01 per minute. At 1,000 hours of audio per month, cloud STT costs $360–$600 per month indefinitely. A single RTX 3090 (approximately $900 new) running Faster-Whisper processes that volume easily and pays for itself within 2–3 months. The crossover point where local deployment becomes cheaper than cloud varies by volume and hardware utilisation, but for any application processing more than a few hundred hours of audio per month, local STT is almost always more cost-effective. The operational overhead — model management, hardware maintenance, uptime monitoring — is the real cost of local deployment and should be factored in, but for most teams already running GPU infrastructure for other purposes, adding STT workloads is a marginal cost.