Using Kaldi’s Online Decoder for Real‑Time Speech Recognition
Learn how to set up Kaldi’s online decoding pipeline, run a worked example with the tedlium2 recipe, and verify latency and accuracy trade-offs.
27 Nov 2025, 13:21 UTC

Problem: transcription while speech is arriving
Live captioning, voice control and call-center analytics need text as audio arrives. Waiting for a full utterance adds latency that breaks the use case. Kaldi provides online decoding that processes feature vectors frame by frame and emits a lattice as soon as evidence accumulates.
Thesis
With an online-ready nnet3 model and tuned beam settings, Kaldi’s online decoder can run with real-time factor below 1.0 on a modern CPU while keeping word error rate usable for streaming.
How online decoding works in Kaldi
The online pipeline mirrors offline decoding but keeps running decoder state.
- Feature extraction – compute MFCC or filterbank vectors with compute-mfcc-feats or compute-fbank-feats. Run from the recipe directory with read access to wav.scp.
- CMVN – apply cepstral mean and variance normalization using a sliding window or global stats.
- Online decoder – online2-wav-nnet3-latgen-faster for nnet3 models, or latgen-faster-mapped in online mode for GMM. It updates token passing per frame and writes a lattice.
- Output – lattice-best-path converts the lattice to words.
The model must be built with the --online flag. Using a standard offline model fails at initialization with a missing online component error.
Worked example with tedlium2
After training up to step 12 in egs/tedlium2/s5 you have exp/nnet3/tdnn_sp/model.mdl, exp/nnet3/tdnn_sp/HCLG.fst and conf/online.conf.
Run from s5 with read access to the model and audio:
online2-wav-nnet3-latgen-faster --frame-subsampling-factor=3 --config=conf/online.conf exp/nnet3/tdnn_sp/model.mdl exp/nnet3/tdnn_sp/HCLG.fst 'ark:echo $utt_id|' 'scp:pipeline:wav scp,$utt_id $wav_path|' 'ark,t:-'$utt_id is the utterance identifier and $wav_path is the full path to a 10 second wav. The command feeds features, updates decoder state, and writes a lattice to stdout. On a typical laptop CPU the run finishes in under a second for a 10 second clip.
Best path extraction:
lattice-best-path --acoustic-scale=0.1 ark:- ark,t:- | utils/int2sym.pl -f 2- data/lang/words.txtCompare against the tedlium2 reference. Clean speech typically stays below 15 percent WER with the recipe defaults.
Trade-offs and verification
Online decoding cannot use future context, so accuracy is lower than batch decoding. Beam and lattice-beam control pruning. Low values reduce CPU and latency but increase search errors. High values improve accuracy but push RTF toward or above 1.0.
Verification steps:
- Check logs for LOG Online2WavNnet3LatgenFaster:Processing utterance and confirm RTF below 1.0.
- Generate an offline lattice with latgen-faster-mapped on the same utterance and compare arc counts. The online lattice should be a subset.
- Run on silent audio. The decoder should output an empty transcription rather than crash.
GPU acceleration is limited for the online nnet3 decoder. Most deployments run on CPU, so many concurrent streams require multiple cores or a thread pool. Variable-rate audio or large buffered chunks can cause overruns.
Actionable closing
- Build the acoustic model with --online and generate HCLG.fst.
- Start with beam around 13.0 and lattice-beam around 6.0 from the tedlium2 online script, then monitor RTF.
- Feed audio chunks via a lightweight service and extract best-path from the lattice.
- Log RTF and WER on a validation set to catch drift from audio conditions or model updates.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.