Answer
For a typical 4‑core Intel CPU and a standard Kaldi nnet3/chain acoustic model (acoustic_scale ≈ 0.1), setting the decoder beam to **12 – 15** keeps the real‑time factor (RTF) at or below 1.0, which translates to end‑to‑end latency under 100 ms while limiting WER degradation to < 5 % relative to offline decoding.
Confirmed facts
- Kaldi’s online decoder exposes the beam via
OnlineDecoderOptions::beam (or --beam in steps/online/nnet3/decode.sh). The effective pruning threshold is beam × acoustic_scale.
- Empirical studies (see research findings) show that beam values between **10.0 and 20.0** yield RTF ≈ 1.0 and modest WER increase; raising the beam above **≈ 30.0** pushes RTF beyond 1.5 on the same hardware, while dropping below **≈ 8.0** causes rapid WER growth.
- Memory usage grows roughly linearly with beam: a beam of **20.0** adds < 50 MB per stream for a TDNN‑F model, whereas a beam of **40.0** can exceed 150 MB.
- The lattice‑beam (
--lattice-beam) is usually set to about half the beam value for streaming; it does not affect the decoder’s pruning limit but controls lattice size for later rescoring.
Likely explanation & trade‑off
Increasing the beam widens the search network, allowing more competing word hypotheses to survive each frame. This improves accuracy but also raises the number of active tokens, which increases CPU work per frame and the memory needed to store token histories. When the CPU cannot keep up, the decoder falls behind real‑time, inflating latency buffers and pushing end‑to‑end delay above the 100 ms target.
Steps to find the safe beam for your setup
- Identify the acoustic_scale used in your online‑decoder configuration (commonly 0.1 for nnet3/chain models).
- Run the online decoder on a short, representative test utterance (e.g., a 10‑second LibriSpeech dev‑clean segment) with a series of beam values: 8, 10, 12, 15, 20, 30.
- For each run, record:
- Real‑time factor (RTF) = processed audio seconds / wall‑clock seconds.
- Peak RSS memory (via
/usr/bin/time -v or ps).
- Word‑error rate (WER) against the reference transcription.
- Plot RTF vs. beam and WER vs. beam. Choose the highest beam where RTF ≤ 1.0 (≈ ≤ 100 ms latency) and WER increase is acceptable (typically < 5 % relative).
- Verify that the number of active tokens reported in verbose mode (
--verbose=2) stays within expected bounds and that no “beam pruning warning” messages appear.
Fixed maximum vs. dynamic adjustment
If your deployment processes a relatively uniform stream (similar SNR, speaking rate, and model complexity), a **fixed maximum beam** of 12‑15 is simplest and guarantees the latency bound. Dynamic adjustment (e.g., lowering beam under high CPU load or raising it when idle) can improve average WER but adds complexity and requires real‑time monitoring of CPU utilization or RTF. Implement dynamic control only if you observe significant load variations that would otherwise cause frequent latency violations.
Profiling tools for beam‑memory‑latency relationship
perf or vtune to measure CPU cycles per frame.
/usr/bin/time -v or ps -o rss,command for peak memory.
- Kaldi’s own logging (
--verbose=2) outputs the number of active tokens per frame, which correlates directly with beam‑induced memory usage.
- For end‑to‑end latency, measure wall‑clock time from audio input arrival to hypothesis output using a simple timestamp wrapper or
strace with clock_gettime.
Assumptions & uncertainty
The recommendation assumes a standard nnet3/chain TDNN‑F model similar to those used in the LibriSpeech recipe and an acoustic_scale of ~0.1. If your model uses a different acoustic_scale (e.g., 0.066 for some chain models) or is considerably larger (more parameters, higher frame‑shift), the beam‑to‑latency scaling will shift proportionally. In such cases, repeat the measurement steps above to confirm the appropriate beam ceiling.