Online decoding beam width limit for low‑latency Kaldi streaming
27K reputation · 08 Mar 2024, 06:03 UTC
Goal
Determine a beam‑width setting for Kaldi’s online decoding that guarantees sub‑100‑ms latency on commodity CPUs while preserving acceptable word‑error‑rate (WER).
Unresolved Constraint
Kaldi’s online decoder exposes a user‑configurable beam width but does not document an upper bound that ensures low latency. The current implementation applies a static beam width; increasing it improves accuracy but can cause significant latency and memory spikes.
Memory Considerations
Memory consumption for large beam widths is not linear and depends on the number of active states. Empirical measurement is required to understand the trade‑off between beam width, latency, and peak memory usage.
Specific Questions
- What beam‑width threshold yields <100 ms end‑to‑end latency on a typical 4‑core Intel CPU for standard Kaldi acoustic models?
- Should we implement a fixed maximum beam width or a dynamic adjustment that reacts to CPU load or network conditions?
- How does beam width affect peak memory usage for streaming decoding, and what profiling tools best capture this relationship?