Balancing Latency and Flexibility: Static vs. Dynamic Decoding in Kaldi
Explore the trade‑offs between Static and Dynamic decoding in Kaldi. Learn how HCLG graphs impact ASR latency, memory usage, and grammar flexibility.
08 Aug 2025, 16:42 UTC

The Search Space Dilemma
When deploying an Automatic Speech Recognition (ASR) system, you face a fundamental trade‑off: do you pre‑calculate every possible word sequence your system might encounter, or do you calculate them on the fly? In Kaldi, this is the choice between Static Decoding (using a pre‑compiled HCLG graph) and Dynamic Decoding (composing the grammar at runtime).
The takeaway is simple: if your vocabulary and grammar are fixed, static decoding provides the lowest possible latency. If your application requires a changing set of keywords or a dynamic user‑specific vocabulary, dynamic decoding is the only viable path, despite the increased CPU overhead.
The HCLG Architecture
Kaldi uses Weighted Finite State Transducers (WFSTs) to map acoustic signals to words. To do this, it composes four distinct transducers into one massive search graph:
- H (HMM): Maps HMM states to transitions.
- C (Context‑dependency): Maps context‑dependent phones to mono‑phones.
- L (Lexicon): Maps phones to words.
- G (Grammar): Maps words to sequences based on a language model (usually an N‑gram).
When these are composed into an HCLG graph, the decoder doesn’t have to \"think\" about the grammar during inference; it simply traverses a pre‑optimized path of weights to find the most likely sequence.
Static Decoding: The High‑Throughput Path
In static decoding, the HCLG graph is built offline. This process involves determinization (removing redundant paths) and minimization (reducing the number of states). Because the graph is already optimized, the Viterbi decoder can move through the search space with minimal computation.
This is ideal for production environments where the domain is narrow—such as a voice‑command system for a microwave or a standardized medical transcription tool—where the language model rarely changes.
Dynamic Decoding: Handling Fluid Grammars
Dynamic decoding skips the full offline composition. Instead, it composes the G transducer (the grammar) with the HCL components at runtime. This allows you to inject new words or change the probability of certain phrases without rebuilding the entire system.
The cost is computational. Composing transducers during the decoding process increases CPU usage and introduces latency, as the system must perform operations that were previously handled during the offline build phase.
Practical Example: Building a Static Graph
To implement static decoding, you typically use the utils/compose_lm.sh script. This script handles the heavy lifting of composing the H, C, L, and G transducers.
# Run from the kaldi directory as a user with write permissions to the data folder
# Required: L.fst, G.fst, and the H and C transducers already built
utils/compose_lm.sh data/train/lang.fst data/train/lexicon.txt model/final.fst graph/HCLG.fst
Expected Check: After execution, verify the size of graph/HCLG.fst. If the file size is unexpectedly small, the composition may have failed; if it is several gigabytes, you may need to increase your system RAM for the decoder.
Risk: Running this on a massive N‑gram model without sufficient memory can lead to OOM (Out of Memory) crashes during the determinization phase.
Trade‑offs and Limitations
| Feature | Static (HCLG) | Dynamic |
|---|---|---|
| Inference Speed | Very Fast | Slower |
| Memory Usage | High (loads full graph) | Lower (loads components) |
| Flexibility | Rigid (requires rebuild) | Flexible (runtime updates) |
| Setup Time | Long offline build | Minimal offline build |
A critical limitation of the static approach is the RAM ceiling. For massive vocabularies (e.g., general‑purpose web speech), the HCLG graph can exceed available system memory, forcing a move toward dynamic decoding or more aggressive pruning of the G transducer.
Verification and Result Checking
To verify that your chosen decoding strategy is working correctly, use the src/latgen2 tool. This tool generates a lattice from the graph and allows you to check if the most likely paths align with your expected transcriptions.
If you notice the system is consistently missing keywords that are present in your G transducer, check your weight pushing. Weight pushing ensures that the lowest weights are encountered as early as possible in the search, allowing the beam search to prune unlikely paths more efficiently without losing the correct result.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.