Switching to Kaldi’s Chain Models: TDNN‑F with LF‑MMI for Faster, Accurate ASR
Learn how Kaldi’s chain models with TDNN‑F and lattice‑free MMI replace legacy DNN/HMM hybrids, cutting training time and model size while keeping or improving word error rates.
29 Oct 2025, 05:00 UTC

The problem: legacy DNN/HMM hybrids are slow and hard to scale
If you have been training ASR systems with Kaldi’s older nnet1/nnet2 recipes, you likely notice two pain points: long training times because each iteration requires generating lattices for the MMI objective, and a tangled topology that mixes a separate transition model (GMMs) with the neural network. Adding more data or trying a new language often means rebuilding the alignment step from scratch.
Thesis: the chain model unifies acoustics and simplifies training
Kaldi’s chain approach replaces the hybrid DNN/HMM stack with a single neural network that directly outputs phone posteriors and a transition model. The lattice‑free MMI (LF‑MMI) loss is computed on the denominator graph without materializing lattices, making the objective GPU‑friendly. When combined with a factorized TDNN (TDNN‑F) layer, the model shrinks by 30‑50 % while preserving accuracy.
What changed under the hood
Chain topology
In a chain model the neural network predicts a sequence of phone‑level outputs that are already aligned to a denominator finite‑state transducer (FST). There is no separate GMM‑HMM alignment step; the transition probabilities are absorbed into the network’s output layer.
TDNN‑F factorization
A standard TDNN layer applies a full weight matrix across time‑delayed frames. TDNN‑F replaces that matrix with two low‑rank matrices (W₁ × W₂) where the inner dimension is the rank. This cuts parameters dramatically; typical recipes use ranks of 32‑64 for English.
LF‑MMI objective
Instead of first generating lattices and then computing the MMI loss, LF‑MMI evaluates the numerator and denominator directly on the training minibatch. The denominator graph is built once from the phone lexicon and language model, and the loss is back‑propagated through the network. This reduces memory usage and lets you leverage GPU parallelism.
Worked example: training a TDNN‑F chain model on WSJ
Assume you have Kaldi installed at $KALDI_ROOT and you have write access to the egs/wsj/s5 directory.
Navigate to the WSJ example:
cd $KALDI_ROOT/egs/wsj/s5Run the chain recipe, skipping data preparation if already done (stage 0‑12) and starting at chain training (stage 13). Adjust
--gpuif you have a GPU../run.sh --stage 13 --train-stage -10 --gpu 1Required permissions: read access to the corpus, write access to
exp/. The script will create directories likeexp/chain/tdnn_f_1a_sp.Monitor training progress. Look for the LF‑MMI loss decreasing and the learning‑rate schedule being applied (see
exp/chain/tdnn_f_1a_sp/log/train.log). Divergence or NaN values indicate that the learning‑rate or regularization needs tuning.After training finishes, decode the test set:
steps/nnet3/decode.sh --acwt 1.0 --post-decode-acwt 10.0 \ exp/chain/tdnn_f_1a_sp/graph_tgsmall data/test_dev93 \ exp/chain/tdnn_f_1a_sp/decode_test_dev93_tgsmallCheck the word error rate (WER) in the scoring file:
cat exp/chain/tdnn_f_1a_sp/decode_test_dev93_tgsmall/scoring_kaldi/best_werCompare this number to the WER reported for the nnet2 baseline in
RESULTS(typically around 8‑9 % on WSJ). You should see equal or lower WER with a smaller model file (exp/chain/tdnn_f_1a_sp/final.mdl).
Trade‑offs and limitations
LF‑MMI is sensitive to learning‑rate scheduling, L2 regularization, dropout, and frame‑subsampling. The provided recipes embed heuristics that work for English; for a new language or noisy domain you may need to rerun a short learning‑rate finder or adjust
--l2-regularizeand--dropout-schedule.The chain model assumes a fixed phone set and lexicon. Adding new words requires rebuilding the denominator FST and, in most cases, retraining or at least fine‑tuning the acoustic model. This is less flexible than pure end‑to‑end systems that can simply rescore with an external language model.
TDNN‑F rank is a hyper‑parameter. Too low a rank (e.g., <8) can hurt accuracy; too high a rank erodes the compression benefit. The default ranks in the WSJ recipe (32‑64) are a good starting point, but you should validate on a held‑out set for your target language.
While the neural network forward/backward pass runs on the GPU, data preparation, i‑vector extraction, and FST operations remain CPU‑bound. Expect overall training speed‑ups of roughly 2× when moving from nnet2 to chain on a modern GPU, but not an order‑of‑magnitude gain.
Practical verification steps
Model size:
ls -lh exp/chain/tdnn_f_1a_sp/final.mdl. Compare with the nnet2 model (exp/nnet2_online/nnet/final.mdl).Real‑time factor (RTF): decode a short wav file with the online chain decoder and measure wall‑clock time versus audio length.
online2-wav-nnet3-latgen-faster --frame-subsampling-factor=3 \ --online=true --do-endpointing=false \ exp/chain/tdnn_f_1a_sp/final.mdl \ exp/chain/tdnn_f_1a_sp/graph_tgsmall/HCLG.fst \ 'ark:echo utterance-id1 utterance-id1|' \ 'scp:echo utterance-id1 /path/to/audio.wav|' \ ark:/dev/nullCheck the printed elapsed time; RTF < 1 indicates real‑time capable decoding on your CPU.
Actionable closing
If you are starting a new ASR project or looking to replace an aging nnet2 system, try the chain recipe as a baseline:
Run the WSJ chain recipe (or the equivalent for your language) to confirm that the build works and that you can reproduce the reported WER.
Inspect the training log for loss convergence; if you see instability, adjust the learning‑rate schedule or increase dropout before scaling up to more data.
Measure model size and RTF on your target hardware; if the model is still too large, experiment with a lower TDNN‑F rank and re‑evaluate WER.
When you need to add vocabulary, plan to rebuild the denominator FST and consider a short fine‑tuning pass rather than training from scratch.
By adopting the chain model with TDNN‑F and LF‑MMI you gain a simpler topology, faster GPU‑friendly training, and decoding performance that matches or exceeds the older hybrid approach—without sacrificing the flexibility to tune for your specific domain.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.