Optimizing Kaldi’s LF‑MMI: From Lattice Generation to Real‑World WER Gains
Learn how to reliably train Kaldi’s LF‑MMI model, from GMM‑based lattice generation to hyper‑parameter tuning, with a concrete example and practical diagnostics. Reduce WER, avoid common pitfalls, and get actionable steps for production pipelines.
04 Oct 2025, 23:56 UTC

The Concrete Problem – Why LF‑MMI Matters
When you finish a Kaldi recipe with a maximum‑likelihood (ML) neural net you often see a 5–10% relative WER drop after switching to lattice‑free MMI (LF‑MMI). The key question for any engineer is: how do I reliably get that boost without drowning in hyper‑parameter noise?
LF‑MMI replaces the ML objective with a discriminative loss that looks at phone‑level probabilities over lattices produced by a GMM‑HMM system. Because the lattices carry the HMM’s decision boundaries, their quality directly influences the neural net’s learning path. A weak GMM can seed a bad lattice and make the network converge to a worse model than the baseline.
Building the Pipeline – From GMM to Neural
Kaldi’s nnet3-lfmi-train script automates the sequence:
- Train a GMM‑HMM with
run.shto getfinal.mdland alignments. - Generate lattices with
lattice-faster-mapped-gmmorlattice-faster-gmm. - Run
nnet3-lfmi-trainto optimize the neural net on those lattices.
Because the GMM drives the lattice quality, it is worth spending a few hours on a clean, well‑tuned GMM before launching LF‑MMI. A common pitfall is reusing a GMM that was trained on a very small subset of data – the lattices will be sparse and the MMI gradient will be noisy.
Tuning the Objective – Hyperparameters That Pay Off
LF‑MMI in Kaldi exposes several knobs that interact strongly with the neural architecture:
- Learning‑rate schedule – The default exponential decay often works, but for TDNNs a linear decay from
1e-3to1e-4across 20 epochs gives more stable phone‑accuracy curves. - Weight‑tying – Enabling weight sharing between the output layer and the last hidden layer (
--weight-tying=true) reduces parameters by ~30% and can improve generalisation. - iVector adaptation – Adding speaker vectors to the input (via
--ivector-dir) can reduce WER by 1–2% on noisy corpora, but increases per‑example memory. - L2 regularisation – In Kaldi 6.x the flag changed from
--l2-regularizeto--l2-regularization. A value of1e-4is a good starting point for LSTMs.
Below is a minimal conf/train_lfmi.conf snippet that reflects these choices for a TDNN:
# train_lfmi.conf
--learning-rate 1e-3
--learning-rate-decay 1e-4
--weight-tying true
--ivector-dir exp/ivectors
--l2-regularization 1e-4
--num-epochs 20
--batch-size 128
Running the Training
Execute from the recipe directory:
./run.sh --stage 6 # triggers nnet3-lfmi-train
Permissions: you need write access to exp and read access to data. The script will create exp/tdnn_1b_lfmi containing the checkpoints.
Diagnosing a Stuck Model – What to Inspect
After the first epoch you can sanity‑check the training progress by looking at log/progress.bar and log/train_lfmi.log:
- Phone accuracy per epoch – Should rise steadily. A flat line suggests a bad lattice or too high a learning rate.
- Lattice likelihoods – Inspect the
log/lattice_likelihoods.txtfile. Values that are consistently very low (< -10) indicate the network is not learning the correct phone boundaries. - Gradient norms – Extremely high or low values can reveal exploding or vanishing gradients. If you see
grad_norm=0.0, your batch size may be too large for the GPU memory, causing truncation.
When you spot an issue, a quick fix is often to reduce the batch size or adjust the learning‑rate schedule. If the lattice likelihoods are the problem, revisit the GMM training: add more Gaussians or increase the number of EM iterations.
Trade‑offs and Limitations
- Memory and I/O – LF‑MMI requires storing full lattices on disk. On a 100‑hour corpus this can exceed 50 GB. Staging the feature extraction and lattice sharding (e.g.,
--lattice-dir exp/lattices/part-01) can mitigate peak disk usage. - Sequential pipeline – The GMM → lattice → LF‑MMI steps cannot run in parallel, which lengthens the overall training time. A common practice is to pre‑compute lattices on a separate machine and transfer them via
scp. - Domain shift – If you move to a new acoustic domain (e.g., from clean to noisy), the GMM may not model the new data well, and the lattices will be unreliable. In that case, retrain the GMM with more data or add data‑specific Gaussian splits.
Actionable Checklist
- Train a high‑quality GMM (≥ 2000 Gaussians, 5 EM iterations).
- Generate lattices with
lattice-faster-gmmand inspectlattice_likelihoods.txt. - Configure
train_lfmi.confwith learning‑rate schedule, weight‑tying, iVector dir, and L2 regularisation. - Run
nnet3-lfmi-trainand monitorprogress.barfor steady phone‑accuracy improvements. - Evaluate on a held‑out dev set; compare WER to the ML baseline. If the WER does not improve, revisit lattice quality or hyper‑parameters.
- Optional: experiment with batch‑size reduction or gradient clipping if you hit GPU memory limits.
By following this disciplined approach you can reliably harness LF‑MMI’s discriminative power while keeping training time and resource usage in check.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.