Speaker‑Adaptive Chain Model Training in Kaldi with i‑Vector Augmentation
Learn how to augment Kaldi chain models with speaker i‑vectors: mechanism, a complete configuration example, limits, and common pitfalls to avoid.
19 Mar 2026, 17:40 UTC

To improve speaker adaptation in a Kaldi chain model, append speaker i‑vectors to each frame’s feature vector before feeding the data to the TDNN‑F network. This gives the discriminative chain objective direct access to speaker information while keeping the training pipeline compatible with standard Kaldi scripts.
How the mechanism works
The chain framework replaces the traditional cross‑entropy objective with a sequence‑level loss that directly optimizes phone‑level accuracy. During training, a numerator graph (derived from the forced‑alignment transcription) and a denominator graph (a phone‑level language model) are built for each utterance. The acoustic model scores each frame; the i‑vector, which encodes the speaker characteristics of the recording, is concatenated to the MFCC‑based feature vector at every frame. Because the i‑vector is constant across an utterance, the network learns to modulate its internal activations based on speaker identity, yielding better generalization especially when test speakers differ from the training set.
Worked configuration example
Below is a minimal, reproducible snippet that shows where to place the i‑vector step in a typical WSJ recipe. Assume you have already extracted MFCC+CMVN features and trained a GMM tri3b model.
# 1. Train i‑vector extractor (run once)
steps/online/nnet2/train_ivector_extractor.sh \
--cmd "$train_cmd" \
--nj 8 \
data/train \
lang \
exp/nnet3/ivector_extractor
# 2. Extract i‑vectors for the training set
steps/online/nnet2/extract_ivectors_online.sh \
--cmd "$train_cmd" \
--nj 8 \
data/train \
exp/nnet3/ivector_extractor \
exp/nnet3/ivectors_train
# 3. Train the chain TDNN‑F model with i‑vector augmentation
steps/nnet3/chain/train.py \
--stage 0 \
--train-set train \
--gmm tri3b \
--nnet3-affix \
--ivector-dir exp/nnet3/ivectors_train \
--ivector-dim 100 \
--lda-dim 40 \
--hidden-dim 512 \
--minibatch-size 128 \
--effective-lrate 0.001 \
--num-jobs-initial 2 \
--num-jobs-final 4 \n --cmd "$train_cmd" \
exp/nnet3/chain/tdnnf_sp
Where to run: execute the commands in the Kaldi root directory (e.g., /opt/kaldi/egs/wsj/s5). You need write permission to the exp/ subtree to store models, logs, and i‑vector archives. The --ivector-dim flag must match the dimensionality used when training the extractor (see step 1).
Limits and practical considerations
- GPU memory: The denominator lattice generation for chain training can consume >8 GB VRAM on a medium‑sized WSJ setup; reduce
--minibatch-sizeor use gradient checkpointing if your GPU is smaller. - System RAM: Storing lattices for all utterances may require tens of gigabytes; consider splitting the training set into chunks with
--num-jobs-initialand--num-jobs-final. - Learning‑rate schedule: The default schedule in
train.pyworks for many recipes, but if the objective plateaus early, lower the initial--effective-lrateor increase the dropout via--dropout-schedule. - I‑vector consistency: The extractor must be trained on features with the same sample rate and bandwidth as the MFCCs (usually 16 kHz, 80 Hz‑8 kHz). Mismatched bandwidth leads to misaligned speaker information and degraded WER.
Common mistakes and how to avoid them
- Mismatched i‑vector dimension: Setting
--ivector-dim 130while the extractor outputs 100‑dim vectors causes a shape error in the network’s first layer. Fix: verify the extractor’s output withivector-extractor --dimand match it exactly. - Skipping re‑alignment after epochs: The chain trainer expects frame‑level alignments that stay in sync with the subsampled TDNN‑F outputs. If you change the subsampling factor (
--frame-subsampling-factor) without regenerating alignments, the denominator graph timing breaks. Solution: runsteps/align_fmllr.shorsteps/nnet3/align.shafter any architecture change. - Using speed‑perturbed data without updating i‑vectors: When you apply speed perturbation (e.g.,
steps/data/augment_data.sh), the speaker characteristics remain the same but the utterance length changes. If you extract i‑vectors only on the original data and reuse them, the i‑vectors will be mis‑aligned with the perturbed frames. Correct approach: either extract i‑vectors on each perturbed copy or apply the same perturbation to the i‑vector archive usingsteps/online/nnet2/copy_ivector_dir.sh --perturb true.
Quick verification steps
Before committing to a full training run, perform a sanity check on a small subset:
- Create a 10‑minute subset:
utils/subset_data_dir.sh --shortest data/train 1000 data/train_10min - Extract i‑vectors for this subset (reuse the extractor).
- Run one epoch of chain training:
steps/nnet3/chain/train.py --stage 0 --train-set train_10min ... --num-epochs 1 exp/nnet3/chain/tdnnf_sp_small - Inspect the training log; the objective should increase (or at least not decrease) each iteration.
- Decode a held‑out dev set with
steps/nnet3/decode.sh --nj 4 exp/tri3b/graph data/dev exp/nnet3/chain/tdnnf_sp_small/decode_devand compare WER to a GMM baseline. A relative reduction of ≳10 % indicates the i‑vector‑augmented chain model is learning.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.