Kaldi online2 Streaming Decoder and nnet3 Models: Adapt Features at Runtime or Fail on Pipeline Mismatch?
0 reputation · 16 Jul 2024, 18:32 UTC
0 reputation · 16 Jul 2024, 18:32 UTC
Kaldi's online2 decoder performs streaming recognition by extracting features on the fly, and it relies on the acoustic model having been trained with a matching feature configuration — the same feature type, dimensionality, and normalization treatment. When an nnet3 model is loaded, the feature options stored alongside it act as the contract the live pipeline must satisfy, and a mismatch aborts decoding rather than being silently tolerated.
The open integration question sits at this boundary. nnet3 recipes keep feature options in the model directory, but there is no runtime mechanism that reshapes or renormalizes incoming streaming features to satisfy a model whose pipeline differs — for example, when decoding uses binaries built from a different Kaldi revision, or normalization options the decoder build predates. A runtime adapter has been proposed as a graceful-degradation path, but whether it belongs in the decoder, and how it should behave, remains unresolved.
Specific questions:
A thoughtful contribution can make all the difference. Be the first to share one.
Use comments to ask for clarification. Post a solution as an answer.
No question comments on this page.