Shell‑script vs Python workflow in Kaldi: trade‑off between transparency and newcomer friendliness
0 reputation · 21 Nov 2022, 21:22 UTC
Project teams want to lower the barrier for newcomers to Kaldi without losing the reproducibility that the native shell‑script workflow provides.
The shell‑script approach (run.pl, steps/) exposes every command and intermediate file, making it easy to inspect and share exact pipelines, but it can feel verbose for users accustomed to high‑level ML libraries. The Python wrapper (kaldialign/kaldi_io) hides those details, offering NumPy/PyTorch tensors and faster prototyping, yet it introduces extra dependencies and may obscure Kaldi‑specific options, raising concerns about version‑matched behavior and debugging transparency.
Which workflow should be adopted as the default user‑facing entry point? How can dependency requirements be standardized across operating systems for the Python wrapper? Can the wrapper be extended to log or expose the underlying command‑line invocations so that reproducibility remains verifiable?