awk length() and substr() Limits Under Byte and Character Locales
0 reputation · 17 Sept 2024, 04:35 UTC
POSIX awk defines string functions such as length(), substr(), index() and match() in terms of characters when the active locale uses a multibyte encoding, and in terms of bytes otherwise. Field splitting with a single-character FS follows the same locale-sensitive path in implementations built with multibyte support.
That makes output reproducibility a design choice rather than a fixed property. A report generator that pads or truncates columns using length() can align differently under LC_ALL=C and under a UTF-8 locale, and the difference also affects the text a screen reader announces. Multibyte handling depends on how the installed binary was compiled, not only on its release number, so the standard alone does not predict behavior. Documented behavior therefore needs confirmation against each implementation's own manual and build options.
The unresolved decision: pin a fixed locale for byte-exact, reproducible output, or accept locale-aware character semantics for correctness with non-ASCII input.
- Which of the two goals, byte-exact reproducibility or character-correct counting, should take precedence in a user-facing awk pipeline?
- Should display-width formatting be delegated to a tool with an explicit width model instead of awk's locale-dependent
length()? - Can a workflow detect at runtime whether the installed awk honors multibyte locales?