Rexx SUBSTR: 1-Based Indexing Is a Feature, Not an Off-by-One Bug
Rexx SUBSTR counts from 1, not 0. That makes fixed-width extraction arithmetic-free — and breaks ported code. Here's how to use it, plus the padding and byte-counting traps.
10 Mar 2026, 19:16 UTC

Port a fixed-width parser from Python to Rexx and the first test that fails is usually the simplest one. SUBSTR('Hello', 2, 3) returns ell, not llo. Rexx counts string positions from 1, and that single design choice is exactly what makes SUBSTR a good fit for column-oriented data — and exactly what breaks code copied from a zero-based language.
The thesis here is narrow: treat SUBSTR's 1-based, length-driven contract as the interface you design around, not as a quirk to work around. Once you do, fixed-width extraction stops needing arithmetic, and the remaining risks are concentrated in two places — padding edge cases and byte-versus-character counting.
What SUBSTR actually promises
The built-in function form is SUBSTR(string, start [, length] [, pad]), available in classic Rexx interpreters (z/OS TSO/E Rexx, Regina, ooRexx) and stable across them for decades.
- start is 1-based. Position 1 is the first character. There is no position 0.
- length is optional. Omit it and you get everything from
startto the end of the string. - The window is padded, not clipped. If
start + length - 1runs past the end of the string, the result is filled with the pad character, which defaults to a blank. - start past the end yields pad characters (or a null string when no pad is supplied). The exact boundary conditions are worth confirming on your interpreter rather than assuming.
One thing not to build on: a negative length. It is not a portable "trim from the end" idiom, and some interpreters reject it outright. Compute the length explicitly with LENGTH() instead of hoping the runtime interprets a negative value the way you meant.
A worked example: fixed-width records
Suppose each record is 26 characters: an 8-character ID field, an 8-character date, and a 10-character amount. Run this in any Rexx interpreter as a script or from the command line:
record = 'A1024 202601150000123.45'
id = SUBSTR(record, 1, 8)
date = SUBSTR(record, 9, 8)
amount = SUBSTR(record, 17, 10)
SAY STRIP(id) STRIP(date) STRIP(amount)
The expected output is A1024 20260115 0000123.45. The offsets are the layout itself: field one starts at column 1, field two at column 9, field three at column 17. No -1 corrections, no start:end slice pairs where you have to remember whether the end is exclusive.
Note that STRIP() is doing real work here — it removes the leading and trailing blanks that the padded ID field carries. SUBSTR extracts the window; it does not trim it. Keeping those two responsibilities separate makes the code easier to reason about when a field's width changes.
This example has not been executed as part of writing this article. Run it yourself and compare against the expected line before you trust it in a pipeline.
When PARSE beats SUBSTR
Rexx has a second, declarative way to slice a record, and for a fixed layout it is often the better choice:
PARSE VAR record id 9 date 17 amount
The numbers are absolute starting columns, and they are 1-based too. This one line reads like the layout specification, which is a real maintenance advantage when a non-programmer has to confirm the column positions.
The trade-off is flexibility. PARSE needs literal column numbers written into the source. SUBSTR takes an expression, so it wins when offsets are computed at runtime — variable-width fields, a layout table read from a config file, or a loop walking a record with a moving cursor. A reasonable rule: use PARSE for a layout that is fixed and known, and SUBSTR when the offsets are data.
Bytes, characters, and the Unicode caveat
This is the limitation most likely to bite silently. In byte-oriented Rexx interpreters, SUBSTR counts bytes, not characters. A UTF-8 string where one character occupies three bytes will be sliced mid-character, and you get garbled output rather than an error. Unicode-enabled builds count characters instead, so the same expression can behave differently across interpreters.
Before relying on character semantics, test it. Take a short multi-byte string, extract a known number of characters with SUBSTR, and inspect the result — if you see replacement characters or broken byte sequences, your interpreter is counting bytes and you need a different strategy for that data.
An actionable way to adopt this
Refactor one extraction at a time rather than rewriting a whole parser. For each field, keep a small golden sample of input records and the output you currently produce, replace the manual slicing with a single SUBSTR call, and diff the results. Pay particular attention to records shorter than the declared layout, since that is where padding behavior shows up.
Then verify the result with three checks: the first three characters of 'Hello' via SUBSTR('Hello', 1, 3) should be Hel; a short record should produce pad characters rather than a crash; and a multi-byte sample should round-trip intact. If all three pass on your interpreter, the 1-based contract is working for you instead of against you.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.