Using Rexx PARSE for Reliable String Tokenization
Learn how the PARSE instruction simplifies tokenizing strings in Rexx scripts, improves readability, and works across implementations, with a worked example and notes on limitations.
13 Aug 2025, 17:56 UTC

The problem with manual tokenization
When a Rexx script needs to split a line into fields, developers often chain SUBSTR and POS calls. This approach is verbose, easy to mis‑count offsets, and prone to off‑by‑one errors, especially when the input format changes.
Why PARSE is a better engineering choice
The PARSE instruction consolidates tokenization into a single statement. By supplying a template that describes positional, keyword, or pattern‑based slots, PARSE assigns the matching parts to variables in one step. This reduces script length, improves readability, and centralises the parsing logic, making it easier to review and maintain.
Worked example: parsing a comma‑separated line
Consider a line that holds three fields: a name, an age, and a city, separated by commas. The following Rexx fragment shows how PARSE extracts each field into a variable.
/* parse_example.rexx */
line = 'Alice,30,New York'
PARSE VAR line name ',' age ',' city .
/* name, age, city now hold the three fields */
SAY 'Name:' name
SAY 'Age:' age
SAY 'City:' city
The template "name ',' age ',' city ." tells PARSE to take everything up to the first comma into NAME, then skip the comma, take up to the next comma into AGE, skip that comma, take the rest into CITY, and ignore any trailing characters (the final period). When you run the script, the SAY instructions will display the assigned values.
Checking the result across implementations
To verify that PARSE behaves the same in different Rexx engines, you can execute the same script under Regina and ooRexx.
- Save the script to a file, e.g.,
parse_example.rexx. - Run it with Regina:
rexx parse_example.rexx(Regina’s interpreter is usually invoked asrexx). - Run it with ooRexx:
rex parse_example.rexx(ooRexx provides therexcommand). - In each case, inspect the output of the three SAY statements. If the lines appear identical, the PARSE instruction produced the same variable assignments.
No special privileges are required; the commands need only read access to the script file and execute permission on the interpreter. The risk is minimal, but if the input line does not match the template, placeholders may receive empty strings, which could silently propagate incorrect data.
Limitations and practical mitigation
- Complex templates – A PARSE template with many nested patterns can become hard to read. Keep templates simple; if the logic grows, consider breaking the parsing into multiple PARSE steps or moving the logic to a subroutine.
- Unicode support – Older Rexx releases treat PARSE as byte‑oriented. When processing multilingual text, convert the string to a known encoding (e.g., UTF‑8) before parsing, or use a newer build that offers Unicode awareness.
- Mismatched templates – If the template expects more delimiters than present, extra variables get empty values. Always validate critical fields after parsing, for example by checking that a variable is not empty or matches an expected pattern.
To confirm that your parsing succeeded, add a quick validation step after the PARSE block:
IF name = '' THEN DO
SAY 'Error: name field missing'
EXIT 1
END
Actionable closing
Adopting PARSE for string tokenization makes Rexx scripts shorter, clearer, and more portable across Regina, ooRexx, and IBM Object Rexx. Start by replacing ad‑hoc SUBSTR/POS chains with a single PARSE VAR statement, keep the template straightforward, and validate the results. This change reduces the chance of off‑by‑one bugs and eases future maintenance.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.