Using Rexx PARSE to Tokenize Input Without Regular Expressions
Learn how Rexx’s PARSE instruction can tokenize input in a single readable statement, with a worked example, verification steps, and notes on limitations.
16 Aug 2025, 00:19 UTC

Problem: Splitting a line into meaningful parts
When processing text files, logs, or user input in a Rexx script, you often need to break a string into separate pieces—for example, extracting a date, a timestamp, a severity level, and a message from a log line. Writing a loop with SUBSTR, INDEX, or manual character counting works, but it quickly becomes hard to read and maintain, especially as the format evolves.
Thesis: PARSE gives a single‑statement, readable way to tokenize
The Rexx PARSE instruction can split a string using whitespace, delimiters, or positional templates and assign the results to variables in one line. This reduces boilerplate, improves clarity, and works consistently across major Rexx implementations (IBM Object Rexx, Regina, ooRexx).
How PARSE works
PARSE takes a source string and a template. The template can contain:
- Variable names that receive the next whitespace‑delimited token.
- Literal strings or delimiters that must match exactly.
- Pattern keywords such as
WORD,NUMBER,LENGTH, or a numeric length specification.
When the instruction runs, Rexx scans the source left‑to‑right, matches each template element, and stores the matched substring in the corresponding variable. If a part of the template fails to match, the associated variable receives a null string.
Worked example: parsing a log entry
Consider a simple log line with the format YYYY-MM-DD HH:MM:SS LEVEL Message. The following script shows how PARSE extracts each component:
/* rexx_logparse.rex */
/* Sample log line */
line = '2026-10-09 14:32:05 INFO User logged in'
/* Template: date, time, level, rest of line as message */
PARSE VAR line date time level message
/* Display the results */
say 'Date :' date
say 'Time :' time
say 'Level :' level
say 'Message:' message
To run the script:
- Save the code to a file, e.g.,
rexx_logparse.rex. - Ensure you have a Rexx interpreter installed (Regina or ooRexx are common open‑source choices).
- Execute the file with the interpreter:
rexx rexx_logparse.rex(on Unix‑like systems) orrexx.exe rexx_logparse.rexon Windows. - No special permissions are needed beyond read access to the file and execute permission for the interpreter.
Expected output (verification):
Date : 2026-10-09 Time : 14:32:05 Level : INFO Message: User logged in
If the output matches the expected values, the PARSE instruction has correctly tokenized the input.
Trade‑off: flexibility versus simplicity
PARSE excels at straightforward, predictable formats. It does not support arbitrary pattern matching like regular expressions, so complex or variable‑width delimiters may require additional logic or a different approach. Moreover, the treatment of multibyte or Unicode characters depends on the interpreter’s locale settings; some versions operate on bytes rather than characters, which can cause unexpected splits when processing UTF‑8 data. To verify behavior with non‑ASCII text, run the same script with a line containing accented characters and compare the results across interpreters.
Actionable closing
For most everyday tokenization tasks—splitting command‑line arguments, parsing configuration lines, or extracting fields from structured logs—the PARSE instruction offers a clear, maintainable solution without the overhead of regular‑expression libraries. Start by replacing manual SUBSTR loops with a PARSE statement, test with a few representative inputs, and adjust the template as the format evolves. If you encounter locale‑specific issues, consult your interpreter’s documentation on character handling or consider normalizing input to a known encoding before parsing.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.