Parsing Structured Text Files in Rexx Using Template Patterns
Learn how to use the Rexx PARSE instruction to extract data from delimited and fixed-width text files using template patterns, stem arrays, and validation checks.
09 Feb 2026, 00:59 UTC

The Problem: Extracting Data from Rigid Text Formats
Processing legacy logs, CSVs, or fixed-width data files often requires splitting a single string into multiple usable variables. While many languages rely on complex regular expressions or split functions that return arrays, Rexx uses the PARSE instruction. The challenge lies in designing a template that accurately captures data while handling delimiters or specific column offsets without crashing the script when a line is malformed.
Prerequisites
- A Rexx interpreter installed (such as ooRexx or Regina).
- A text file with a consistent record structure (either delimited or fixed-width).
- Read permissions for the target data file.
Implementing Delimiter-Based Parsing
When data is separated by a character (like a comma or pipe), the PARSE instruction uses literal strings in the template to identify where one field ends and the next begins.
/* Example: Parsing a CSV line: "John,Doe,30,Engineer" */
line = "John,Doe,30,Engineer"
PARSE VAR line firstName ',' lastName ',' age ',' occupation
SAY "Name: " firstName lastName " | Age: " age
In this operation, the VAR keyword tells Rexx to parse the variable line. The commas in quotes are treated as delimiters; Rexx looks for the first comma, assigns everything before it to firstName, and then searches for the next comma to assign the next value.
Implementing Fixed-Width Parsing
For files where data is defined by column position rather than delimiters, you can specify numeric positions within the template. This is common in mainframe-style data exports.
/* Example: Fixed-width line: Col 1-5 ID, 6-15 Name, 16-20 Code */
line = "12345Smith A101 "
PARSE VAR line 1 5 id 6 15 name 16 20 code
/* Remove trailing blanks from fixed-width captures */
name = STRIP(name)
The numbers 1 5 tell Rexx to extract characters from position 1 through 5 and assign them to the variable id. Note that fixed-width parsing often captures trailing spaces, requiring the STRIP() function to clean the resulting variables.
Complete File Processing Workflow
To process an entire file, combine the LINEIN function with a loop and a PARSE template. Run this script from your Rexx command line or IDE.
/* Process data.txt and store in a stem array */
filename = "data.txt"
count = 0
/* Open file for reading */
do while (line = LINEIN(filename))
count = count + 1
/* Template: ID, Value, Status */
PARSE VAR line id ',' value ',' status
/* Basic validation: ensure we have at least an ID and Value */
if id = "" | value = "" then do
SAY "Warning: Malformed data on line " count
iterate
end
/* Store in stem array for later use */
record.count.id = id
record.count.val = value
record.count.stat = status
end
SAY "Successfully parsed " count " lines."
Diagnostic Checks and Verification
To verify that your template is capturing the correct slices of data, use the TRACE instruction. Running TRACE ?R (which shows Rexx variables and their values during execution) allows you to see exactly what is being assigned to each variable during the PARSE step.
Verification Checklist:
- Empty Fields: Test a line like
"Value1,,Value3". Rexx will assign an empty string to the middle variable. - Short Lines: If a line ends before the template is satisfied, the remaining variables are assigned empty strings.
- Trailing Delimiters: Ensure the template does not expect a trailing delimiter if the file lines do not end with one.
Recovery and Error Handling
Because PARSE does not throw a runtime error when a pattern is not found (it simply assigns empty strings), you must implement manual validation:
- Null Check: Immediately after the
PARSEinstruction, check if critical variables are empty usingif var = "" then.... - Logging: When a line fails validation, write the original
linevariable to an error log to identify pattern mismatches. - Iterate: Use the
ITERATEcommand to skip the current malformed line and proceed to the next record without terminating the script.
Limitations
The PARSE instruction is highly efficient for consistent structures but struggles with variable-length delimiters (like multiple spaces) or nested delimiters (like commas inside quoted strings). For complex CSVs containing quoted commas, a dedicated CSV parser or a loop that scans for quotes is required, as a simple PARSE VAR line ',' field will split on every comma regardless of context.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.