Using SPSS SPLIT FILE to Run Separate Analyses by Subgroup
Learn how to use SPSS SPLIT FILE LAYERED BY to run separate regressions (or any procedure) for each subgroup without creating new data files, plus verification steps and pitfalls to avoid.
02 Aug 2025, 21:19 UTC

Quick Answer
You can make SPSS run every subsequent procedure (e.g., REGRESSION, DESCRIPTIVES) separately for each value of a variable without creating new data files. Activate split‑file processing with SPLIT FILE LAYERED BY, run your analysis, then turn it off with SPLIT FILE OFF. The output will contain a separate block for each subgroup.
How SPLIT FILE Works
The split variable tells SPSS to treat each distinct value as a temporary layer. Before issuing the split command, the dataset must be sorted by that variable; otherwise SPSS ignores the split and analyzes the whole file. When LAYERED is used, all output appears in the same Viewer window, grouped by layer. SEPARATE would instead send each layer’s output to a distinct block, which can be handy for exporting.
Worked Example
Run the following syntax in the SPSS Syntax Editor (you need permission to edit and run syntax; no special admin rights are required). The example assumes a dataset with variables dept (numeric or string), salary, yearseduc, and exp.
SORT CASES by dept.
SPLIT FILE LAYERED BY dept.
REGRESSION
/DEPENDENT salary
/METHOD=enter yearseduc exp.
SPLIT FILE OFF.
Explanation:
SORT CASESorders the file bydept; this step is mandatory.SPLIT FILE LAYERED BY deptactivates split‑file mode. SPSS will now treat each uniquedeptvalue as a layer.REGRESSIONis executed once per layer; the Output Viewer shows a regression table labeled, for example,dept = 1,dept = 2, etc.SPLIT FILE OFFdeactivates the split so that later commands analyze the full dataset.
Verification Steps
After running the block, confirm that the split was active and that the output matches expectations:
- In the Syntax Editor, issue
SHOW SPLIT.– while the split is active SPSS returns something likeSPLIT FILE LAYERED BY dept; afterSPLIT FILE OFFit returnsNO SPLIT FILE. - Inspect the Output Viewer: each regression table should have a title indicating the split value (e.g.,
Regression Analysis - Dependent Variable: salary [dept = 3]). - To check that missing split values were excluded, compare the N reported in each table’s title with the frequency count of
dept(obtainable viaFREQUENCIES VARIABLES=dept.). Any discrepancy indicates cases with missingdeptwere omitted from the split output.
Limitations and Common Mistakes
- Sorting requirement: If you forget to sort, SPSS will either ignore the split or produce an error, and the procedure will run on the entire dataset.
- Only one split active: Issuing a new
SPLIT FILEcommand replaces the previous one without warning. Running two different splits sequentially without turning the first off can lead to confusion. - Missing values: Cases with system‑missing or user‑missing values in the split variable are excluded from all split‑file output, which can reduce your effective sample size unexpectedly.
- Forgetting to turn off: Leaving
SPLIT FILEactive causes every subsequent command (even unrelated ones likeDESCRIPTIVES) to be repeated per layer, generating unnecessary output and slowing performance.
Rolling Back the Split
Because split‑file mode changes how SPSS interprets commands, the only way to “undo” it is to issue SPLIT FILE OFF. This returns the session to the default state where procedures analyze the whole file. No data are altered, so there is no risk of losing information; the rollback merely resets the analytical context.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.