Direct answer
Stata does not offer a native option to send collapse results to a separate memory buffer while keeping the original dataset active. collapse always replaces the in-memory dataset with the aggregated table. To run multiple different aggregations from the same source without repeatedly reloading the original file from disk, keep one copy of the raw data available and reset to it between aggregations using preserve/restore.
Confirmed behavior
collapse aggregates observations by the grouping variables in by() and keeps only the statistics you request. Every other observation-level variable is dropped from memory. The operation is destructive in memory, so any identifier or raw variable not named in the varlist is lost unless you saved or preserved the data first.
Recommended pattern for repeated aggregations
- Call
preserve once before the first aggregation to snapshot the current state. - Run
collapse, save or analyze the summary, then restore. - Repeat with different statistics or groupings.
preserve
collapse (mean) price, by(group_var)
save summary_mean.dta, replace
restore
preserve
collapse (sum) volume, by(group_var)
save summary_sum.dta, replace
restore
Alternative for very large data
If the preserve snapshot itself is a concern, save the raw data once to a temporary file and reload that copy between runs. This still avoids re-reading or re-importing the original source.
save temp_original.dta, replace
collapse (mean) price, by(group_var)
save summary_mean.dta, replace
use temp_original.dta, clear
collapse (sum) volume, by(group_var)
save summary_sum.dta, replace
Keeping detail and summary together
If you need both the aggregated values and the raw rows in one dataset, collapse into a separate file and merge it back onto the original using the grouping variables as keys, rather than collapsing in place.
Verification
- Run
describe before and after collapse to compare variable lists and observation counts; a sudden drop signals unintended loss. - Use
count or isid group_var on the collapsed data to confirm one row per group.
Fact vs. assumption
Confirmed: collapse overwrites memory and no new-buffer option exists. Assumption worth checking on your version: preserve snapshots the dataset efficiently, so restore is generally cheaper than re-importing the source, but exact behavior can vary by Stata version.
One diagnostic that would change the recommendation: do you need all summaries simultaneously in a single dataset, or only sequentially? If simultaneously, prefer the collapse-to-file-and-merge approach over repeated preserve/restore cycles.