Stata Frames: Stop Juggling Datasets and Start Linking Them
Stata 16's Frames feature lets multiple datasets coexist in memory with named frames, observation-level links, and cross-frame variable access — replacing save/load cycles with in-memory joins. Here's how to use them, where they shine, and the memory and compatibility trade-offs to watch.
03 Nov 2025, 02:00 UTC

The Problem: One Dataset at a Time
If you've used Stata before version 16, you know the drill: load a dataset, merge in a lookup table, run your analysis, save results, clear memory, load the next dataset, repeat. Every context switch means a save/load cycle. When you're joining survey responses to geographic codes, then pulling in model coefficients for a report, the friction adds up.
Stata 16 introduced Frames — a named-frame architecture that lets multiple datasets coexist in memory simultaneously. Each frame holds its own variables, observations, and macros. You switch between them with frame change, link them observation-by-observation with frlink, and pull variables across frames with frget. The mental model shifts from "which dataset is loaded?" to "which frame am I working in?"
How Frames Change the Workflow
Instead of a single active dataset, you now have a workspace of frames. The default frame is called default. Create a new one with frame create lookup. Each frame is isolated: variables, observations, and even local macros defined inside frame lookup { local x = 1 } stay invisible to other frames.
Linking is where the power lives. frlink creates observation-level relationships — 1:1, 1:m, or m:1 — between frames based on key variables. Once linked, frget pulls variables from the linked frame into the current one, and frun executes commands in another frame's context without switching.
This integrates cleanly with collect for table building and putdocx/putpdf for reporting. You can stage cleaned data in one frame, model results in another, and formatted tables in a third, then export everything without ever clearing memory.
Worked Example: Survey Data with Geographic Lookup
Suppose you have a survey dataset (survey.dta) with respondent IDs and county codes, and a separate geographic lookup table (county_codes.dta) with county codes, names, and region assignments. In the old model, you'd merge, analyze, then maybe drop the merge variables. With Frames:
* Load survey data into default frame
use survey.dta, clear
* Create a frame for the lookup table
frame create geo
frame geo {
use county_codes.dta, clear
keep county_code county_name region
}
* Link 1:1 on county_code
frlink 1:1 county_code, frame(geo)
* Pull county_name and region into the survey frame
frget county_name region, frame(geo)
* Now run analysis with geographic variables available
regress income i.region age education
* Store results in a third frame
frame create results
frame results {
estimates store model1
estimates table model1, b se p
}
Three frames coexist: default (survey + pulled variables), geo (clean lookup), results (estimation output). No temporary files, no repeated merges, no clearing.
Trade-offs and Gotchas
- Memory is additive. A 500 MB dataset in three frames consumes ~1.5 GB RAM plus Stata's overhead. On 32-bit Stata (still used in some legacy environments), this hits hard limits fast.
- Cartesian explosions are silent. A many-to-many
frlinkfollowed byfrgetcan multiply observations into the millions without warning. Always verify link cardinality withfrlink describefirst. - Frame-local macros don't appear in global macro lists. Debugging
frame myframe { local x = 1 }means you mustframe myframe { macro list }to see them. - Backward compatibility breaks. Multi-frame
.dtafiles saved in Stata 16+ won't open correctly in Stata 15 or earlier — only the active frame loads. - Old community commands may assume a single dataset. Pre-2019 ado-files often break when run inside a non-default frame.
When to Adopt Frames
Frames shine when you regularly juggle related datasets: master data + lookup tables, longitudinal panels + cross-sectional supplements, or analysis data + reporting staging areas. The frlink/frget pattern replaces merge/save/load cycles with in-memory joins that preserve both source frames intact.
Start small: create a lookup frame for your most-used reference table, link it once, and pull variables as needed. Measure memory with memory before and after. If you're on Stata 16+ with 64-bit and 8+ GB RAM, the overhead is usually negligible for typical social-science datasets.
For programmatic workflows, the Mata API (st_frame(), st_frun(), st_fstore()) lets you build frames, run estimation, and retrieve results without touching the interactive session — useful for plugin-style automation.
Bottom line: if you've ever thought "I wish I could keep both datasets open," Frames is that wish granted. Just respect the memory budget and verify your links.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.