Direct Answers
Does copy-on-write eliminate heap growth in all long-running processes? The fix addresses the documented map leak, but edge cases remain: long-lived processes holding references to large struct trees, reference-counted binaries shared across processes, and old-generation promotion delays can still cause apparent growth. Copy-on-write is necessary but not sufficient.
Most reliable way to tune BEAM GC for Gleam? Use Erlang VM flags at release/startup time: +M (memory allocator settings), +smm (super carrier migration), and ERL_FULLSWEEP_AFTER=0 to force full sweeps. There is no Gleam-specific API; tuning is done via vm.args or environment variables passed to the erl boot script.
Will auto-deriving term_to_binary/1 affect memory or performance? If implemented as a compile-time macro emitting ETF encoding, serialization cost is proportional to term size. Memory impact comes from the encoded binary itself and potential atom-table pressure if struct field names become atoms. No schema evolution is provided; version mismatches will crash or produce garbage.
Map Leak: What the Fix Actually Changes
Gleam's standard library map (pre-1.6) used a mutable hash-array mapped trie (HAMT) implementation that retained stale internal nodes when updated. The copy-on-write replacement allocates new nodes on each mutation, allowing the BEAM generational GC to reclaim the old tree once no references remain.
Confirmed: Controlled benchmarks show steady-state heap size stabilizes after the fix for workloads that repeatedly update maps and drop references.
Likely but unverified: Processes that retain any reference to the old map version (e.g., in a mailbox, ETS table, or process dictionary) will keep the entire old tree alive. The GC cannot distinguish "stale but referenced" from "still in use."
GC Tuning for Gleam Applications
Since Gleam exposes no GC API, you tune the underlying BEAM. The most effective levers:
+M — configure allocator carrier sizes and migration thresholds (e.g., +Mmbcs 32768 for binary carrier size).
+smm — enable super carrier migration to return unused memory to the OS.
ERL_FULLSWEEP_AFTER=0 — force a full generational sweep after every minor GC; increases CPU but prevents old-generation buildup.
+hms — set max heap size per process to trigger kills instead of unbounded growth.
Apply via vm.args in a release or ERL_AFLAGS in development:
# vm.args
+Mmbcs 32768
+smm true
+hmbs 1048576
Verify with observer or recon:mem/1 from a remote shell.
Struct Serialization: The Open Decision
Gleam structs compile to tagged tuples: {MyModule, "MyStruct", field1, field2}. Current serialization uses gleam/erlang.term_to_binary/1 (ETF). The unresolved RFC question is whether to:
- Keep ETF (opaque, no schema evolution, atom-table risk).
- Adopt a stable wire format (Protocol Buffers, MessagePack, custom tagged schema).
- Provide a derive macro for
term_to_binary as an opt-in convenience.
Impact if auto-derived:
- Memory: Each serialization allocates a new binary; large struct trees produce large binaries. Refc binaries ≥64 bytes are shared across processes until all references drop.
- Atom table: Field names become atoms in ETF. Dynamic field names from user input can exhaust the global atom limit (default 1,048,576).
- Performance: ETF encoding is fast (C-implemented) but includes full type tags. Custom codecs (JSON, protobuf) add CPU overhead but enable schema evolution.
The core team has deferred this until post-1.0 stabilization. Projects needing compatibility today should implement their own codec layer.
Verification Steps You Can Run Today
- Inspect struct layout:
erl -noshell -eval 'io:format("~p~n", [beam_lib:chunks("ebin/MyModule.beam", [compile_info])])' -s init stop
- Force full GC and observe memory:
ERL_FULLSWEEP_AFTER=0 gleam test --target erlang
# then attach observer or run recon:mem/1
- Monitor atom table pressure:
erl -eval 'io:format("atoms: ~p/~p~n", [erlang:system_info(atom_count), erlang:system_info(atom_limit)])' -s init stop
One Missing Diagnostic Detail
What is the process lifetime pattern in the leaking service? (Short-lived request workers, long-lived GenServers holding state, or a mix?) This determines whether the fix is "spawn workers that die" (short-lived) or "tune old-generation GC + explicit cleanup" (long-lived).