Using Clojure Transducers to Avoid Intermediate Sequence Allocation
Learn how Clojure transducers let you compose map/filter steps without allocating intermediate lazy sequences, reducing GC pressure in data pipelines.
14 Jan 2026, 09:07 UTC

Why intermediate sequences hurt performance
When you chain map, filter or take over a lazy seq, each step builds a new seq object that holds the results of the previous step. For a million-item range this means allocating large numbers of temporary seq nodes just to throw them away after the next step. The garbage collector spends time cleaning those objects, and the overall runtime grows with the number of passes.
What a transducer actually is
A transducer is a higher-order reducing function transformer. It takes a reducing function rf – the operation that combines a single input with an accumulator (e.g. +, conj, merge) – and returns a new reducing function that first transforms the input according to a step (map, filter, etc.) and then calls the original rf. Because the transformation logic is wrapped inside the reducing function, the same transducer can be fed any reducible source: a vector, a list, a Java Iterable, or a core.async channel.
Composing and running a transducer pipeline
Transducers compose with the standard comp function. The example below builds a pipeline that increments each number and then keeps only the even ones.
(def xf (comp (map inc) (filter even?)))
To execute the pipeline eagerly we use transduce, supplying a reducing function and an initial accumulator:
(transduce xf + 0 (range 1 1000000))
The call walks the range once, applies inc and filter even? to each element, and feeds the resulting values directly into +. No intermediate lazy seqs are created between the steps – though note the pipeline is not literally allocation-free: the reducing function and the source itself may still allocate.
For comparison, the lazy-seq version looks like this:
(reduce + 0 (filter even? (map inc (range 1 1000000))))
Both expressions produce the same sum, but the transducer form avoids building the temporary seqs that the lazy version creates at each step.
Checking the difference in practice
You can verify the allocation difference yourself with a simple benchmark:
- Wrap each expression in
(time ...)in a REPL, or useSystem/nanoTimearound the call. - Run the benchmark with a sufficiently large range (e.g., 10 million) and compare the reported elapsed time. Run each version several times to let the JIT warm up before drawing conclusions.
- Watch garbage-collection activity (for example with
-verbose:gcon the JVM); the transducer run should generally show less allocation pressure from intermediate seqs.
If you use a profiler such as VisualVM, you can look for instances of clojure.lang.LazySeq being created by the transformation steps. In the transducer run you should see substantially fewer of these from the pipeline itself, though the exact counts depend on the source collection and JVM behavior, so treat profiling results as indicative rather than absolute.
Trade-offs and limitations
Transducers excel when you need a pure, side-effect-free transformation that can be reduced to a single accumulator. They are less convenient when you want to:
- Insert debugging prints between steps (you would need to wrap
printlnin a reducing function). - Produce a lazy sequence as the final result; you must wrap the transducer with
sequenceorintoif you need a concrete collection. - Express algorithms that naturally rely on retaining intermediate results, such as overlapping partitions or sliding windows, without extra bookkeeping.
There is also a readability cost: developers unfamiliar with reducing functions often find transducer-based code harder to follow and debug than a classic map/filter chain. For quick REPL exploration or code where clarity outweighs raw performance, lazy seqs remain a perfectly fine choice.
Actionable takeaway
Next time you find yourself chaining several sequence operations over a large collection, try extracting the transformation logic into a transducer with comp and feeding it to transduce. Measure the runtime and GC impact on your own workload; in tight loops and data-processing pipelines you will often see a noticeable reduction in intermediate allocation, but confirm it with a benchmark rather than assuming it.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.