Using apoc.periodic.iterate for Large‑Scale Batched Updates in Neo4j
Learn how to safely update millions of nodes with Neo4j’s APOC apoc.periodic.iterate, complete with a real Cypher example, best‑practice configuration, and common pitfalls to avoid.
06 Jul 2025, 11:52 UTC

Why batched updates matter in Neo4j
When a graph contains millions of nodes, a single Cypher statement that touches every node can exceed the transaction size limit (default 100 MB) or exhaust the JVM heap. The apoc.periodic.iterate procedure splits a large workload into smaller, independent transactions that run one after the other or in parallel. This keeps memory usage predictable, gives you fine‑grained control over retry logic, and allows you to monitor progress.
Core syntax and a concrete example
The procedure has the following signature:
CALL apoc.periodic.iterate(outerQuery, innerQuery, options) YIELD batches, total
- outerQuery – a Cypher query that returns a list of items to process. Each row becomes the context for the inner query.
- innerQuery – a Cypher statement that performs the actual update. It can reference columns from the outer query via
{columnName}placeholders. - options – a map to tune batch size, parallelism, timeout, and other knobs.
Below is a minimal working example that flags every Person node as processed. Adjust batchSize to match your environment’s heap and throughput.
CALL apoc.periodic.iterate(
"MATCH (p:Person) RETURN p",
"SET p.processed = true",
{
batchSize: 1000,
parallel: true,
iterateList: true
}
) YIELD batches, total
RETURN batches, total;
Run this on a Neo4j 4.4+ instance with the APOC plugin installed. After execution, total should equal the number of Person nodes in the database.
Understanding the options map
| Option | Default | Purpose |
|---|---|---|
batchSize | 1000 | Number of rows processed per inner transaction. |
parallel | false | Run inner queries concurrently using a thread pool. |
iterateList | false | Flatten list results from the outer query so each element is processed individually. |
timeout | 0 (no timeout) | Maximum time (seconds) for each inner transaction. |
retry | 0 | Number of times to retry a failed inner transaction. |
queryTimeout | 0 | Maximum time for the outer query. |
When parallel is true, each inner transaction runs on a separate thread. This increases throughput but demands that the inner query be thread‑safe (e.g., it should not modify the same node from two threads simultaneously).
Best‑practice configuration
- Choose a realistic batchSize – Start with 500–1000 rows. Monitor
dbms.transaction.activeand heap usage; if you see a spike, reduce the size. - Enable iterateList only when needed – If the outer query returns a list of collections (e.g.,
RETURN collect(n)), setiterateList:trueso each element is processed separately. - Use idempotent inner queries – Because retries may re‑execute the inner statement, ensure that setting a property or creating a relationship does not duplicate data.
- Index relevant properties – If the inner query contains a
WHEREclause, make sure the property is indexed to keep each transaction fast. - Test on a subset first – Run the procedure on 1 % of the data, verify
total, and inspect a few nodes before scaling up. - Monitor metrics – Use Neo4j’s
dbms.metricsor external monitoring to watch transaction counts, commit times, and heap consumption during the run.
Common pitfalls and how to avoid them
- Too large a batchSize – Causes
OutOfMemoryErroror transaction aborts. Reduce the size and observe the impact. - Parallel updates on the same node – If the inner query touches a node that other threads also modify, a deadlock can occur. Keep updates isolated or set
parallel:false. - Missing iterateList:true – When the outer query returns lists, the inner query receives a list instead of a single node, leading to unexpected results or failures.
- Using apoc.periodic.iterate for read‑only workloads – For simple reads, plain Cypher or
CALL apoc.periodic.commitis more efficient. - Ignoring retry logic – Without a
retryoption, transient failures will abort the entire operation. Set a reasonable retry count (e.g., 3) if you expect occasional lock timeouts.
Verification checklist after execution
- Run
CALL db.stats.nodes()and compare the count of nodes withprocessed=trueto thetotalreturned by the procedure. - Execute
PROFILE CALL apoc.periodic.iterate(...)on a small sample to confirm that the inner query uses index lookups. - Inspect a random sample of nodes:
MATCH (p:Person) WHERE p.processed=true RETURN p LIMIT 5. - Check
dbms.metrics.active_transactionsanddbms.memory.heap.usedHeapMemorybefore and after the run to ensure memory stayed within limits.
When to choose apoc.periodic.iterate
Use this procedure when you need to:
- Process millions of nodes or relationships in a single operation.
- Maintain control over transaction size and retry behavior.
- Parallelise updates to improve throughput on multi‑core hardware.
For simple, small‑scale updates, a single SET or MERGE statement is usually sufficient and faster.
Conclusion
apoc.periodic.iterate is a powerful tool for large‑scale graph updates, but it requires careful tuning of batch size, parallelism, and retry logic. By following the configuration guidelines above, verifying the result, and monitoring metrics, you can perform reliable, efficient updates on massive Neo4j datasets without hitting transaction or memory limits.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.