Diagnosing and Fixing Goroutine Leaks in Go Services
Learn how to identify and resolve goroutine leaks in Go. This guide covers pprof diagnostics, common blocking channel patterns, and context-aware fixes to prevent memory exhaustion.
16 Oct 2025, 00:16 UTC

The Leak Condition
A goroutine leak occurs when a goroutine is started but never terminates, remaining in memory indefinitely. In long-running services, this manifests as a steady, linear increase in Resident Set Size (RSS) memory and a rising goroutine count, even when request traffic remains constant or drops. Because each goroutine allocates a minimum stack (typically 2KB), thousands of leaked routines can lead to Out-of-Memory (OOM) crashes.
Quick Diagnostic Table
| Symptom | Likely Cause | Diagnostic Indicator |
|---|---|---|
| Linear memory growth | Blocked channel send/receive | pprof shows many routines at the same line of code |
| High goroutine count | Missing termination signal | runtime.NumGoroutine() increases after every request |
| Slow performance degradation | Context cancellation failure | Stacks blocked on chan receive despite request timeout |
Step-by-Step Investigation
-
Monitor Baseline Counts:
Check the current number of active goroutines. If you have
net/http/pprofimported, navigate to/debug/pprof/goroutine?debug=1in your browser. Look for a high number of goroutines blocked on the same function call. -
Capture a Stack Dump:
Use the pprof tool from your terminal to analyze the distribution of goroutines. Run this command from your local machine targeting the service:
# Run as a user with network access to the service go tool pprof http://localhost:6060/debug/pprof/goroutineOnce inside the pprof interactive shell, type
topto see which functions are holding the most goroutines. -
Identify the Block Point:
Look for stacks that end in
chan receiveorchan send. If you see 500 goroutines all stopped atmain.go:42waiting on a channel, you have found the leak site.
Common Leak Patterns and Fixes
Pattern 1: The Abandoned Sender
This occurs when a goroutine sends a result to an unbuffered channel, but the receiver has already timed out or returned, leaving the sender blocked forever.
Incorrect Implementation:
func process(data string) {
ch := make(chan bool)
go func() {
// Simulate work
ch <- true // Blocks forever if the receiver exits early
}()
select {
case <-ch:
return
case <-time.After(1 * time.Second):
return // Receiver leaves; sender is now leaked
}
}
The Fix: Use Buffered Channels or Contexts If the number of responses is known, use a buffered channel of size 1. This allows the sender to complete the operation and exit even if the receiver is gone.
func process(data string) {
// Buffer of 1 prevents the sender from blocking
ch := make(chan bool, 1)
go func() {
ch <- true
}()
select {
case <-ch:
return
case <-time.After(1 * time.Second):
return
}
}
Pattern 2: Ignoring Context Cancellation
Goroutines that perform long-polling or wait on external events often forget to monitor the ctx.Done() channel, meaning they persist even after the HTTP request that spawned them has been closed.
The Fix: Context-Aware Select
Always wrap blocking channel operations in a select block that includes the context.
func worker(ctx context.Context, jobs <-chan int) {
for {
select {
case job := <-jobs:
// Process job
case <-ctx.Done():
// Exit cleanly when parent request is cancelled
return
}
}
}
Verification and Limitations
To verify the fix, write a test that checks the goroutine count before and after the function execution:
func TestForLeaks(t *testing.T) {
initial := runtime.NumGoroutine()
process("test")
// Give the scheduler a moment to clean up
time.Sleep(10 * time.Millisecond)
final := runtime.NumGoroutine()
if final > initial {
t.Errorf("Leak detected: started with %d, ended with %d", initial, final)
}
}
Limitations
- Over-buffering: While buffered channels prevent leaks, setting buffers too high can mask architectural bottlenecks and lead to sudden OOM errors under extreme load.
- Panic Risks: Be cautious when closing channels to signal termination. Sending to a closed channel or closing a channel twice will cause a runtime panic.
- Scheduler Latency:
runtime.NumGoroutine()may not drop immediately after a function returns; a short sleep orruntime.GC()may be necessary for accurate test measurements.
Escalation Criteria
If goroutine counts continue to rise despite implementing context cancellation and buffered channels, escalate to a full heap profile (/debug/pprof/heap) to determine if the leak is caused by objects being held in global slices or maps, preventing the garbage collector from reclaiming the memory associated with the goroutine stacks.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.