Internal Pool Scaling vs. External Pooler Deployment: Choosing the Right Approach for ArgoCD Connection Exhaustion
0 reputation · 07 Jul 2020, 20:46 UTC
Diagnosing Intermittent Connection Pool Exhaustion
In a highly‑available ArgoCD cluster, sync waves and frequent repository polling can trigger “pool exhausted” or “dial tcp” errors. The root cause is the mismatch between the per‑pod connection pools ArgoCD maintains for Redis and PostgreSQL and the hard limits those backends enforce.
ArgoCD exposes two documented mitigation paths: 1) tune the internal pool sizes via argocd-cm keys such as redis.connectionPoolSize and db.pool.maxOpenConnections; 2) add an external pooler (Redis Cluster or PgBouncer) to multiplex connections. Each path has trade‑offs: internal scaling is straightforward but risks hitting backend maxclients or max_connections; an external pooler adds operational overhead but can provide better observability and failover.
Currently ArgoCD does not expose pool utilization metrics, so operators infer exhaustion from logs or backend counters. The decision point is which approach delivers reliable performance without exceeding backend limits.
What metrics or instrumentation can be added to expose per‑pod pool usage? How can the optimal internal pool size be calculated given backend limits and expected traffic? Does enabling a PgBouncer transaction pool interfere with ArgoCD’s use of prepared statements, and what configuration is required to mitigate that?