Ballerina Client-Side HTTP Resilience: Per-Client Timeout, Retry, and Circuit Breaker
Configure per-client timeout, retry, and circuit breaker in Ballerina's http:Client to bound latency and absorb transient failures without a service mesh.
11 Sept 2025, 18:14 UTC

Requirements: Resilient Outbound Calls Without a Mesh
When a Ballerina service calls downstream HTTP APIs, each call can hang, return errors, or degrade under load. The requirement is to bound latency, absorb transient failures, and prevent cascading overload—without introducing a service mesh, sidecar, or external resilience library. The solution must be declared once per downstream dependency, enforced at compile time, and observable in production.
Smallest Suitable Design: One Configured Client Per Dependency
Create a single module-level http:Client for each downstream service. Configure timeout, retry, and circuit breaker in the client's ClientConfiguration record. This keeps resilience policy collocated with the dependency it protects, avoids per-call boilerplate, and lets the compiler validate field names.
import ballerina/http;
import ballerina/config;
// Module-level client for the payments service
configurable string paymentsBaseUrl = ?;
configurable int requestTimeoutMs = 3000;
configurable int retryCount = 2;
configurable int retryIntervalMs = 200;
configurable float retryBackoff = 2.0;
configurable int breakerThreshold = 5;
configurable int breakerWindowSec = 10;
configurable int breakerCoolDownSec = 30;
http:Client paymentsClient = new (paymentsBaseUrl, {
timeout: requestTimeoutMs,
retry: {
count: retryCount,
interval: retryIntervalMs,
backoff: retryBackoff,
statusCodes: [500, 502, 503, 504]
},
circuitBreaker: {
rollingWindow: breakerWindowSec,
threshold: breakerThreshold,
resetTime: breakerCoolDownSec
}
});
Place this in a dedicated module (e.g., clients/payments.bal). Other modules import the client and call paymentsClient->post(...) without repeating resilience logic.
Trust and Data Boundaries
The http:Client runs in-process. The security boundary is the TLS connection to the downstream. Credentials (API keys, tokens, mTLS certificates) must come from config.toml or environment variables—never hardcoded or logged. Ballerina's config module reads these at startup; the client configuration references them via configurable variables.
Retries shift delivery semantics toward at-least-once. If the downstream is not idempotent (e.g., a payment POST that charges a card), a retried request whose first attempt succeeded will double-apply the effect. Only enable retry for:
- Idempotent operations (GET, PUT, DELETE)
- POST endpoints that implement idempotency keys or deduplication
For non-idempotent calls, omit the retry field entirely or set count: 0.
Error Handling: Union Types Over Panics
Every client call returns a union type including error. Use check or explicit matching to map downstream failures to deliberate responses.
public function chargePayment(PaymentRequest req) returns http:Response | error {
http:Response resp = check paymentsClient->post("/charges", req);
return resp;
}
// In a resource, map errors to controlled degradation
resource function post payments(@http:Payload PaymentRequest req) returns error? {
http:Response|error result = chargePayment(req);
if (result is error) {
// Breaker open, timeout, or non-retryable 5xx
return respondWith503("Payment service temporarily unavailable");
}
return result;
}
This ensures breaker trips surface as HTTP 503 (or your chosen degradation response) rather than unhandled panics.
Operational Checks: Built-In Observability
Ballerina ships observability (logs, metrics, distributed tracing) that activates with minimal configuration. Enable it in Cloud.toml or via bal run --observability-included.
[observability]
metrics = true
tracing = true
logging = true
Key metrics to watch per client:
http_client_request_duration_seconds— downstream latency distributionhttp_client_requests_totalwithstatuslabel — error rateshttp_client_retries_total— retry frequencyhttp_client_circuit_breaker_state— open/closed/half-open transitions
Alert on sustained open circuits (e.g., breaker open > 1 minute) rather than individual retries. A flapping breaker (rapid open/close cycles) indicates thresholds set too tightly.
Failure Modes
| Failure Mode | Cause | Mitigation |
|---|---|---|
| Retry amplification | Downstream struggling; retries multiply load | Limit retry count (1–2), use exponential backoff, restrict to idempotent calls |
| Breaker flapping | Threshold too low or window too short | Increase threshold or rollingWindow; require sustained error rate |
| Timeout exceeds caller budget | Client timeout > upstream service's own deadline | Set client timeout < caller's latency budget; propagate deadlines via headers if needed |
| Double-apply on non-idempotent POST | First attempt succeeded but response lost; retry re-sends | Disable retry for non-idempotent endpoints; use idempotency keys |
| No timeout configured | Default timeout may be infinite or very long | Always set explicit timeout; verify default for your Ballerina version |
Conditions That Would Change the Design
Move resilience out of per-client configuration when:
- Polyglot fleet: Multiple languages/services need consistent policies—use a service mesh (Istio, Linkerd) or API gateway.
- Fleet-wide breaker state: You need a circuit breaker that shares state across service replicas—requires external coordination (e.g., Redis-backed breaker).
- Hedged requests: You want to send duplicate requests to multiple instances and use the first response—not expressible in current
http:Clientconfig. - Shift to async messaging: Outbound calls become event publishing; resilience moves to broker acknowledgments, dead-letter queues, and consumer retries.
Verification Checklist
- Pin Ballerina version (e.g.,
2201.8.6inBallerina.toml). Read theballerina/httpmodule API docs forClientConfiguration,Retry, andCircuitBreakerrecords. - Compile with
bal build. The compiler validates record fields—field-name drift from older examples fails fast. - Local reproduction: Point the client at a test endpoint that returns 503 or delays (e.g.,
httpbin.org/status/503). Observe retries in logs, then fast-fail once breaker opens. - Staging observability: Enable built-in observability; confirm downstream latency, error, retry, and breaker metrics appear before wiring alerts to them.
Limitations
- Breaker state is per client instance in-process; it does not coordinate across replicas of your service.
- Exact record field names and defaults differ across Swan Lake update releases; confirm against your pinned version's library docs.
- Pre-Swan-Lake (Ballerina 1.x) code samples use a different HTTP API and will not compile on current versions.
Practical Result Check
After deployment, verify in staging:
- Call the downstream with induced 5xx errors; confirm client retries
retryCounttimes with backoff. - After
thresholderrors withinrollingWindowseconds, confirm breaker opens and subsequent calls fail fast. - After
resetTimeseconds, confirm breaker half-opens, allows a probe request, and closes on success. - Confirm metrics
http_client_circuit_breaker_statetransitions appear in your monitoring stack.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.