Scalingo Auto‑Scaling for Stateless Web Services: Minimal Architecture, Trust Boundaries, and Operational Safeguards
Deploy a stateless web service on Scalingo and let the platform auto‑scale it with minimal configuration. Learn the trust boundaries, operational checks, and failure modes that can force a redesign.
06 Jan 2026, 16:21 UTC

Problem: How to Let Scalingo Handle Traffic Spikes Without Manual Intervention
When your stateless web service receives unpredictable traffic, you want the platform to add or remove instances automatically. Scalingo offers a built‑in horizontal auto‑scaling feature that evaluates CPU, memory, and latency metrics every 30 seconds. This article walks through the smallest architecture that enables that feature, explains the trust and data boundaries, and lists the operational checks and failure modes that might force you to redesign.
Requirements
- Scalingo account with API access.
- Docker image of your stateless application.
- Scalingo CLI installed and authenticated:
scalingo login. - API token with
scalingo:autoscalescope stored securely.
The Minimal Design
At its core you need only two files in your repo:
# scalingo.yml
name: myapp
services:
web:
dockerfile: Dockerfile
scale:
auto: true
# scalingo-autoscale.yml
services:
web:
min: 1
max: 5
thresholds:
cpu: 75
memory: 70
latency: 200
Deploying with scalingo push will create the service and register the auto‑scaling policy. The platform polls the metrics API every 30 s and triggers scale‑up when any threshold exceeds the limit, or scale‑down when all are below a lower bound (default 50 %). No custom code is required.
Trust and Data Boundaries
- Authentication: Scalingo authenticates scaling actions with the user’s API token. Keep the token out of the image and store it in a vault or the Scalingo secret store.
- Network isolation: All inbound traffic is routed through Scalingo’s load balancer. The application must not accept connections directly from the Internet; this prevents exposure to unfiltered traffic.
- Data persistence: The design assumes the service is stateless. Any session data, cache, or state must be stored in an external store (Redis, database, etc.) to survive pod termination.
Operational Checks
- Health‑check endpoint
- Implement
/healththat returns200 OKonly when the instance is ready. - Configure Scalingo to use this endpoint in the service definition:
health_check: /health. - During a scale‑up, if the endpoint fails, Scalingo rolls back the new instance.
- Implement
- Metrics visibility
- Expose a Prometheus exporter (or use the built‑in metrics endpoint) so you can see CPU, memory, and latency values.
- Verify in the console that the thresholds are being evaluated:
scalingo logs --service webshows scale events.
- Audit logs
- Check
scalingo audit-logsto confirm that scale actions match the policy and that rollbacks are logged.
- Check
- Billing monitor
- Auto‑scaling can increase billable dyno usage. Keep an eye on the dashboard or set a budget alert.
Failure Modes and When to Re‑Design
- Health‑check failures: If the
/healthendpoint is slow or returns errors during a scale‑up, the instance is removed. Ensure the endpoint is lightweight and cached where possible. - Missing or delayed metrics: If the metrics API stalls, the policy may not trigger, leading to over‑ or under‑provisioning. Add a fallback or use a custom exporter that pushes metrics reliably.
- Stateful workloads: If the service stores data locally, scaling will cause data loss. Move to an external database or add a persistent volume.
- Short‑lived traffic spikes: Auto‑scaling evaluates every 30 s; a burst lasting <30 s may not be captured. For such patterns, consider a higher baseline or a custom scaling script.
- Multi‑region deployment: Scalingo’s auto‑scaling is region‑specific. If you deploy across regions, you’ll need separate policies and may need to adjust latency thresholds.
- Custom metrics: If you need to scale on queue depth or business metrics, you must expose those via a Prometheus exporter and update the policy file.
Practical Example
Deploy a minimal Node.js app:
# Dockerfile
FROM node:20
WORKDIR /app
COPY package*.json ./
RUN npm ci
COPY . .
EXPOSE 3000
CMD ["node", "index.js"]
# index.js
const http = require('http');
const port = process.env.PORT || 3000;
http.createServer((req, res) => {
if (req.url === '/health') {
res.writeHead(200);
res.end('OK');
} else {
res.writeHead(200);
res.end('Hello World');
}
}).listen(port);
Push to Scalingo:
scalingo push
Generate traffic (e.g., with ab -n 10000 -c 200 http://myapp.scalingo.io/) and watch the console for scale events. Verify the new instance count in scalingo services and inspect the audit log.
Conclusion
With just a scalingo.yml and a scalingo-autoscale.yml, you can offload traffic spikes to Scalingo’s platform. The key is to keep the service stateless, secure the API token, expose healthy endpoints, and monitor metrics and logs. When your workload changes—stateful data, custom metrics, or multi‑region needs—you’ll need to extend the design accordingly.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.