Diagnosing Unexpected Scaling Behavior in Discloud
A step‑by‑step guide to diagnose and correct unexpected auto‑scaling behavior in Discloud, with checks, fixes, and escalation paths.
02 Apr 2026, 09:08 UTC

Recognizable Condition
You notice that your Discloud virtual machines are either scaling up too aggressively (incurring unexpected costs) or not scaling up when load increases (causing latency spikes).
Cause/Diagnostic Table
| Symptom | Possible Cause |
|---|---|
| Instances add rapidly despite modest traffic | Scaling threshold set too low or metric aggregation window too short |
| Instances stay constant while CPU > 80% | Health‑check failures causing instances to be marked unhealthy and removed before scaling decision |
| Billing shows higher per‑second usage than expected | Scale‑in cooldown too short, leading to frequent terminate‑and‑restart cycles |
Ordered Checks
- Review scaling policy – In the Discloud control panel, navigate to Instances → Scaling Policies and note the target metric, threshold, evaluation period, and cooldown values.
- Inspect metric collection – Open the Metrics Dashboard for the instance group and verify the raw values of CPU, memory, and request latency over the last 10 minutes.
- Check health‑check status – Under Instances → Health, look for any instances repeatedly transitioning to
Unhealthy. - Examine recent scaling events – In the Activity Log, filter for
ScaleUpandScaleDownentries and note timestamps. - Validate billing granularity – Export the billing CSV for the past hour and compare the summed
usage_secondsagainst the wall‑clock time reported in the activity log.
Fixes Tied to Findings
- If thresholds are too low: Increase the target CPU threshold (e.g., from 50 % to 70 %) or lengthen the evaluation period (e.g., from 30 s to 90 s).
- If health checks are failing: Adjust the health‑check grace period or fix the underlying application issue causing the failures.
- If cooldown is too short: Raise the scale‑in cooldown (e.g., from 60 s to 300 s) to prevent rapid terminate‑restart cycles.
- If metric lag is observed: Ensure the monitoring agent is running on all instances; restart it if needed (
sudo systemctl restart discloud‑agent).
Example: Adjusting a Scaling Policy via CLI
Assuming you have the Discloud CLI installed and are logged in with a user that has admin rights on the project:
# Show current policy (replace <policy-id> with your policy)
discloud scaling-policy get --id <policy-id>
# Update target CPU utilization to 70% and evaluation period to 90s
discloud scaling-policy update \
--id <policy-id> \
--target-metric cpu_utilization \
--target-value 70 \
--evaluation-period 90 \
--cooldown 300
Where to run: Any terminal with network access to the Discloud API. Permissions: API token with scaling:write scope. Expected check: After the command returns success, revisit the Metrics Dashboard; you should see fewer scale‑up events under similar load.
Escalation Criteria
- If after applying the above fixes the scaling behavior remains abnormal for more than two consecutive evaluation periods.
- If health‑check failures persist despite configuration changes (indicating a deeper application or infrastructure issue).
- If billing discrepancies exceed 10 % of expected usage after verifying the activity log.
In these cases, open a support ticket with Discloud, providing:
- Policy ID and current configuration
- Relevant metric graphs (CPU, memory, latency)
- Activity log excerpt showing the scaling events
- Billing export for the affected period
Limitations
The diagnostic steps rely on the visibility provided by the Discloud control panel and CLI. If custom metrics are not exported to Discloud’s monitoring system, threshold‑based checks will not reflect those signals. Verify that any application‑specific metrics are pushed via the Discloud agent before relying on them for scaling decisions.
Practical Way to Check the Result
After making a change, wait at least one full evaluation period (as set in the policy) and then:
- Check the Activity Log for the absence of unwanted scale‑up/down events.
- Confirm that the metric stays within the target band (e.g., CPU between 60 %‑80 %).
- Compare the billing line item for the instance group to the expected usage seconds based on the observed uptime.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.