Autoscale Azure App Service Based on Custom Latency Metrics from Application Insights
Learn how to drive Azure App Service autoscale with a custom Application Insights latency metric, including a ready‑to‑use JSON template, CLI steps, and verification tips.
03 Jul 2025, 03:16 UTC

Problem: Latency spikes cause slow responses
When an App Service plan runs at a fixed size, sudden traffic bursts can push request latency above acceptable limits. Users see slow pages, and the only way to react is to manually increase the instance count—a slow, error‑prone process.
Thesis: Use a custom Application Insights metric to drive autoscale
By exposing a latency metric (e.g., the 95th percentile of request duration) from Application Insights to Azure Monitor, you can create an autoscale rule that adds or removes instances based on real‑world performance rather than on generic CPU or memory thresholds.
How custom metrics work with Azure Monitor
Application Insights already stores request duration as a custom metric under the namespace microsoft.insights/components. Azure Monitor can read this namespace directly when you configure an autoscale setting.
Key concepts:
- Custom metric: a value emitted by your application (here, request duration) that is stored in Application Insights and made available to Monitor.
- Autoscale rule: a condition that evaluates a metric over a time grain and triggers a scale action when the condition is true.
- Cooldown: the minimum time between successive scale actions, preventing thrashing.
Worked example: Create the autoscale setting with Azure CLI
Assume you have an App Service plan named myplan in resource group my-rg and an Application Insights component with ID /subscriptions/<sub-id>/resourceGroups/my-rg/providers/microsoft.insights/components/myai.
- Sign in and set the subscription (requires
Microsoft.Insights/autoscaleSettings/writepermission):az account set --subscription <sub-id> - Create a JSON rule file (
autoscale-latency.json):{ "location": "West US 2", "properties": { "enabled": true, "name": "latency-autoscale", "targetResourceUri": "/subscriptions/<sub-id>/resourceGroups/my-rg/providers/Microsoft.Web/serverfarms/myplan", "profiles": [ { "name": "latency-profile", "capacity": { "minimum": "2", "maximum": "10", "default": "2" }, "rules": [ { "metricTrigger": { "metricName": "request duration", "metricNamespace": "microsoft.insights/components", "metricResourceUri": "/subscriptions/<sub-id>/resourceGroups/my-rg/providers/microsoft.insights/components/myai", "timeGrain": "PT1M", "statistic": "Percentile95", "timeWindow": "PT5M", "timeAggregation": "Average", "operator": "GreaterThan", "threshold": 2.0 }, "scaleAction": { "direction": "Increase", "type": "ChangeCount", "value": "1", "cooldown": "PT5M" } }, { "metricTrigger": { "metricName": "request duration", "metricNamespace": "microsoft.insights/components", "metricResourceUri": "/subscriptions/<sub-id>/resourceGroups/my-rg/providers/microsoft.insights/components/myai", "timeGrain": "PT1M", "statistic": "Percentile95", "timeWindow": "PT5M", "timeAggregation": "Average", "operator": "LessThan", "threshold": 0.8 }, "scaleAction": { "direction": "Decrease", "type": "ChangeCount", "value": "1", "cooldown": "PT5M" } } ] } ] } } - Deploy the autoscale setting:
az monitor autoscale create --resource-group my-rg --name latency-autoscale --config autoscale-latency.json
After deployment, Monitor evaluates the 95th‑percentile request duration every minute, looking at the last five minutes. If the average of that percentile exceeds 2 seconds, the plan scales out by one instance; if it falls below 0.8 seconds, it scales in.
Trade‑offs and limitations
- Ingestion cost: Each request duration value is a custom metric; high‑traffic apps can generate many data points, increasing Azure Monitor charges.
- Evaluation lag: The metric is aggregated over the chosen time window (5 minutes in the example). Sudden spikes may not trigger scaling until the window elapses.
- Upper bound: Ensure the App Service plan’s maximum instance count is high enough to allow scale‑out; otherwise the rule will be throttled.
Practical way to check the result
- In the Azure portal, go to the App Service plan → Scale out (App Service plan) → Run history. You should see scale events that align with latency spikes.
- Open Metrics Explorer, chart the custom metric
request duration(Percentile95) alongside theInstance countmetric. Verify that instance count rises after the metric crosses the 2‑second threshold and falls after it drops below 0.8 seconds. - Run a short load test (e.g., with Azure Load Testing) and watch the metric and instance count in near‑real time.
Actionable closing
Start by defining a latency SLO for your service, expose the corresponding percentile as a custom metric in Application Insights, and then create an autoscale profile using the JSON template above. Monitor the cost of custom metric ingestion and adjust the time grain or cooldown if you observe excessive scaling actions. This approach ties scaling directly to user‑experienced performance, reducing guesswork and manual intervention.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.