Eliminating Lambda Cold Starts with Provisioned Concurrency
Learn how to eliminate Lambda cold starts using Provisioned Concurrency. This guide covers implementation via aliases, scaling strategies for peak traffic, and how to verify results using CloudWatch.
31 Oct 2025, 22:46 UTC

The Latency Spike Problem
In a serverless architecture, the "cold start" is a well-known performance hurdle. When an AWS Lambda function hasn't been used for a while, or when it needs to scale up to handle a burst of traffic, AWS must initialize a new execution environment. This involves downloading your code, starting the runtime, and running your initialization code outside the handler.
For a background task, a 500ms or 2-second delay is negligible. For a customer-facing REST API with a strict Service Level Agreement (SLA), those spikes result in a poor user experience and erratic p99 latency metrics. The goal is to ensure that the execution environment is already "warm" before the request arrives.
How Provisioned Concurrency Works
Provisioned Concurrency keeps a specified number of execution environments initialized and ready to respond immediately. Unlike the standard on-demand model—where AWS manages the lifecycle based on incoming traffic—Provisioned Concurrency allocates a pool of environments that have already completed the Init phase.
Crucially, you cannot apply Provisioned Concurrency to the $LATEST version of a function. Because the $LATEST version changes every time you update your code, AWS requires you to target a specific Version or an Alias. This ensures that the environments being kept warm are running a stable, immutable snapshot of your application.
Practical Implementation: API Warm-up
Imagine a scenario where a retail API experiences a massive surge of traffic every morning at 9:00 AM. To avoid a wave of cold starts for the first few hundred users, you can use a combination of a Lambda Alias and Application Auto Scaling.
Configuration Example
To set up a baseline of 10 warm environments for a production alias, run the following command via the AWS CLI. This assumes you have already published a version and created an alias named prod.
# Run this from a terminal with AdministratorAccess or lambda:PutProvisionedConcurrencyConfig permissions
aws lambda put-provisioned-concurrency-config
--function-name MyRetailApi
--qualifier prod
--provisioned-concurrent-executions 10
Expected Check: After running this, you can verify the status using get-provisioned-concurrency-config. The status will move from IN_PROGRESS to READY once the environments are initialized.
Scaling for Peak Hours
Static provisioning is expensive if your traffic is cyclical. You can integrate with Application Auto Scaling to adjust the warm pool based on a schedule or a target tracking policy (e.g., keeping utilization at 70%).
| Scaling Method | Use Case | Benefit |
|---|---|---|
| Scheduled Scaling | Predictable peaks (e.g., 9 AM start) | Zero cold starts at the exact start time. |
| Target Tracking | Unpredictable but gradual growth | Balances cost and performance automatically. |
Trade-offs and Limitations
Provisioned Concurrency is not a "silver bullet" and introduces specific operational trade-offs:
- Cost: You pay for the provisioned capacity from the moment it is configured, regardless of whether the function is actually executing requests.
- The "Burst" Ceiling: If you provision 10 environments but receive 15 concurrent requests, the first 10 will be instant. The remaining 5 will trigger standard on-demand scaling, meaning those 5 users will still experience a cold start.
- Deployment Lag: When you update your alias to point to a new version, the provisioned concurrency must be re-applied to the new version, which takes time to initialize.
Verifying the Result
To confirm that Provisioned Concurrency is working, use AWS X-Ray or examine your CloudWatch Logs. In a standard cold start, the log will show an Init Duration segment. For requests served by Provisioned Concurrency, the Init Duration is absent from the request log because the initialization happened before the request arrived.
Rollback: If costs spike or the feature is no longer needed, remove the configuration to return to standard on-demand scaling:
aws lambda delete-provisioned-concurrency-config
--function-name MyRetailApi
--qualifier prod0 replies
A thoughtful contribution can make all the difference. Be the first to share one.