Scaling Event-Driven Workflows with Azure Cosmos DB Change Feed
Learn how to decouple your Azure Cosmos DB writes from downstream processing using the Change Feed and Change Feed Processor (CFP) to build scalable, event-driven architectures.
03 Jul 2026, 17:02 UTC

The Challenge of Synchronizing Data in Real-Time
When your application grows, you often find that a single database write isn't enough. You might need to update a search index, send a notification, or aggregate data into a reporting dashboard every time a document changes. The naive approach—triggering these actions directly within your application code—creates tight coupling and risks data loss if a downstream service is offline during the write operation.
The Azure Cosmos DB Change Feed solves this by acting as a persistent log of changes. Instead of your application managing multiple API calls, it writes once to Cosmos DB, and a separate processor asynchronously reacts to those changes. The key takeaway is that the Change Feed enables a reliable, decoupled event‑driven architecture that scales independently of your primary write path.
How the Change Feed Processor (CFP) Works
While you can manually pull changes from the feed, the Change Feed Processor (CFP) library is the standard for production. It manages the complexity of distributing the workload across multiple worker instances.
To function, the CFP requires a Lease Container. This is a separate Cosmos DB container that acts as a state store. It tracks "checkpoints", which are markers indicating how far each worker has progressed through the change log. If a worker instance crashes, another instance reads the lease and picks up exactly where the previous worker left off, ensuring no events are missed.
Implementation: The Pull Model vs. Push Model
You generally have two architectural paths depending on how much control you need over the execution environment:
- Push Model (Azure Functions): The simplest path. You use a Cosmos DB Trigger. Azure manages the lease container and scaling logic behind the scenes. This is ideal for lightweight tasks like sending an email or updating a cache.
- Pull Model (SDK): You implement the CFP library within a custom .NET or Java application (e.g., running in AKS). This is necessary when you need fine‑grained control over batch sizes, custom retry logic, or long‑running processing tasks that would timeout in a serverless function.
Worked Example: Implementing a Custom Processor
In this scenario, we use the .NET SDK to process changes from a Orders container and move them to a Analytics container. This requires the Microsoft.Azure.Cosmos NuGet package.
// Run this in a worker service with 'Contributor' access to the Cosmos DB account
var client = new CosmosClient(connectionString);
var leaseContainer = client.GetContainer("DatabaseId", "leases");
var sourceContainer = client.GetContainer("DatabaseId", "orders");
var processor = sourceContainer.GetChangeFeedProcessorBuilder("OrderProcessor",
async (IReadOnlyCollection changes, CancellationToken cancellationToken) =>
{
// Logic to process the batch of changes
foreach (var order in changes)
{
Console.WriteLine($"Processing order: {order.Id}");
// Perform idempotent update to Analytics container here
}
})
.WithInstanceName("Worker-01")
.WithLeaseContainer(leaseContainer)
.Build();
await processor.StartAsync();
Verification: To verify the setup, insert a document into the orders container. You should see the Processing order: [ID] output in your console. To test scaling, start a second instance of the application with a different InstanceName; you will observe the lease container redistribute the partitions between the two workers.
Critical Trade‑offs and Limitations
| Limitation | Technical Impact | Practical Workaround |
|---|---|---|
| No Delete Tracking | The feed only tracks inserts and updates. | Implement a "soft delete" by adding a isDeleted: true property to the document. |
| At‑Least‑Once Delivery | Events may be delivered more than once during worker failovers. | Ensure your processing logic is idempotent (processing the same event twice has no additional effect). |
| Poison Pills | A malformed document that causes a crash will block the partition. | Wrap the processor logic in a try‑catch block and move failing documents to a Dead Letter Queue (DLQ). |
Closing Action Plan
If you are building a system that requires high reliability and scalability, avoid synchronous multi-service writes. Start by identifying your "source of truth" container and creating a dedicated lease container. For rapid prototyping, use Azure Functions; for high-throughput enterprise workloads, implement the CFP library in a containerized environment to maintain full control over your processing lifecycle.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.