Using Traefik Mesh Weighted Traffic Splitting for Safe Canary Deployments
Learn how to shift a fraction of traffic to a new service version with Traefik Mesh’s TrafficSplit resource, verify the split with Prometheus metrics, and understand the operational constraints.
22 Jun 2026, 20:54 UTC

Problem: Rolling out a new API version without risking all users
When you need to expose a v2 release of a service while keeping the majority of traffic on the stable v1, you want a mechanism that:
- does not require code changes or redeploying the applications,
- can be adjusted at runtime, and
- provides observable feedback that the split is behaving as expected.
Traefik Mesh’s TrafficSplit custom resource satisfies these requirements by letting the sidecar proxies route a configurable percentage of requests to each destination service.
Thesis: A declarative TrafficSplit resource enables safe, observable canary rollouts
By defining a TrafficSplit that assigns weights to two MeshServices (v1 and v2), the data plane enforces the split. The built‑in Prometheus exporter exposes request counters per destination, allowing you to confirm the ratio in real time.
Worked example: 90/10 split between v1 and v2 of a web API
Assume both versions are already deployed in the same namespace and registered in Consul as api-v1 and api-v2. The steps below illustrate the creation of the necessary MeshService and TrafficSplit objects.
1. Define MeshServices
apiVersion: mesh.traefik.io/v1alpha1
kind: MeshService
metadata:
name: api-v1
namespace: production
spec:
service: api-v1 # matches the Consul service name
---
apiVersion: mesh.traefik.io/v1alpha1
kind: MeshService
metadata:
name: api-v2
namespace: production
spec:
service: api-v2
Apply the manifests with kubectl apply -f meshservices.yaml (requires cluster admin or delegated MeshService permissions).
2. Create the TrafficSplit
apiVersion: mesh.traefik.io/v1alpha1
kind: TrafficSplit
metadata:
name: api-canary
namespace: production
spec:
service: api-v1 # the base service that clients call
backends:
- destination: api-v1
weight: 90
- destination: api-v2
weight: 10
Apply with kubectl apply -f traffic-split.yaml. The sidecar proxies for any pod that references the api-v1 MeshService will now route ~90% of requests to the v1 endpoint and ~10% to v2.
3. Verify the split
Generate load against the virtual hostname exposed by Traefik Mesh (e.g., api.production.mesh) using a tool like hey:
hey -z 1m -c 10 http://api.production.mesh/endpoint
While the test runs, query Prometheus:
sum by (destination) (rate(mesh_traffic_split_requests_total{destination=~"api-v1|api-v2"}[1m]))
You should see approximately a 9:1 ratio between the api-v1 and api-v2 series. If the Traefik Mesh dashboard is enabled, the split percentage for the service will also be displayed live.
Trade‑off and limitation
The weighted routing logic resides in the data plane and therefore depends on two assumptions:
- Mutual TLS (mTLS) must be enabled between sidecars; if mTLS is disabled, the proxy falls back to sending all traffic to the first destination.
- Both destination services must be present in the same Consul catalog (or the same namespace when using the built‑in service discovery). Cross‑namespace splits require an additional
ServiceMeshPolicyto authorize traffic.
If Consul connectivity is lost, the proxies cannot resolve the destination services and will default to a single version, potentially breaking the canary intent.
Actionable closing
To adopt this pattern safely:
- Confirm mTLS is enforced across the mesh (
traefikmesh --set global.mtls.enabled=true). - Deploy both service versions and register them in Consul.
- Apply the MeshService and TrafficSplit manifests.
- Run a controlled load test and watch the
mesh_traffic_split_requests_totalmetric for the expected ratio. - Gradually increase the weight for v2, observing the metric at each step, until you reach 100% or decide to roll back.
By treating the TrafficSplit as the single source of truth for the split ratio, you can adjust the canary without touching application code, while still having real‑time observability to catch any deviation early.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.