Using Traefik Mesh TrafficSplit for Safe Canary Deployments
Learn how Traefik Mesh TrafficSplit enables weighted canary releases without code changes, including setup, verification via Prometheus, and key limitations.
09 May 2026, 02:24 UTC

Problem: releasing a new version without risking user impact
Teams often need to shift a small fraction of live traffic to a new service version to validate behavior before a full rollout. Doing this safely requires a mechanism that can split requests based on weight, work across protocols (HTTP, gRPC, TCP), and respect the mesh’s mutual TLS settings—all without changing application code.
Thesis: Traefik Mesh’s TrafficSplit CRD provides declarative, layer‑7 weighted routing that integrates directly with the existing sidecar proxy
TrafficSplit lets you define what percentage of requests go to each MeshService version. The control plane programs the sidecar proxies (Envoy‑based) to enforce those weights using the same routing rules Traefik Proxy already applies for HTTP, TCP, and gRPC. When mTLS is enabled, only pods presenting valid mesh‑issued certificates receive traffic, preserving security.
How TrafficSplit works
The TrafficSplit custom resource references one or more MeshService objects and assigns a weight (0‑100) to each. The mesh control plane reconciles the resource and updates the proxy configuration. Traffic is split at the request level, not at the connection level, which gives fine‑grained control for HTTP/gRPC but requires additional TLS termination considerations for pure TCP.
Prerequisites: services must be MeshServices
Only services annotated with traefikmesh.io/mesh-service: "true" are recognized by the control plane. A plain Kubernetes Service without this annotation is ignored, so you must first promote each version to a MeshService.
apiVersion: traefikmesh.io/v1alpha1
kind: MeshService
metadata:
name: myapp-v1
namespace: production
spec:
service:
name: myapp-v1 # references a regular K8s Service
namespace: production
---
apiVersion: traefikmesh.io/v1alpha1
kind: MeshService
metadata:
name: myapp-v2
namespace: production
spec:
service:
name: myapp-v2
namespace: production
Defining a TrafficSplit for weighted canary
Once both versions are MeshServices, you can create a TrafficSplit that sends, for example, 90 % of traffic to v1 and 10 % to v2.
apiVersion: traefikmesh.io/v1alpha1
kind: TrafficSplit
metadata:
name: myapp-split
namespace: production
spec:
service: myapp # the logical service name clients use
backends:
- service: myapp-v1
weight: 90
- service: myapp-v2
weight: 10
Apply the resource with kubectl apply -f traffic-split.yaml. The control plane will reconcile; you can watch progress with kubectl get trafficsplit myapp-split -o yaml and look for the status.currentWeight fields.
Worked example: validating the split with Prometheus metrics
- Deploy two versions (v1 and v2) of a simple HTTP service, each exposing a distinct response header or body (e.g.,
X-Version: v1). - Create MeshService objects for each deployment as shown above.
- Apply the TrafficSplit with 90/10 weighting.
- Generate traffic from a test pod or using
kubectl run -i --tty curl --image=appropriate/curl -- sh -c 'while true; do curl -s http://myapp.production.svc.cluster.local/; sleep 0.2; done'. Run this in the same namespace or with appropriate service‑lookup. - Check the metric
mesh_traffic_split_requests_totalin Prometheus. The querysum by (backend) (mesh_traffic_split_requests_total{trafficsplit="myapp-split"})should return values roughly in a 9:1 ratio. - Observe drift: if the ratio deviates significantly, investigate proxy reconciliation delays or mis‑matched MeshService selectors.
No output is invented; the actual numbers will depend on your traffic volume and the timing of proxy updates.
Trade‑offs and limitations
- MeshService requirement: TrafficSplit ignores plain Kubernetes Services. You must annotate or convert each version.
- Layer‑7 focus: Weighted routing is applied at HTTP/gRPC/TCP level. Pure TCP splits (e.g., raw TLS passthrough) need extra termination configuration and are less granular.
- Reconciliation latency: Changing weights triggers a control plane update; large abrupt shifts can cause a brief burst of traffic to the new version until all sidecars pick up the new config.
- Observability dependency: Verifying the split relies on exposing the Prometheus metric; if metrics scraping is disabled, you must rely on logs or external tracing.
Actionable closing: adopting TrafficSplit safely
- Start with a low weight (e.g., 5 %) on the new version and monitor the
mesh_traffic_split_requests_totalmetric for at least a few minutes. - If the observed ratio matches the target and error rates remain stable, increase the weight in incremental steps (10 %, 25 %, 50 %).
- After each step, check both the metric and application‑level SLOs (latency, error rate).
- If you need to roll back, simply set the weight of the new version to 0 and re‑apply; the sidecars will revert to sending all traffic to the stable version.
- Document the MeshService annotations and TrafficSplit definitions in your GitOps repository so the split can be reproduced across environments.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.