Deploying Mattermost Enterprise in a High‑Availability Kubernetes Cluster
Learn how to spin up Mattermost Enterprise on Kubernetes with true high‑availability, using a shared PostgreSQL cluster, Helm, and Prometheus monitoring. Get a concrete example, trade‑offs, and next‑step checklist to keep your chat platform running 24/7.
01 Jun 2026, 22:27 UTC

Problem: Keeping Mattermost Alive When a Pod Dies
In a production chat platform, a single pod failure can bring the entire service down if the database or web tier isn’t replicated. Mattermost Enterprise offers a built‑in high‑availability (HA) mode, but the community edition only runs a single instance. The challenge is to set up a resilient cluster that survives node failures, pod crashes, and database switchover without manual intervention.
Takeaway: By deploying Mattermost Enterprise with a shared PostgreSQL cluster, a StatefulSet of web and worker replicas, and a load‑balanced ingress, you can achieve zero‑downtime scaling and automatic recovery. Helm makes the heavy lifting, while Prometheus gives you the visibility you need to act before users notice.
Architecture Overview
The HA design relies on three core components:
- PostgreSQL HA – A single source of truth that all Mattermost pods read from. Use Patroni, RDS Multi‑AZ, or a cloud‑managed cluster.
- Mattermost Web & Worker Pods – A StatefulSet with a configurable replica count. Each pod mounts its own persistent volume for logs and uploads.
- Load Balancer & Ingress – A TCP/HTTPS LB (e.g., ELB, NGINX Ingress) that forwards traffic to the web replicas, ensuring even load distribution.
All pods share the same PostgreSQL endpoint; if the primary fails, a new primary is elected, and the pods automatically reconnect. Kubernetes readiness and liveness probes guarantee that only healthy pods receive traffic.
Deploying with Helm
The official mattermost/mattermost-enterprise chart bundles all the necessary Kubernetes manifests. Below is a minimal yet production‑ready values.yaml snippet you can drop into a Helm release.
# values.yaml
replicaCount: 3
# PostgreSQL connection – use a secret for credentials
postgresql:
host: "postgresql-ha.example.com"
port: 5432
database: "mattermost"
username: "mmuser"
passwordSecret: "mattermost-postgres-credentials"
# Persistence – one volume per pod
persistence:
enabled: true
size: 20Gi
storageClass: "fast"
# Load balancer settings – external LB or Ingress
service:
type: LoadBalancer
annotations:
service.beta.kubernetes.io/aws-load-balancer-backend-protocol: https
# TLS – terminate at Ingress
tls:
enabled: true
certSecret: "mattermost-tls"
# Monitoring – expose Prometheus metrics
metrics:
enabled: true
serviceMonitor:
enabled: true
Install the chart in a dedicated namespace:
# kubectl create namespace mattermost
# helm repo add mattermost https://helm.mattermost.com
# helm repo update
# helm install mattermost-enterprise mattermost/mattermost-enterprise \
-n mattermost -f values.yaml
Verify the rollout:
# kubectl get pods -n mattermost
# kubectl get svc -n mattermost
All three web replicas should show Running and Ready. The service should expose an external IP (or DNS record) that your users can reach.
Testing High‑Availability
After deployment, simulate a failure and confirm the cluster recovers:
- Stop one web pod:
kubectl delete pod mattermost-web-0 -n mattermost. Kubernetes will recreate it automatically. - Check that the remaining pods stay healthy:
kubectl get pods -n mattermost. - Verify database connectivity from the pods:
kubectl exec mattermost-web-1 -n mattermost -- psql -h ${POSTGRES_HOST} -U ${POSTGRES_USER} -d ${POSTGRES_DB} -c 'SELECT 1;'. Expect a successfulSELECT 1;response. - Perform a smoke test: create a channel in the Mattermost UI, post a message, then delete the pod that handled the request. The message should still appear after the pod restarts.
For database HA, confirm replication status:
# kubectl exec -it mattermost-web-1 -n mattermost -- psql -h ${POSTGRES_HOST} -U ${POSTGRES_USER} -d ${POSTGRES_DB} -c 'SELECT pg_is_in_recovery();'
A result of false indicates the pod is connected to the primary; true means it’s a replica. All Mattermost pods should connect to the primary.
Monitoring with Prometheus
The chart exposes a metrics service. If you have a Prometheus Operator, a ServiceMonitor will automatically scrape the metrics. Add the following to your Prometheus scrape config if you’re not using the Operator:
scrape_configs:
- job_name: "mattermost"
static_configs:
- targets: ["mattermost-enterprise-metrics.mattermost.svc.cluster.local:9100"]
In Grafana, import the Mattermost dashboard (ID 12345 on Grafana Labs) to visualize:
- Web pod CPU/memory usage
- Database replication lag (seconds)
- Message queue depth
- HTTP response latency
Set alerting rules for:
- Replication lag > 5 s
- Web pod restarts > 3 in 5 min
- Database connection failures
Trade‑Offs and Limitations
- PostgreSQL HA is the single point of failure. If your DB cluster goes down, the entire Mattermost service is offline. Use a managed HA solution or a Patroni cluster with fencing.
- TLS termination can be at the ingress or inside the pods. If you terminate at the ingress, ensure you forward the
X‑Forwarded‑Protoheader so Mattermost generates correct URLs. - Storage per pod – the default PersistentVolumeClaim (PVC) strategy creates one volume per pod. For large teams, consider a distributed file system (e.g., Ceph, NFS) or a cloud object store for uploads.
- Load balancer limits – some cloud LB providers cap the number of backend instances. Scale the web replicas within those limits or use an Ingress controller that supports session affinity.
Next Steps
- Secure the cluster: enable RBAC, network policies, and encrypt secrets with
sealed-secretsor a cloud KMS. - Automate rollouts: use
helm upgradewith--atomicand a CI pipeline to deploy new Mattermost versions. - Integrate with your logging stack: ship pod logs to Loki or ELK for forensic analysis.
- Plan for disaster recovery: take regular snapshots of the PostgreSQL cluster and PVCs, and test restore procedures.
- Consider scaling: as traffic grows, increase
replicaCountand adjust the load balancer accordingly.
By following this guide, you’ll have a Mattermost Enterprise deployment that can survive pod, node, and even database failures, while giving you the metrics and alerts to stay ahead of issues. Happy chatting!
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.