Diagnosing VTTablet NOT_SERVING in Vitess Clusters
When a VTTablet stops serving reads in a Vitess cluster, this guide walks through the four most common causes—type mismatch, cell misconfiguration, replication lag, and process crashes—with an ordered diagnostic table and targeted fixes.
20 Feb 2026, 20:59 UTC

Recognizable condition
You check a Vitess cluster and a VTTablet reports NOT_SERVING or OFFLINE. VTGate stops routing reads to that shard, queries time out, or applications return tablet not serving errors. This state usually stems from one of four well‑documented conditions, each with a distinct diagnostic path.
Cause and diagnostic table
| Cause | Symptom | Quick check |
|---|---|---|
| Tablet type mismatch (e.g., replica configured as primary) |
VTGate routes queries away; tablet not serving or timeout errors in application logs |
vctl tablet_explain <keyspace>/<shard>/<tablet> and confirm the type field matches the intended role |
| Cell label misconfiguration | vttablet registers in the wrong topology cell; status shows OFFLINE or NOT_SERVING across the cluster |
vctl tablet_explain and compare the cell value against the expected cell for the keyspace/shard |
| Replication lag exceeds max_lag threshold | vttablet automatically enters NOT_SERVING mode to protect consistent reads from stale data |
Check vttablet tablet.log for ReplicationLag; compare against the configured max_lag value |
| vttablet process crash or restart loop | In‑memory state resets; tablet repeatedly cycles through OFFLINE and NOT_SERVING |
Inspect process health: systemctl status vttablet or ps aux | grep vttablet; look for repeated restart patterns |
Ordered diagnostic checks
- Confirm the tablet’s serving status and type. Run
vctl tablet_explain <keyspace>/<shard>/<tablet>on any Vitess admin workstation. Verify theserving,type, andcellfields. Iftypedoes not match the expected role (primary vs replica), the mismatch is the root cause. - Check cell assignment. If the type is correct, confirm that the
cellfield aligns with the keyspace’s cell topology. A mismatched cell prevents vttablet from registering in the correct replication group, producing anOFFLINEstatus. - Inspect replication lag. Open
vttablet tablet.log(typically at/var/log/vitess/vttablet/on the tablet host). Locate theReplicationLagentry. Compare it to the shard’smax_lagconfiguration (set viavtctl AlterShard). If lag exceeds the threshold, the tablet will self‑serve asNOT_SERVINGuntil lag subsides. - Verify vttablet process health. If the first three checks show no configuration issue, confirm the vttablet process is running without crash loops. Use
systemctl status vttabletor examine operating system logs for OOM killer events, signal handling, or goroutine starvation. A restart loop will reset in‑memory topology and require repair.
Fixes tied to findings
- Tablet type mismatch: Correct the tablet type using
vtctl FixTabletType <keyspace>/<shard>/<tablet> --tablet-type <primary|replica>. After the type is updated, runvctl RefreshTabletto re‑register the tablet in topology. VTGate will then route reads according to the corrected role. - Cell label misconfiguration: Relabel the tablet to the correct cell with
vtctl RelabelTablet <keyspace>/<shard>/<tablet> --cell <correct-cell>. Follow withvctl RefreshTabletso that VTGate and the topology server recognize the new cell assignment. - Replication lag above max_lag: Investigate the source replica’s binary log flow, network connectivity, and disk I/O. Once lag drops below
max_lag, vttablet will automatically transition back toSERVING. No manual topology change is required; the decision is driven by the built‑in lag guard. - vttablet process crash or restart loop: After stabilizing the host environment (e.g., memory, disk space, signal handling), repair the tablet state with
vtctl RepairTablet <keyspace>/<shard>/<tablet>. This command re‑inserts the tablet into the topology with its last known serving state. If the process continues to crash, diagnose host‑level issues before retrying.
Escalation criteria
Proceed to escalation if:
- The tablet remains
NOT_SERVINGafter correcting type, cell, and replication lag. - Goroutine dumps or core dumps from the vttablet process show panics, deadlocks, or unhandled signals that correlate with the status change.
- Multiple tablets in the same shard are affected, suggesting a broader topology or replication infrastructure issue.
At this stage, open a Vitess issue with vtctl DebugTablet output, tablet logs from multiple time windows, and the cluster’s topology snapshot.
Limitations and practical verification
The replication‑lag‑driven NOT_SERVING guard behavior changed between Vitess releases. In Vitess 16.x, the evaluation logic was refined to include a grace period before entering NOT_SERVING; earlier releases may trigger the state immediately upon crossing max_lag. Always cross‑reference your cluster’s Vitess version with the release notes for the exact lag‑threshold behavior.
After applying any fix, verify the resolution by running:
vctl tablet_explain <keyspace>/<shard>/<tablet>
Confirm that serving reads true, the type matches the intended role, and the cell matches the expected topology cell. Then issue a consistent read query through VTGate and observe that it returns without tablet not serving errors.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.