Diagnosing and Correcting Tablet Distribution Skew in YugabyteDB
Detect and correct tablet distribution skew in YugabyteDB by checking tablet counts per node, enabling the load balancer, and triggering a rebalance.
11 Jun 2026, 23:01 UTC

Recognizable Condition
One or more nodes consistently show high CPU, disk I/O, or network usage while peers stay idle.
Cause and Diagnostic Indicators
| Possible Cause | What to Look For |
|---|---|
| Hot keys or range partitioning imbalance | Specific tables show uneven tablet counts |
| Recent node add/remove without rebalance | Tablet count per node deviates after topology change |
Ordered Checks
- Log in to any master node (or a client with yb-admin) as the
yugabyteuser or a role withyb-adminexecution rights. - Run the tablet listing command:
This prints a table with columns: tablet_id, table_id, state, leader_term, etc. Look for theyb-admin --master_addresses master1:7100,master2:7100,master3:7100 list_tabletsnodecolumn (or derive fromleader_termmapping). - Calculate average tablets per node: total tablets ÷ number of nodes. Identify any node with >20 % more tablets than the average.
- Optional: cross‑check with node metrics (CPU, disk I/O, network) from your monitoring system to confirm the hot node matches the tablet skew.
Fixes Tied to Findings
- If the skew is due to a recent topology change and the load balancer is disabled, enable it:
This requires the ability to modify cluster‑wide flags (typically theyb-admin --master_addresses master1:7100,master2:7100,master3:7100 change_config_flagserver true --flag_name=enable_load_balancer --flag_value=trueyugabyteadmin user). - Start an immediate rebalance:
The command returns immediately; the balancer runs in the background.yb-admin --master_addresses master1:7100,master2:7100,master3:7100 start_load_balancer - If hot keys are the root cause, consider schema changes (e.g., sharding by a composite key) or using YugabyteDB’s tablet split feature after verifying split thresholds.
Escalation Criteria
- Tablet distribution remains outside the 10 % tolerance after a full rebalance cycle (monitor for at least 15 minutes).
- Node health deteriorates during rebalance: disk usage > 90 %, latency spikes > 2× baseline, or cluster becomes unavailable.
- Repeated skew events occur despite enabled load balancer and no topology changes.
In these cases, open a support ticket with Yugabyte, attaching:
- Output of
yb-admin list_tabletsbefore and after rebalance. - Node metric snapshots (CPU, disk I/O, network, disk usage).
- Relevant YSQL/YCQL latency dashboards.
Verification
- Re‑run
yb-admin list_tabletsand compute the new average; each node should be within ±10 % of that average. - Monitor node utilization for a minimum of 15 minutes post‑rebalance; verify that the previously hot node’s metrics now align with the cluster median.
- Check query latency graphs; latency should return to the baseline observed before the skew event.
Limitations and Practical Check
The load balancer only moves tablets; it does not change the underlying data distribution caused by hot keys. If a single key drives the majority of traffic, tablet‑level rebalance will not alleviate the load. In that scenario, consider:
- Redesigning the primary key to spread writes.
- Using uniform hash sharding (if applicable) or increasing the number of tablets per table.
To confirm that the balancer is active, you can query the flag:
yb-admin --master_addresses master1:7100,master2:7100,master3:7100 get_config_flagserver true --flag_name=enable_load_balancer
It should return true. If it returns false, the balancer will not run.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.