Pod distribution drift occurs when soft topology spread constraints in Amazon EKS become uneven after node availability changes, leading to imbalanced workloads across Availability Zones. The Kubernetes descheduler automatically identifies and evicts overloaded pods to rebalance distribution without downtime or requiring hard constraints. This post explains how to deploy and configure the descheduler to maintain even pod spread in EKS clusters.


Why pod distribution drifts in Amazon EKS
Soft topology spread constraints in Kubernetes prefer even pod distribution across zones but do not enforce it strictly. When a node becomes unavailable due to spot interruption or scaling events, pods on that node are rescheduled elsewhere, often concentrating in remaining zones. Once the node returns, the scheduler does not automatically move pods back to restore balance, causing drift over time.
This drift leads to uneven resource utilization, where some zones run hot while others remain underused. For stateful or latency-sensitive workloads, this imbalance can increase costs and reduce fault tolerance. Monitoring shows pod counts per zone diverging from the ideal spread, even when constraints are defined.
The issue is not a misconfiguration but a limitation of the default scheduler: it optimizes for placement at pod creation time, not ongoing rebalancing. Without intervention, clusters gradually accumulate distribution skew that worsens with frequent node churn.
How the Kubernetes descheduler restores balance
The Kubernetes descheduler is a cluster add-on that runs as a pod and evaluates node utilization and pod distribution against policies. It identifies pods that violate desired spread constraints and evicts them, allowing the scheduler to reschedule them to underutilized zones. This process respects pod disruption budgets and avoids downtime.
Unlike hard constraints, which can block scheduling during capacity shortages, the descheduler works reactively and only when imbalance exceeds a threshold. It uses the same topology spread logic as the scheduler but acts as a corrective force. Evicted pods are typically replaced within seconds by healthy replicas.
In Amazon EKS, the descheduler integrates with existing workloads and requires no changes to pod specifications. It runs periodically, evaluates spread policies, and triggers evictions only when rebalancing improves overall cluster alignment with zone distribution goals.
Deploying the descheduler in Amazon EKS
Deploy the descheduler using the official AWS EKS add-on or via Helm chart from the Kubernetes-sigs repository. Configure a policy that targets topology spread constraints, such as 'RemovePodsHavingTooManyRestarts' or a custom policy focused on zone imbalance. Set eviction frequency and thresholds to match your workload’s tolerance for disruption.
Ensure the descheduler has RBAC permissions to read nodes, pods, and topology spread constraints, and to create eviction events. Use a service account with minimal required privileges. Monitor descheduler logs to verify it is detecting and acting on imbalanced pods without over-evicting.
Validate the setup by simulating a node drain in one zone and observing pod redistribution. After the node returns, confirm the descheduler gradually restores even spread. Combine with CloudWatch metrics to track zone-level pod counts over time and ensure drift remains within acceptable bounds.
What to do next
To maintain balanced workloads in Amazon EKS, deploy the Kubernetes descheduler as a corrective mechanism for topology spread drift. Monitor zone-level pod distribution, tune eviction policies to avoid unnecessary churn, and treat the descheduler as a complementary tool to—not a replacement for—sound capacity planning and node group design.
FAQ
Does the descheduler cause downtime when evicting pods?
No, the descheduler respects pod disruption budgets and evicts pods only when healthy replicas can be scheduled elsewhere, ensuring no downtime for replicated workloads.
Can I use the descheduler with stateful workloads like databases?
Use caution with stateful sets; the descheduler should be configured to avoid evicting pods with non-replicated state or local storage unless replacement is guaranteed.
Source: Fix pod distribution drift in Amazon EKS with the Kubernetes descheduler (AWS).



