Scaling in Kubernetes Safely on On-Prem KaaS Across 1,300+ Clusters and 40,000+ Nodes - S. Yoshimura

CNCF
AI summary

This talk presents a custom controller for safely scaling in underutilized Kubernetes worker nodes across 1,300+ clusters and 40,000+ nodes in an on-premises KaaS platform. It covers the engineering challenges of capacity optimization at scale, using observed usage, seven-day peak metrics, minimum replica guarantees, and machine-group-level safety checks. Platform engineers and SREs managing large Kubernetes clusters will learn practical safeguards including graceful shutdown expectations and node deletion behavior to minimize user impact.