tag: Autoscaling · 3 items
- Platform/SRE — Learn: Covers a real incident pattern — GPU pods pending during traffic spikes — and predictive scaling approaches; worth reading to inform GPU cluster design, but no GA tool, deadline, or breaking change anchors an action now.
- CI/CD — Skip
- Leader — Skip
- Platform/SRE — Learn: EKS Provisioned Control Plane clusters now get significantly faster HPA-driven scaling with no configuration changes required — worth knowing if you run large clusters with many HPA objects, but no action needed.
- CI/CD — Skip
- Leader — Skip
- Platform/SRE — Learn: A solid walkthrough for extending HPA with custom signals (queue depth, connection counts) via a Prometheus-compatible exporter — useful design reference but no operational change required to existing clusters.
- CI/CD — Skip
- Leader — Skip