Learn
archived
Kubeflow + Cilium: debugging 60% GPU idle in distributed training
- Platform/SRE — Learn: Practical post-mortem on how Cilium networking caused GPU underutilization in Kubeflow training jobs — worth reading for anyone operating GPU clusters or eBPF-based CNIs where pod-to-pod latency affects collective communication.
- CI/CD — Skip
- Leader — Skip
This entry was curated and judged by AI (Claude) with automated enrichment
(CISA KEV / EPSS / public PoC). Verify against the original source before
acting. Found a bad verdict?
Report it —
confirmed errors go to the corrections log.