Quota Isn't Enough: Topology Aware Scheduling on Modern GPU Clusters
About this session
Rack-scale systems are here! Fragmentation is widening the gap between “I have GPU quota” and “my job is actually running.” NVLink domains, rack boundaries, and the physical layout of the data center now shape whether a large training job can start, where it lands, and how efficiently it runs. On modern GPU clusters, quota is no longer enough. Topology has become a scheduling requirement.
Kueue’s Topology Aware Scheduling (TAS) brings that reality into Kubernetes. In this session, you’ll learn how TAS works and the configuration choices that matter, including preferred versus required topologies and PodSet slices. You’ll also discover how topology-aware placement affects workload performance. Finally, you’ll walk away with the production lessons that do not show up in the happy path: sticky scheduling, quota debugging, fragmentation, and node hot swaps.
Join us to understand how topology-aware scheduling helps AI teams turn expensive GPU capacity into actual training progress.
Speakers
Share this session


Get in the room. San Francisco, September 29.
Fully Connected 2026 is where the engineers, leaders, and operators running AI in production come together for three days of depth, access, and real conversation.