Monal Gupta

Signal

2022 · Python · CUDA · Kubernetes

Signal allocates GPU time across tenants with wildly different request profiles, from millisecond embeddings to multi-minute batch jobs. It implements a deficit round-robin scheduler with preemption and warm-model affinity.

Model swap costs are amortized by routing requests to nodes that already hold the weights resident, which raised cluster utilization from 44% to 81%.