Livestream Tue 20 Jan

Infrastructure for AI: Managing GPU Workloads with Kubernetes

Hosted by O'Reilly Media

Saiyam Pathak speaks at O'Reilly's Infrastructure & Ops Superstream on managing GPU workloads at scale with Kubernetes. Covering GPU sharing, Kai Scheduler, multi-tenancy with vCluster, and scalable inference with vLLM.

Register free →
When
Tue 20 Jan
Format
Livestream
Speaker
Saiyam Pathak
gpukubernetesinferencemulti-tenancyai infrastructure