Jobs
No opportunities available
There are currently no job opportunities available. Please check back later.
Infrastructure SRE - HPC
Infrastructure SRE - HPC
Sarvam AISarvam is seeking an Infrastructure SRE to manage a large, multi-vendor GPU fleet, ensuring the reliability of both long-running training jobs and low-latency inference services. This specialized role focuses on solving complex reliability problems beyond standard Kubernetes administration, dealing with challenges in parallel filesystems, RDMA fabrics, and heterogeneous hardware. Candidates should have 5+ years in infrastructure or SRE, with 2+ years operating GPU clusters at scale, and proficiency in Python or Go for tooling development. The ideal candidate will be a specialist with deep domain expertise in one of five focus areas, capable of triaging issues across the entire GPU fleet.