Jobs
No opportunities available
There are currently no job opportunities available. Please check back later.
Senior Site Reliability Engineer - Fleet
Senior Site Reliability Engineer - Fleet
LambdaAbout the role
Lambda, a leader in AI cloud infrastructure, is seeking a Senior Site Reliability Engineer - Fleet to build and operate large-scale HPC clusters for AI workloads. This role involves automating cluster lifecycle, troubleshooting complex issues across various networking and hardware components, and participating in on-call rotations. The ideal candidate will have 7+ years of SRE/DevOps experience, a strong understanding of modern AI infrastructure, and proficiency in tools like Ansible, Terraform, Python, and Go. Join us to help build the world's best AI cloud and make compute as ubiquitous as electricity.
Similar jobs
Browse Jobs by Role
Chief of StaffProduct MarketingForward Deployed EngineerForward DeployedSoftware EngineerProduct ManagerData ScientistDesignSalesMarketingOperationsEngineering ManagerFinanceCustomer SuccessHR & People OpsHardware EngineerStrategyAccountingFounding EngineerGrowthDevOps & InfrastructureMachine Learning Engineer
