Responsibilities
Operate and scale Kubernetes platforms across cloud providers, manage HPC infrastructure and GPU job scheduling, and maintain platform stability. Define service-level indicators and objectives, build monitoring and alerting, respond to incidents, and coordinate with networking, storage, security, and AI/ML teams.
Requirements
Requires at least four years of infrastructure engineering, cloud platform, or HPC experience, with substantial hands-on Kubernetes operations experience. Candidates should be proficient in Terraform, have working knowledge of AWS services including EC2, S3, EFS, and FSx for Lustre, and use Python for tooling and automation.