Responsibilities
Operate Kubernetes clusters at meaningful scale, including sizing node pools, debugging schedulers, troubleshooting CNI, and performing rolling upgrades across large fleets. Write and review infrastructure-as-code and develop tooling and automation.
Requirements
Requires 4+ years of experience in infrastructure engineering, cloud platforms, or HPC, with substantial hands-on Kubernetes cluster operations experience. Candidates should be proficient in Terraform, have working knowledge of AWS services including EC2, S3, EFS, and FSx for Lustre, and use Python for tooling and automation.