Job Summary
We are seeking an experienced Kubernetes (K8s) / High-Performance Computing (HPC) Engineer to design, deploy, optimize, and manage scalable compute platforms supporting AI/ML, engineering simulations, and data-intensive workloads.
The ideal candidate will have strong expertise in Kubernetes, containerization, Linux administration, HPC environments, and cloud-native infrastructure, with hands-on experience supporting distributed computing and GPU-accelerated workloads.
Key Responsibilities
- Design, deploy, and manage Kubernetes clusters and containerized applications using Docker, Helm, and related orchestration technologies.
- Support HPC clusters and distributed computing environments using workload schedulers such as Slurm and Volcano.
- Configure and optimize GPU infrastructure, particularly NVIDIA-based environments, for AI/ML workloads and high-performance computing.
- Administer Linux-based systems and develop automation scripts using Python and Bash.
- Implement and maintain CI/CD pipelines, infrastructure automation, monitoring, and performance optimization solutions.
- Deploy and manage cloud-native infrastructure on platforms such as Azure AKS, AWS EKS, and Google Kubernetes Engine (GKE).
- Troubleshoot infrastructure and workload performance issues to ensure platform reliability, scalability, and efficiency.
Required Qualifications
- 5+ years of experience in Kubernetes and cloud infrastructure engineering.
- At least 2+ years of hands-on experience with HPC or large-scale computing platforms.
- Strong knowledge of Kubernetes, Docker, Helm, and container orchestration.
- Experience with HPC clusters, workload schedulers such as Slurm or Volcano, and distributed computing.
- Hands-on experience with NVIDIA GPU infrastructure, AI/ML workloads, and performance tuning.
- Strong Linux administration and Python/Bash scripting skills.
- Experience with CI/CD, infrastructure automation, monitoring, and cloud platforms such as Azure AKS, AWS EKS, or GCP GKE.
Preferred Qualifications
- Knowledge of high-speed networking technologies, including RDMA and InfiniBand.
- Experience with parallel storage systems and large-scale compute environments.
- Background supporting engineering simulations, semiconductor workloads, AI model training, or scientific computing platforms.
Additional Requirement
Candidates must be available to start immediately and successfully complete the mandatory background check.
Interested candidates: Please share your updated resume, current location in Canada, availability to start, and relevant Kubernetes/HPC experience.