We are looking for an experienced Kubernetes (K8S) / HPC Engineer to design, deploy, and manage scalable compute platforms supporting AI/ML, simulation, and data-intensive workloads.
🔹 Key Skills:
- 5+ years of Kubernetes / Cloud Infrastructure experience
- 2+ years of HPC / large-scale compute experience
• Strong hands-on experience with Kubernetes, Docker & Helm
• Experience with Slurm / Volcano workload schedulers
• Strong Linux administration and Python/Bash scripting
• Experience with NVIDIA GPU infrastructure & AI/ML workloads
- CI/CD, monitoring & infrastructure automation
• Experience with AWS EKS, Azure AKS, or GCP GKE
- AWS Trainium / Inferentia / Neuron
⭐ Nice to Have:
- Large-scale AI/ML or scientific computing environments
- Engineering simulation/semiconductor workloads