DevOps, Kubernetes and Site Reliability Engineer
📍 Location: Montreal, QC
🏢 Work arrangement: Day 1 onboarding and 3 days on site per week
💼 Experience: 5 to 7 years
Position summary
Our client is looking for a DevOps, Kubernetes and Site Reliability Engineer to join an Operations Technology team within a global financial services environment.
The successful candidate will help design, automate, deploy, and support reliable application platforms and CI/CD capabilities. This role involves close collaboration with application development, infrastructure, network, cybersecurity, database, and production support teams.
Main responsibilities
- Design, maintain, and improve CI/CD pipelines and the infrastructure supporting build, testing, release, and deployment activities.
- Develop and maintain automated deployment solutions for applications running in Kubernetes and containerized environments.
- Create and manage Kubernetes deployment artifacts, including YAML files, Helm charts, configuration, secrets, services, and ingress.
- Develop automation using Python, Bash, KornShell, Ansible, and related tools to reduce manual operational work.
- Troubleshoot application, Linux, container, Kubernetes, network, and deployment issues, including production incidents and root-cause analysis.
- Improve reliability through monitoring, alerting, dashboards, operational metrics, technical documentation, runbooks, production-readiness reviews, and continuous improvement initiatives.
Qualifications
- 5 to 7 years of experience in DevOps, SRE, production engineering, platform engineering, infrastructure engineering, or a related field.
- Strong hands-on experience with Kubernetes, Docker or Podman, including deployment, configuration, scaling, services, ingress, secrets, and operational support.
- Strong Linux and UNIX troubleshooting skills, with practical experience using Bash, KornShell, Python, or a comparable programming language.
- Experience with CI/CD and artifact-management tools such as Jenkins, Artifactory, GitHub, GitHub Actions, or equivalent platforms.
- Strong knowledge of Git, GitHub workflows, YAML, Ansible, software delivery lifecycles, release automation, and automated testing or security checks.
- Good understanding of networking and production support, including DNS, TCP/IP, HTTP/HTTPS, TLS, proxies, firewalls, routing, load balancing, monitoring, incident response, and SRE practices.
Experience with OpenShift, Helm, GitOps, cloud platforms, Grafana, infrastructure as code, database technologies, financial services, or other regulated environments is an advantage. A bachelor’s degree in Computer Science, Engineering, Information Technology, or a related field is preferred.