Summary
- Location: Montreal (day 1 onboarding / onsite presence required 3x/week)
- Duration: 12 Months Contract
- Schedule: On-call weekend support might be required (1-2 hours for weekend deployments). It is on a rotation and always communicated in advance.
Responsibilities
- Design, build, maintain, and enhance CI/CD pipelines and supporting build, test, release, and deployment infrastructure.
- Develop and maintain automated deployment solutions for applications running in Kubernetes and containerized environments.
- Partner with application development teams to improve build, test, deployment, and release processes.
- Support the deployment of applications, configuration changes, patches, and platform upgrades across development, testing, and production environments.
- Build and maintain Kubernetes deployment artifacts, including YAML configuration, Helm charts, or equivalent packaging and configuration mechanisms.
- Develop automation using Python, Linux shell scripting, and related tools.
- Manage and improve GitHub repositories, branching strategies, pull-request workflows, access controls, and automated repository processes.
- Integrate automated testing, code-quality validation, security scanning, dependency checks, and linting into CI/CD pipelines.
- Troubleshoot application, infrastructure, container, Kubernetes, network, and deployment issues in complex environments.
- Investigate production incidents, identify root causes, and implement preventive or corrective engineering solutions.
- Apply SRE principles to improve system reliability, availability, scalability, observability, and operational readiness.
- Define and improve monitoring, alerting, dashboards, operational metrics, and production support procedures.
- Automate routine operational activities to reduce manual effort and operational risk.
- Collaborate with infrastructure, network, database, cybersecurity, release management, and application teams.
- Create and maintain technical documentation, operational runbooks, deployment procedures, and troubleshooting guides.
- Participate in design reviews, production-readiness reviews, incident reviews, and continuous-improvement initiatives.
- Participate in an after-hours production support and on-call rotation when required.
Requirements
- Minimum of 5 years of relevant experience in DevOps, SRE, production engineering, platform engineering, infrastructure engineering, or a related discipline.
- Strong hands-on experience with Kubernetes, including application deployment, configuration, troubleshooting, scaling, services, ingress, secrets, and operational support.
- Strong hands-on experience with Docker/Podman and containerized application environments.
- Strong Linux and UNIX system administration and troubleshooting skills.
- Strong experience with Linux shell scripting, such as Bash or KornShell.
- Hands-on programming and automation experience using Python or a comparable language.
- Strong understanding of CI/CD concepts, software delivery lifecycles, release automation, and deployment strategies.
- Hands-on experience with CI/CD and artifact-management tools such as Jenkins and Artifactory, or equivalent platforms.
- Hands-on experience with Git and GitHub, including repository management, pull requests, branching strategies, release workflows, and automated checks.
- Experience developing automation using YAML, Ansible, or an equivalent automation framework.
- Strong understanding of networking concepts, including DNS, TCP/IP, HTTP/HTTPS, TLS, proxies, firewalls, routing, load balancing, and network troubleshooting.
- Experience supporting applications and resolving production issues in a fast-paced environment.
- Understanding of SRE practices, including monitoring, incident response, root-cause analysis, service reliability, operational readiness, and continuous improvement.
- Experience integrating code-quality tools, security scanning, automated testing, and policy controls into CI/CD pipelines.
- Ability to troubleshoot issues across application, operating system, container, network, infrastructure, and database layers.
- Experience working in an Agile development environment and using tools such as Jira.
- Strong written and verbal communication skills.
- Ability to collaborate effectively with globally distributed engineering and support teams.
- Ability to prioritize work, manage multiple tasks, and deliver results with limited supervision.
- Ability to participate in an after-hours on-call support rotation.
Preferred Skills
- Experience with enterprise Kubernetes platforms such as OpenShift or another managed Kubernetes environment.
- Experience provisioning on-demand environments using virtual machines and containers.
- Experience working with Azure, AWS, GCP, or another cloud platform.
- Experience with infrastructure-as-code and configuration-management technologies.
- Experience with Kubernetes package-management and deployment tools such as Helm.
- Experience with GitOps deployment models and tools.
- Knowledge of service mesh technologies, container networking, ingress controllers, and API gateways.
- Experience with observability and telemetry platforms, including metrics, logs, traces, dashboards, and alerting.
- Experience developing or maintaining Grafana dashboards.
- Familiarity with production incident management, problem management, change management, service operations, and release management.
- Understanding of high availability, disaster recovery, capacity management, and production resiliency.
- Experience with relational database technologies such as DB2, Sybase, or Oracle.
- Experience working in financial services or another regulated enterprise environment.
- Familiarity with secure software supply-chain practices, secrets management, certificate management, and vulnerability remediation.
- Education: Bachelor’s degree in Computer Science, Engineering, Information Technology, or a related discipline, or equivalent practical industry experience.
This role is for an existing vacancy.