Pay Range: CAD 70-75/hr
We are seeking an experienced DevOps / Site Reliability Engineer (SRE) to design, implement, and maintain highly available, scalable, secure, and reliable infrastructure and application environments. The role will focus on cloud infrastructure, automation, CI/CD, observability, incident management, and continuous improvement of production systems.
Key Responsibilities
- Design, build, and maintain highly available and scalable cloud infrastructure across AWS, Azure, or GCP environments.
- Develop and manage Infrastructure as Code (IaC) using Terraform, Ansible, CloudFormation, or similar technologies.
- Build, maintain, and optimize CI/CD pipelines using Jenkins, GitHub Actions, GitLab CI, Azure DevOps, or similar tools.
- Automate application deployments, infrastructure provisioning, configuration management, and operational processes.
- Implement and manage Kubernetes and containerized environments using Docker, Kubernetes, Helm, and related technologies.
- Establish and maintain monitoring, logging, alerting, and observability solutions using tools such as Prometheus, Grafana, ELK/EFK, Splunk, Datadog, or New Relic.
- Define and monitor SLIs, SLOs, and SLAs to improve system reliability, availability, and performance.
- Participate in 24x7 production support, incident response, troubleshooting, and root cause analysis (RCA) for critical systems.
- Develop automation and self-healing capabilities to reduce manual operational effort and improve system resilience.
- Identify performance bottlenecks and implement solutions for capacity planning, scalability, availability, and disaster recovery.
- Implement security best practices across cloud infrastructure, CI/CD pipelines, containers, secrets, IAM, and production environments.
- Collaborate with Software Engineering, QA, Security, Cloud, Network, and Product teams to improve application reliability and deployment processes.
- Establish and maintain backup, disaster recovery, business continuity, and high-availability strategies.
- Manage source control, branching strategies, release management, and deployment governance using Git-based platforms.
- Continuously improve infrastructure efficiency, reliability, automation, and cloud cost optimization.
- Document architecture, operational procedures, runbooks, incident reports, and troubleshooting processes.
Required Skills
- 5+ years of experience in DevOps, SRE, Cloud Engineering, or Infrastructure Engineering.
- Strong experience with Linux/Unix administration, networking, system troubleshooting, and production environments.
- Hands-on experience with at least one major cloud platform: AWS, Azure, or GCP.
- Strong knowledge of Terraform and Infrastructure as Code.
- Experience building and managing CI/CD pipelines.
- Hands-on experience with Docker and Kubernetes.
- Proficiency with Python, Bash, PowerShell, or similar scripting languages.
- Experience with monitoring and observability platforms such as Prometheus, Grafana, ELK, Splunk, Datadog, or New Relic.
- Strong understanding of Git, branching, release management, and automated deployment practices.
- Experience with incident management, problem management, RCA, and production support.
- Understanding of SRE principles, SLI/SLO/SLA, error budgets, availability, scalability, and reliability engineering.
- Knowledge of cloud security, IAM, secrets management, vulnerability management, and secure DevOps practices.
Preferred Skills
- Experience with AWS EKS, Azure AKS, or Google GKE.
- Experience with Helm, Argo CD, Flux, or GitOps.
- Experience with Jenkins, GitHub Actions, GitLab CI, or Azure DevOps.
- Knowledge of Ansible, Packer, Vault, Consul, or similar automation tools.
- Experience with Kafka, RabbitMQ, Redis, PostgreSQL, MySQL, or other distributed systems.
- Familiarity with ServiceNow, PagerDuty, Jira, or other ITSM/incident-management platforms.
- Experience implementing zero-downtime deployments, blue/green deployments, canary releases, and automated rollback strategies.
- Knowledge of DevSecOps, vulnerability scanning, container security, and cloud compliance.