What you’ll do
- Define and track SLIs, SLOs, and error budgets with product and engineering teams, and use them to guide release and reliability decisions
- Build and run highly available, scalable services on AWS, Azure, or GCP with Kubernetes, Docker, and infrastructure-as-code (Terraform, Pulumi)
- Write software in Python, Go, or similar languages to automate toil away instead of working around it
- Build observability that tells the truth: metrics, logs, traces, and alerts using Prometheus, Grafana, Datadog, OpenTelemetry, or ELK
- Lead incident response: detect, mitigate, communicate, and learn through blameless postmortems with real follow-through
- Design for failure: capacity planning, load testing, chaos experiments, and disaster recovery drills
- Improve CI/CD pipelines and release safety with progressive delivery, canaries, and automated rollbacks
- Tune performance and cost across infrastructure without sacrificing reliability
- Cut alert noise so on-call pages are rare, meaningful, and actionable
- Write runbooks, architecture reviews, and production readiness checklists teams actually use
- Partner with developers, security, and platform teams to build reliability into services from the start
Find your level
- Level Typical experience What you’ll own
- SRE I - 0-2 years : Monitoring, automation scripts, on-call shadowing, runbooks, and well-scoped reliability projects, learning with a dedicated mentor
- SRE II - 2-5 years : Owning reliability for services end to end, leading incident response, building tooling, improving SLOs, mentoring newer engineers
- Senior SRE - 5+ years : Architecting resilient systems, setting reliability standards, leading cross-team initiatives, shaping on-call and incident culture
- Promotion criteria are written down at every level, with regular career conversations. You can grow as a deep technical expert or toward leadership. Both paths are real here.
What you bring
Must have
- Solid Linux and networking fundamentals (DNS, TCP/IP, load balancing, TLS)
- Programming or scripting ability in Python, Go, Bash, or similar
- Hands-on experience with at least one major cloud provider
- Familiarity with containers and Kubernetes concepts
- A troubleshooting mindset: you follow evidence, stay calm in incidents, and fix root causes
- Clear communication, in writing and on an incident bridge
- Degree or diploma in Computer Science, Engineering, IT, or a related field, or equivalent experience. Home labs, bootcamps, and self-taught engineers are welcome. We hire for skill, not pedigree.
- Legal authorization to work in Canada (citizens, permanent residents, and valid work permit holders)
Nice to have (not required)
- Experience with Terraform, Ansible, Helm, Argo CD, or GitOps workflows
- Observability tooling: Prometheus, Grafana, Datadog, Splunk, OpenTelemetry
- Knowledge of distributed systems, databases, and caching layers (PostgreSQL, Redis, Kafka)
- Certifications: CKA, AWS Solutions Architect or DevOps Professional, Azure AZ-400, Google Professional Cloud DevOps Engineer
- Experience with incident management tools (PagerDuty, Opsgenie) and postmortem practice
- Chaos engineering, performance testing, or capacity planning experience
- Familiarity with security and compliance practices (SOC 2, PIPEDA)
- Industry background in [fintech, SaaS, e-commerce, healthcare, telecom, gaming, etc.]
- Fluency in French, an asset for roles serving Quebec or bilingual teams [if applicable]
- A home lab, GitHub repos, or open-source contributions you can talk about