Senior DevOps Engineer / Senior Cloud Platform Engineer
Location: Downtown Toronto, ON
Work Model: Hybrid – 2 days onsite per week, 3 days remote
Contract: 6 months, with possibility of extension
Start Date: ASAP
End Client: Confidential
About the Role
We are seeking a highly hands-on Senior DevOps Engineer / Senior Cloud Platform Engineer to own and evolve a production AWS cloud platform.
This role requires deep AWS and Kubernetes expertise, with Amazon EKS being a critical requirement. We are specifically looking for someone who has personally designed, built, operated, secured, upgraded, troubleshot, and owned production EKS environments end-to-end.
This is not an application development role with some DevOps exposure, nor is it primarily a people-management position. The successful candidate must have a strong track record in core DevOps / Platform Engineering and remain deeply hands-on technically.
You will help define the DevOps roadmap, establish platform standards and reusable patterns, improve infrastructure automation and reliability, support application releases, and provide production support for critical workloads.
The role will also contribute to the practical adoption of AI within DevOps and infrastructure engineering, including AI-assisted Infrastructure as Code (IaC), automation, runbooks, operational workflows, incident response, and platform tooling.
Critical Screening Requirements
This is not an application developer position with some DevOps exposure. We require someone whose core professional background demonstrates substantial, hands-on DevOps / Platform Engineering experience, particularly with AWS and production Amazon EKS.
The resume must provide concrete project-level evidence, not simply keywords. It should clearly demonstrate:
- EKS environments the candidate personally designed, built, and owned end-to-end
- Cluster architecture, lifecycle management, upgrades, networking, security, and troubleshooting responsibilities
- Terraform/IaC solutions the candidate personally designed or implemented
- Platform standards, reusable modules, guardrails, frameworks, or engineering practices they established
- Production incidents and complex infrastructure issues they personally troubleshot and resolved
- Specific examples of applying AI to DevOps, infrastructure, automation, or platform engineering
- Evidence of technical/platform leadership while remaining hands-on
Simply listing AWS, Kubernetes, EKS, Terraform, DevOps, or AI in a technical skills section without concrete evidence of hands-on implementation and ownership will not be sufficient.
What We Need
Amazon EKS – End-to-End Ownership
- Deep, hands-on experience with Amazon EKS in production environments
- Proven experience personally designing, building, deploying, operating, upgrading, securing, monitoring, and troubleshooting EKS clusters
- Strong Kubernetes architecture and operations expertise
- Experience owning EKS environments throughout their full lifecycle
- Strong knowledge of cluster networking, security, RBAC, IRSA, secrets, observability, capacity management, upgrades, and production troubleshooting
- Experience resolving cluster-level production incidents, not simply application-level deployment issues
Candidates whose Kubernetes experience is primarily limited to deploying applications onto clusters managed by another team will not meet the core requirement.
AWS & Infrastructure as Code
- Strong hands-on AWS infrastructure experience, including multi-account and multi-environment architectures
- Advanced experience with Terraform; Terragrunt experience is highly desirable
- Experience designing reusable IaC modules and environment-driven infrastructure
- Strong understanding of AWS networking, IAM, security, secrets, encryption, tagging, and infrastructure guardrails
- Experience with infrastructure automation, reliability, observability, and production support
- Strong incident management, root-cause analysis, and troubleshooting capabilities
Platform Engineering & Technical Leadership
- Proven experience establishing DevOps/platform standards, tooling, frameworks, guardrails, and engineering practices
- Ability to define how infrastructure should be built, deployed, secured, monitored, and operated
- Experience creating reusable IaC modules, CI/CD standards, automation frameworks, runbooks, reference architectures, and platform patterns
- Experience conducting technical/design reviews and influencing DevOps or platform roadmaps
- Ability to provide technical leadership while remaining strongly hands-on
AI for DevOps / Infrastructure
We are looking for demonstrated practical exposure to applying AI within DevOps, cloud infrastructure, or platform engineering, such as:
- AI-assisted Terraform/IaC authoring and review
- AI-assisted pipeline and automation development
- Infrastructure/platform automation
- Operational documentation and runbook creation or improvement
- Incident summarization, troubleshooting, and triage
- Internal tooling for infrastructure or platform teams
- Evaluation and adoption of AI coding assistants or platform tools
- Internal LLM/RAG solutions for operational knowledge
Candidates should understand the importance of verification, security, compliance, governance, and change control when using AI in production environments.
Key Responsibilities
- Define and drive DevOps priorities across security, reliability, cost, scalability, and delivery velocity
- Establish AWS standards, guardrails, and reusable patterns across networking, identity, secrets, tagging, and infrastructure
- Design, implement, and review infrastructure using Terraform and Terragrunt
- Build and own Amazon EKS environments end-to-end, including:
- Cluster architecture, creation, and lifecycle management
- Version upgrades and patching
- Node groups and capacity management
- VPC/CNI, DNS, ingress, and networking
- RBAC, IRSA, secrets, and security controls
- Add-ons and cluster baseline tooling
- Monitoring, logging, tracing, and observability
- Performance, reliability, and cost optimization
- Production troubleshooting and cluster-level incident resolution
- Develop and maintain CI/CD pipelines and safe promotion processes across development, non-production, and production environments
- Work with AWS services including RDS/Aurora, DynamoDB, S3, CDN, KMS, Secrets Manager, SNS, Lambda, and EventBridge
- Support IAM, VPC, and multi-account/multi-environment AWS architectures
- Partner with development teams on releases, deployment strategies, rollbacks, and post-release validation
- Participate directly in production incident response, root-cause analysis, and preventive remediation
- Improve monitoring, alerting, dashboards, operational documentation, and runbooks
- Partner with security, architecture, and engineering teams on least-privilege access, encryption, backup/DR, and audit-ready operations
- Identify practical opportunities to use AI to improve infrastructure engineering and operational efficiency
Required Qualifications
- 6+ years of experience across DevOps, SRE, Platform Engineering, Systems Engineering, or related infrastructure-focused roles
- 4+ years of hands-on AWS production experience
- Deep Amazon EKS expertise – mandatory
- Demonstrated experience personally building and owning production Kubernetes/EKS environments
- Strong hands-on experience with Terraform and modular Infrastructure as Code
- Strong knowledge of Kubernetes networking, security, RBAC, IRSA, secrets, observability, capacity management, and production troubleshooting
- Strong understanding of CI/CD, artifact promotion, secrets injection, and multi-environment deployment practices
- Experience responding directly to production incidents, conducting RCAs, and implementing long-term remediation
- Experience establishing technical standards, conducting design reviews, and contributing to DevOps/platform roadmaps
- Demonstrated experience or strong practical exposure to AI-assisted DevOps or platform engineering
- Excellent communication skills and ability to collaborate across engineering, security, architecture, and development teams
Preferred Qualifications
- AWS Solutions Architect – Professional or AWS DevOps Engineer – Professional
- CKA or CKS certification
- Terragrunt experience
- Helm / Helmfile experience
- Policy-as-code and Kubernetes cluster baseline tooling
- PostgreSQL / AWS RDS experience
- Experience in regulated, enterprise, or other high-stakes production environments
- Experience defining SLOs, error budgets, or platform KPIs
- AWS cost optimization / FinOps exposure
- Experience with AI coding assistants, internal LLM/RAG solutions, or AI-enabled platform operations