The opportunity
The Site Reliability Engineering Lead (SREL) is accountable for the reliability, health, and operational sustainability of a broader portfolio of applications, data pipelines, platforms, and services. The role proactively identifies systemic and recurring risks and leads or directs technical initiatives, including the design and implementation of monitoring, automation, operational support tooling, and selected code to reduce incidents, manual intervention, cost, and operational burden. The SREL serves as a senior technical advisor and escalation point, translating complex issues, options, risks, and trade-offs into clear business language for stakeholders.
Who You'll Work With
- Reports to: Director, Product / Data Enablement
- Works closely with: Site Reliability Engineering Specialists, Product Engineering Leads (PEL), Data Solutions Leads, Data Platform Leads, Enterprise Architecture, Security, Technology Services, business stakeholders, and third-party vendors.
- Collaborates regularly with Product Engineering, Data Solutions, and Data Platform teams to support smooth transitions from build to run and to identify and resolve recurring reliability and supportability issues. Depending on the nature and complexity of the change, the Lead may implement or direct selected code, configuration, automation, or process changes through development, testing, release, and post-implementation validation, or work with the appropriate team to secure prioritization, ownership, and completion.
- Partners with Technology Services teams including Infrastructure, Cloud, and Platform teams to troubleshoot complex and cross-service issues and to design and implement, or direct the implementation of, observability, automation, resilience, and operational support tooling that improves service quality and reduces operational burden.
- Frequent communication with business stakeholders is required to provide transparency into system health, incident status, systemic risks, and improvement priorities, adapting the level of technical detail to the audience, including senior business stakeholders when required.
What You'll Do
Service Reliability Strategy and Ownership
- Own and continuously improve the reliability, supportability, and operational efficiency of a broad portfolio of applications, data pipelines, platforms, and services including availability, latency, performance, resilience, and cross-service dependencies.
- Define and govern Service Level Availability (SLAs), and related reliability metrics for the portfolio in accordance with enterprise standards and business expectations and drive changes where performance or operational risk warrants.
- Establish and maintain a prioritized reliability roadmap incorporating application, data pipeline and service health plans balancing business criticality, operational risk, support effort, capacity, cost, and technology priorities
- Use deep technical and trend analysis across incidents, alerts, service performance, capacity, storage, cost, and support effort to identify systemic failure modes and lead cross-service improvement initiatives through implementation and outcome validation.
Operational Excellence
- Lead response to high-impact, complex, or cross-service incidents, coordinating technical teams and ensuring timely audience-appropriate stakeholder communication.
- Drive Root Cause Analysis and problem-management activities beyond minimum process requirements, identify systemic technical, architectural, monitoring, process, and organizational contributing factors, and challenge recurring issues that have become normalized as operational support work. Ensure corrective or preventive actions are completed and validated through permanent remediation, automated mitigation, or appropriately authorized risk acceptance.
- Own the end-to-end resolution of selected recurring production defects and supportability issues. Where appropriate, implement or direct code, configuration, scripting, automation, and process changes through technical review, testing, release, and post-implementation validation, and drive changes requiring Product or Data ownership through prioritization and completion.Lead portfolio-level capacity, performance, resilience, and recovery analysis and improve Mean Time to Detect (MTTD) and Mean Time to Restore through monitoring, diagnostics, runbooks, automation, and operational learning.
- Contribute to Disaster Recovery (DR) and resilience plans and participate in recovery testing.
Observability & Automation
- Design and implement, or direct the implementation of, portfolio-wide monitoring, observability, event-management, diagnostic, automation, self-healing, and operational support tools that improve support effectiveness and service quality.
- Assess and rationalize existing alerts, automated emails, service checks, dashboards, automated restarts, runbooks, and integrations to reduce duplication and noise, simplify support processes, and clarify ownership.
- Define target-state standards and reusable patterns for telemetry, dashboards, alerting, event correlation, diagnostics, and automation, and oversee adoption, technical quality, documentation, and sustainable support arrangements.
- Measure the effectiveness of monitoring and automation improvements through reductions in non-actionable alerts, recurring incidents, manual intervention, support effort, Mean Time to Detect, Mean Time to Restore, and avoidable technology cost.
Production Readiness & Governance
- Establish minimum portfolio standards for production readiness, supportability, monitoring, resilience, runbooks, operational documentation, service ownership, and recovery capabilities.
- Lead or provide technical assurance for complex or high-risk production-readiness reviews, engaging Product Engineering, Data Solutions, Data Platform, Architecture, Security, and Technology Services teams, as appropriate, during design reviews to identify and address reliability, resilience, and supportability risks.
- Ensure unresolved production-readiness gaps have documented remediation plans, compensating controls, or appropriately authorized risk acceptance before release.
- Ensure changes implemented or led by the SRE function follow established source-control, peer-review, testing, security, change-management, release, rollback, documentation, and production-validation requirements.
- Monitor and govern vendor service performance against defined SLAs driving escalation and service performance plans through the appropriate vendor and portfolio governance channels.
Stakeholder Management
- Present portfolio health, reliability trends, systemic risks, operational burden, and improvement progress in team and governance forums.
- Act as the senior technical escalation point for major or cross-service production incidents.
- Translate complex technical risks, business impacts, solution options, costs, dependencies, and trade-offs into clear recommendations for business and technology stakeholders, including senior stakeholders when required.
- Build trusted relationships and use operational evidence to influence Product, Data, Technology Services, Architecture, and vendor teams to prioritize corrective and preventive reliability work.
- This role does not carry formal people management accountability but provides functional and technical leadership across the production support function, including setting reliability priorities and standards, coordinating work, and reviewing technical approaches and outcomes.
- The SREL provides work direction, coaching, and mentoring to Site Reliability Engineering Specialists and operational support resources, strengthening diagnostic, automation, and independent problem-solving capability and assuring the technical quality and completion of reliability work performed under its direction. Formal performance management and employment decisions remain with the applicable people leader.
The role has defined autonomy to determine the appropriate remediation approach for reliability and supportability issues.
This May Include Operational Process Changes, Monitoring, Automation, Configuration Changes, Code Changes, Platform Improvements, Or Escalation Into a Product Or Data Team Backlog Including
- Establishing monitoring, alerting, automation, diagnostic, and support-tooling standards and determining technical approaches for approved reliability initiatives.
- Prioritizing portfolio reliability and operational burden reduction initiatives based on business criticality, service risk, recurring incidents, support effort, capacity, cost, and available resources.
- Leading incident and problem management strategies, including determining when recurring issues require permanent remediation, automated mitigation, compensating controls, or formal risk acceptance.
- Determining whether selected changes can be implemented directly or through SRE and operational support resources, or require Product, Data, Technology Services, Architecture, Security, or vendor ownership, and driving the agreed change through development, testing, release, and validation.
- Escalating material financial, regulatory, architectural, funding, or enterprise-risk decisions through the appropriate leadership and governance channels.
What You'll Need
- Bachelor’s degree in Computer Science, Engineering, or related field (or equivalent combination of education and experience).
- 8+ years of experience in Site Reliability Engineering, production engineering, platform engineering, application support, or technology service delivery roles.
- Advanced understanding of application, data, and platform architectures, cloud platforms, and enterprise systems.
- Strong technical leadership, collaboration, communication, facilitation, and influencing skills, with ademonstrated ability to deliver outcomes across multiple teams without direct reporting authority and tailor communications for technical, business, and senior stakeholder audiences.
- Ability to navigate ambiguity, manage competing priorities, exercise sound technical judgement, and lead complex, high-impact outcomes under pressure.
- Advanced knowledge of Site Reliability Engineering principles and practices, including Service Level Availability (SLAs), observability, incident and problem management, capacity and performance engineering, resilience, disaster recovery, operational automation, and reduction of repetitive operational work.
- Advanced knowledge of enterprise monitoring, observability, event-management, diagnostic, automation, and operational support-tooling capabilities and integration patterns, including tools such as Dynatrace.
- Strong understanding of source control, CI/CD pipelines, automated testing, deployment automation, change management, release management, rollback, and production-validation practices.
- Deep understanding of application, data, database, platform, cloud, infrastructure, integration, and distributed-system architectures and their common failure modes.
- Knowledge of capacity management, data and technology lifecycle considerations, cloud storage and cost optimization, resilience, and recovery concepts.
- Hands-on experience operating and improving highly available production systems across multiple applications, data pipelines, platforms, or services.
- Experience leading high-impact and cross-service incident response, deep technical investigation, root cause analysis, problem elimination, and corrective-action programs.
- Experience designing and implementing, or directing the implementation of, monitoring, observability, automation, diagnostic, and operational support capabilities across multiple services.
- Experience implementing or leading code, configuration, scripting, or automation changes through source control, technical review, testing, release, and post-release validation.
- Experience leading ambiguous, cross-service technical initiatives from problem definition and requirements clarification through solution analysis, implementation, adoption, and benefits validation, including influencing Product, Data, Technology Services, Architecture, Security, and vendor teams without formal authority.
- Demonstrated experience applying and guiding the effective use of approved AI-assisted engineering tools, including IDE-integrated assistants and Model Context Protocol (MCP)-enabled integrations, to perform complex cross-service investigations; correlate application, database, code, telemetry, and infrastructure context; develop and validate remediation strategies; and improve incident response, technical quality, and engineering productivity.
- Expertise in delivery methodologies (Agile, Waterfall, DevOps) and IT service management frameworks (ITIL, COBIT).
,
What We’re Offering
- Numerous opportunities for professional growth and development
- Comprehensive employer paid benefits coverage
- Retirement income through a defined benefit pension plan
- The opportunity to invest back into the fund through our Deferred Incentive Program
- A flexible work environment combining in office collaboration and remote working
- Competitive time off
- Our Flexible Travel Program gives you the option to work abroad in another region/country for up to a month each year
- Employee discount programs including Edvantage and Perkopolis
At Ontario Teachers', diversity is one of our core strengths. We take pride in ensuring that the people we hire and the culture we create, reflect and embrace diversity of thought, background and experience. Through our Diversity, Equity and Inclusion strategy and our Employee Resource Groups (ERGs), we celebrate diversity and foster inclusion through events for colleagues to connect for professional development, networking & mentoring. We are building an inclusive and equitable workplace where our talent is respected, accepted and empowered to be themselves. To learn more about our commitment to Diversity, Equity and Inclusion, check out Life at Teachers'.
How To Apply
Are you ready to pursue new challenges and take your career to the next level? Apply today! You may be invited to complete a pre-recorded digital interview as part of your application.
Accommodations are available upon request (peopleandculture@otpp.com) for candidates with a disability taking part in the recruitment process and once hired.
Candidates must be legally entitled to work in the country where this role is located.
Ontario Teachers’ may use AI-based tools to assist in screening and assessing applicants for this position. These tools may help us identify candidates whose skills and experience align with Ontario Teachers’ objectives by analyzing information provided in resumes and applications. Our use of AI does not replace human decision-making.
To learn more about how Teachers’ uses AI with your personal information, please visit our Privacy Centre.
Functional Areas
Information Technology
Vacancy
Current
Requisition ID
7284