Work Location:
Toronto, Ontario, Canada
Hours:
37.5
Line of Business:
Technology Solutions
Pay Details:
$96,900 - $136,800 CADThis role is eligible for a discretionary variable compensation award that considers business and individual performance.
TD is committed to providing fair and equitable compensation opportunities to all colleagues. Growth opportunities and skill development are defining features of the colleague experience at TD. Our compensation policies and practices have been designed to allow colleagues to progress through the salary range over time as they progress in their role. The base pay actually offered may vary based upon the candidate's skills and experience, job-related knowledge, geographic location, and other specific business and organizational needs.
As a candidate, you are encouraged to ask compensation related questions and have an open dialogue with your recruiter who can provide you more specific details for this role.
Job Description:
Role Summary
This first-line role supports the design, governance, assessment, and continuous improvement of Technology Resilience and Availability Management capabilities. The position works across application, infrastructure, cloud, cyber, risk, audit, and business teams to strengthen high availability, disaster recovery, backup and cyber recovery, capacity management, and operational resilience for critical technology services. The successful candidate combines strong technical depth with the ability to influence stakeholders, produce defensible evidence, and drive remediation in a complex regulated environment.
Job Responsibilities
- Lead resilience and availability assessments across applications, infrastructure, data, network, cloud, and third-party services to identify vulnerabilities, single points of failure, recovery gaps, and control weaknesses.
- Design, assess, and improve high-availability and recovery patterns, including active-active architectures, clustering, load balancing, multi-zone or multi-region deployment, automated failover, and resilient dependency design.
- Define, validate, and monitor resilience objectives and measures, including Recovery Time Objective (RTO), Recovery Point Objective (RPO), Maximum Tolerable Downtime (MTD), service-level objectives and indicators, availability targets, and capacity.
- Lead business and application impact analysis, dependency mapping, critical service mapping, and recovery prioritization to align technology capabilities with business resilience requirements.
- Plan and oversee high-availability and failover exercises, recovery-from-backup tests, tabletop scenarios, extended-duration testing, and appropriate failure-injection or chaos-testing practices.
- Assess backup, restore, and cyber-recovery capabilities, including immutable or isolated backups, point-in-time recovery, clean-room recovery, and ransomware recovery scenarios.
- Embed resilience and availability requirements into technology architecture, service operations, capacity management, change management, configuration management, incident and problem management, and the software development lifecycle.
- Use monitoring and observability data to identify availability, performance, capacity, and recovery risks, and translate findings into prioritized remediation actions and measurable improvements.
- Coordinate remediation activities, track risks and actions to closure, and provide clear reporting on capability maturity, test outcomes, control effectiveness, and residual risk.
- Produce organized, traceable, and defensible evidence for internal audit, regulatory examinations, senior management, and board or risk committee reporting.
- Serve as a trusted resilience advisor and central point of coordination across engineering, application owners, infrastructure, cyber, business continuity, technology risk, third-party risk, and operational resilience teams.
- Monitor emerging technology risks, regulatory expectations, cyber threats, and industry practices, and recommend practical enhancements to resilience standards, procedures, controls, and testing methods.
Job Requirements
Resilience, Availability and Recovery
- Demonstrated knowledge of high-availability design patterns, failover strategies, fault tolerance, redundancy, and elimination of single points of failure.
- Experience establishing and assessing RTO, RPO, MTD, availability objectives, SLOs, and recovery or availability metrics.
- Experience with business or application impact analysis, critical service mapping, technology dependency mapping, and recovery sequencing.
- Hands-on experience developing DR plans and runbooks and coordinating technical recovery exercises, failover tests, tabletop exercises, and end-to-end recovery validation.
- Knowledge of enterprise backup and restore, immutable or isolated backup, point-in-time recovery, and cyber-recovery concepts.
- Knowledge of capacity management, performance monitoring, utilization forecasting, and reporting
Platforms, Engineering and Tooling
- Experience with cloud resilience capabilities in AWS, Microsoft Azure, and/or Google Cloud, including multi-region architecture, traffic management, native backup, and disaster recovery services.
- Understanding of on-premises and hybrid technology, including VMware, SAN/NAS storage, Windows, Linux, Active Directory, DNS, and network dependencies.
- Working knowledge of container and Kubernetes high-availability patterns in cloud-native environments.
- Knowledge of database resilience methods for platforms such as Oracle, Microsoft SQL Server, and PostgreSQL, including replication, clustering, backup, restore, and point-in-time recovery.
- Experience using observability and monitoring platforms such as Splunk, Dynatrace, Datadog, or equivalent tools to assess availability, capacity, performance, and recovery outcomes.
- Experience with automation and infrastructure-as-code tools such as Python, or PowerShell.
- Experience with ServiceNow capabilities, including CMDB, incident, problem, change, and related technology risk or control workflows, is strongly preferred.
- Proficiency with Microsoft Word, Excel, and PowerPoint for analysis, evidence management, executive reporting, and program documentation.
Risk, Control and Regulatory
- Strong understanding of technology risk, operational resilience, disaster recovery, business continuity, and control assessment in a regulated environment.
- Working knowledge of relevant frameworks and guidance, including FFIEC Business Continuity Management expectations, OSFI/OCC and Federal Reserve operational-resilience guidance, NIST Cybersecurity Framework, and ITIL practices.
- Experience mapping technology applications and dependencies to important or critical business services.
- Knowledge of third-party resilience, including concentration risk, critical technology service provider testing, contingency planning, and exit strategies.
- Experience preparing evidence and written responses for internal audit, regulators, risk committees, and senior executives.
- Ability to apply major incident, problem, change, and post-incident review practices to improve resilience and reduce recurring disruption.
Competencies
- Advanced stakeholder management and relationship-building skills across engineering, application, infrastructure, cyber, risk, audit, and business teams.
- Clear and concise written and verbal communication, including executive-ready updates, root-cause analyses, postmortems, board or risk materials, and regulator-ready documentation.
- Ability to influence without direct authority and drive adoption of standards, remediation commitments, and sustainable process improvements across a matrix organization.
- Calm, structured leadership during incidents, recovery events, testing exercises, and periods of heightened scrutiny.
- Strong analytical problem-solving, including root-cause analysis, dependency analysis, scenario analysis, risk assessment, and gap-to-control mapping.
- Excellent organization and evidence-management discipline, with the ability to manage multiple priorities, deadlines, and stakeholders without compromising quality.
- Sound judgment under ambiguity and the ability to design and assess severe but plausible scenarios rather than relying only on happy-path recovery assumptions.
- Collaborative mindset suited to global, matrixed, and follow-the-sun operating models.
- Continuous-improvement orientation focused on reducing outages, improving control effectiveness, increasing test fidelity, automating repeatable work, and closing findings sustainably.
Experience and Education
- Typically 8 to 12 years of relevant experience in technology operations, site reliability engineering, infrastructure, disaster recovery, business continuity, availability or capacity management, cyber recovery, or technology risk; experience in financial services or another regulated industry is strongly preferred.
- Bachelor’s degree in Computer Science, Engineering, Information Technology, or a related discipline, or equivalent practical experience.
- Experience leading complex resilience initiatives, technical assessments, remediation programs, or cross-functional testing activities in a large enterprise environment.
Preferred Qualifications
- ITIL 4, CBCI/MBCI, ISO 22301, CISSP, CISA, CRISC, cloud associate or professional certification, or relevant site reliability engineering training.
- Experience applying SRE practices to regulated workloads, including SLOs, error budgets, automation, and toil reduction.
- Experience with cyber-resilience and ransomware-recovery capabilities, including isolated recovery environments and clean-room restoration.
- Exposure to mainframe, payments, core-banking, or other highly critical enterprise platforms.
- Experience supporting regulatory examinations, internal audit reviews, or formal remediation programs related to technology resilience, availability, recovery, or business continuity.
- Experience integrating resilience requirements and control gates into Agile, DevOps, architecture review, change, and software delivery processes.
Who We Are:
TD is one of the world's leading global financial institutions and is the fifth largest bank in North America by branches/stores. Every day, we strive to make every interaction, product, and experience remarkably human and refreshingly simple for over 27 million households and businesses in Canada, the United States and around the world. More than 95,000 TD colleagues bring their skills, talent, and creativity to foster deeper relationships, ensure disciplined execution, and build a simpler, faster banking experience. TD is deeply committed to being a leader in client experience, that is why we believe that all colleagues, no matter where they work, are client facing. Together, we are reimagining what banking can be for our clients, colleagues and communities.
Our Total Rewards Package
Our Total Rewards package reflects the investments we make in our colleagues to help them and their families achieve their financial, physical, and mental well-being goals. Total Rewards at TD includes a base salary, variable compensation, and several other key plans such as health and well-being benefits, savings and retirement programs, paid time off, banking benefits and discounts, career development, and reward and recognition programs. Learn more
Additional Information:
We’re delighted that you’re considering building a career with TD. Through regular development conversations, training programs, and a competitive benefits plan, we’re committed to providing the support our colleagues need to thrive both at work and at home.
Please be advised that this job opportunity is subject to provincial regulation for employment purposes. It is imperative to acknowledge that each province or territory within the jurisdiction of Canada may have its own set of regulations, requirements.
Colleague Development
If you’re interested in a specific career path or are looking to build certain skills, we want to help you succeed. You’ll have regular career, development, and performance conversations with your manager, as well as access to an online learning platform and a variety of mentoring programs to help you unlock future opportunities.
If you’re passionate about helping clients and building deep, lasting relationships, TD offers diverse career paths where you can grow your expertise and make a meaningful impact.
We're committed to your success and foster a respectful workplace where diverse perspectives are valued, everyone has fair opportunities to grow, and you can unlock your full potential to achieve your career goals. Here at TD, we hire and develop the best.
Training & Onboarding
We will provide training and onboarding sessions to ensure that you’ve got everything you need to succeed in your new role.
Interview Process
We’ll reach out to candidates of interest to schedule an interview. We do our best to communicate outcomes to all applicants by email or phone call.
Accommodation
Your accessibility is important to us. Please let us know if you’d like accommodations (including accessible meeting rooms, captioning for virtual interviews, etc.) to help us remove barriers so that you can participate throughout the interview process.
We look forward to hearing from you!
Language Requirement (Quebec only):
Sans Objet