Adobe Site Reliability Engineer (SRE) – CDN & Web Performance
Location: Toronto, ON
Salary: $130,000 per year
Job Type: Full-time, Permanent
Job Summary
We are seeking a highly skilled and proactive Adobe Site Reliability Engineer (SRE) specializing in CDN and Web Performance to join our client's dynamic team.
Job Overview
This role focuses on CDN and edge reliability, website performance, monitoring, incident management, release engineering, and operational resilience. The successful candidate will help establish reliability practices for a vendor-managed platform supporting high-traffic public-facing digital experiences.
WHAT YOU WILL OWN
1. Application Support & Incident Management
- Own end-to-end monitoring of the TCS-controlled environment, including CDN and edge configuration, DNS, certificates, caching, cache invalidation, and third-party integrations such as search, consent management, analytics, personalization, and AI services.
- Build and operate synthetic monitoring outside the corporate network, covering individual templates and languages to identify CDN, DNS, and certificate failures.
- Run smoke tests for dependent interfaces after every change and maintain automated testing suites.
- Participate in the shared on-call rotation as the platform subject-matter expert and lead incident management for customer-facing events.
- Manage vendor escalation procedures with Adobe, including incident severity mapping, escalation contacts, evidence collection, and vendor accountability.
- Troubleshoot platform-specific incidents, including published content that is not visible, cache invalidation failures, and content authoring source outages. Improve diagnostic procedures so the service desk can identify and resolve common issues.
2. Change and Release Reliability
- Design and manage change control for a Git-based production deployment model, ensuring every production change has an approved record without unnecessarily delaying delivery.
- Own the release pipeline as a production control, including branch protection, required checks, linting, performance testing, secret scanning, and supporting evidence.
- Own rollback procedures, including code reverts, content republishing, and cache purging. Rehearse recovery procedures, measure recovery times, and establish clear decision-making authority.
- Establish controls for routine content publishing, ensuring production changes made by content authors have approval evidence, attribution, and retention records.
- Represent the platform at change advisory board meetings and manage release freeze calendars and release notes.
3. Business Continuity and Resilience
- Own TCS's recovery responsibilities. While Adobe operates the delivery infrastructure, TCS must maintain recovery capabilities for content sources, Git repositories, and CDN configurations.
- Manage CDN and edge configuration as code so that lost or corrupted configurations can be restored through controlled redeployment.
- Define Recovery Time Objectives (RTOs) and Recovery Point Objectives (RPOs) with the business, document disaster recovery plans, and execute recovery testing.
- Test realistic failure scenarios, including certificate expiry, cache invalidation failures, WAF misconfiguration, content source unavailability, and repository compromise.
- Maintain operational resilience evidence for third-party technology arrangements, including platform exit and portability plans.
4. Reliability & Performance Engineering
- Establish and maintain service level objectives for availability and page performance. Define Core Web Vitals thresholds for each template, manage error budgets, and report on performance.
- Build an observability practice using Real User Monitoring (RUM), CDN access logs integrated with enterprise SIEM, and external synthetic monitoring.
- Develop monitoring and troubleshooting procedures that account for the absence of traditional origin server logs.
- Manage third-party scripts and tag governance as reliability controls. Measure individual tags' performance impact, establish approval processes, and enforce performance budgets.
- Manage CDN data transfer, asset storage and processing, media delivery, and any TCS-hosted APIs supporting the website.
- Produce reliability and performance reports for business, risk, and technology leadership.
5. Compliance and Control Evidence
- Establish and document operational controls for a platform that TCS does not operate directly.
- Manage SIEM log ingestion and retention, access reviews across repositories, Adobe Admin Console, content sources, and CDN platforms, and maintain audit evidence.
- Support privacy, operational risk, control assessments, and third-party risk management with appropriate operational evidence.
- Maintain accurate configuration management records, support models, and operational assignment groups as the platform evolves.
WHAT WILL YOU DO?
This is a build-and-run role. Approximately half of the first year will focus on establishing a reliability practice.
Key deliverables include:
- Building the observability stack, including Real User Monitoring, Core Web Vitals dashboards and alerts, CDN log ingestion, and external synthetic monitoring.
- Implementing CDN and edge configuration as code with tested recovery procedures.
- Developing operational runbooks, incident response playbooks, operational level agreements, and vendor escalation matrices.
- Establishing a change management model that aligns Git-based deployments and continuous content publishing with TCS change controls.
- Conducting the first disaster recovery exercise and implementing a rehearsed, measured rollback process.
- Defining service level objectives with the business and establishing performance reporting.
WHAT DO YOU NEED TO SUCCEED?
Must-Have Qualifications
- Substantial hands-on experience operating high-traffic public websites behind enterprise content delivery networks.
- Strong expertise in CDN configuration, including origin and cache behaviour, cache invalidation, edge logic, TLS, and DNS. Experience with Akamai or Cloudflare is an advantage.
- Practical Web Application Firewall (WAF) experience, including tuning false positives against production-like traffic and implementing bot management without blocking legitimate search engine crawlers.
- Experience defining service level objectives and error budgets and building observability using logs, real-user telemetry, and synthetic monitoring.
- Strong web performance engineering knowledge, including Core Web Vitals, page loading and rendering behaviour, browser performance analysis, and identifying scripts responsible for performance regressions.
- Familiarity with front-end technologies, particularly JavaScript and CSS, and the ability to investigate reliability and performance issues in browser-delivered applications.
- Experience with Git-based release engineering, CI/CD pipelines, and infrastructure or configuration as code.
- Experience leading incidents affecting customer-facing services and producing technical evidence during incident response.
- Experience working in regulated environments, with an understanding of change control, audit evidence, access management, and third-party risk requirements.
Nice-to-Have Qualifications
- Experience operating vendor-managed or SaaS-delivered platforms, including monitoring, vendor escalation, and supplier accountability.
- Experience with Adobe Experience Manager (AEM), particularly Edge Delivery Services and Assets as a Cloud Service.
- Experience in financial services or another regulated industry.
- Experience supporting digital experiences in both English and French.
- Knowledge of website accessibility and Search Engine Optimization (SEO).
- Scripting or automation experience with Python, Java, or JavaScript.