Senior Site Reliability Engineer – Adobe Experience Manager (AEM), CDN and Edge Delivery
Location: Toronto, Ontario (Hybrid)
Term: Full Time
ABOUT THE ROLE
We are seeking a hands-on Senior Site Reliability Engineer (SRE) with experience supporting Adobe Experience Manager (AEM), particularly Edge Delivery Services and Assets as a Cloud Service, alongside enterprise Content Delivery Network (CDN) platforms.
The role focuses on the reliability, performance, monitoring, and production support of high-traffic public websites powered by Adobe Experience Manager. Key technical areas include CDN and Edge configuration, Web Application Firewall (WAF), bot management, Transport Layer Security (TLS), Domain Name System (DNS), cache invalidation, Continuous Integration and Continuous Delivery (CI/CD), observability, web performance, and production incident management.
The successful candidate will work closely with platform vendors, security, networking, production support, and engineering teams to maintain website availability, optimize performance, improve release reliability, and strengthen operational resilience.
KEY RESPONSIBILITIES
Production Support and Incident Management
- Monitor CDN and Edge configurations, DNS, certificates, caching, cache invalidation, and third-party integrations.
- Build external synthetic monitoring across website templates, languages, and user journeys.
- Implement automated smoke tests to validate production changes.
- Participate in on-call support and technical escalations.
- Lead customer-impacting incident response and coordinate vendor escalations through resolution.
- Troubleshoot content publishing failures, cache invalidation issues, content-source outages, and website availability problems.
- Maintain incident documentation, operational runbooks, and troubleshooting procedures.
Change and Release Reliability
- Manage Git-based production deployments and release processes for website and Edge configurations.
- Implement CI/CD controls, including branch protection, automated checks, code linting, performance testing, and secret scanning.
- Validate rollback procedures, including code reverts, content republishing, and cache purging.
- Improve the reliability and automation of frequent content publishing workflows.
- Coordinate release planning, deployment schedules, and release documentation.
- Minimize production risks through repeatable, controlled, and well-tested deployment practices.
Business Continuity and Operational Resilience
- Develop recovery procedures for content sources, Git repositories, and CDN/Edge configurations.
- Implement configuration as code wherever practical.
- Define Recovery Time Objectives (RTO) and Recovery Point Objectives (RPO) with stakeholders.
- Conduct disaster recovery and restoration exercises.
- Test recovery scenarios involving certificate expiry, cache invalidation failures, Web Application Firewall configuration issues, content-source outages, and repository problems.
- Maintain recovery documentation and ensure operational procedures are tested and current.
Reliability and Web Performance Engineering
- Define and monitor Service Level Objectives (SLOs) for website availability and performance.
- Monitor Core Web Vitals and investigate performance degradation.
- Implement observability using Real User Monitoring (RUM), CDN access logs, external synthetic monitoring, and centralized logging.
- Analyze the performance impact of third-party scripts, tags, and external integrations.
- Establish performance budgets and investigate performance regressions.
- Monitor CDN data transfer and egress, asset storage, media delivery, and platform costs.
- Prepare reliability, availability, and performance reports to identify improvement opportunities.
Documentation and Continuous Improvement
- Maintain operational runbooks, incident response procedures, support documentation, and escalation paths.
- Improve monitoring, alerting, incident response, and operational workflows.
- Maintain platform configuration records and clearly document support ownership.
- Collaborate with engineering and vendor teams to identify recurring issues and reduce incident frequency.
- Automate repetitive operational tasks and troubleshooting activities.
- Drive continuous improvements in platform reliability, production support, and release processes.
WHAT YOU WILL BUILD AND IMPROVE
- Real User Monitoring dashboards and Core Web Vitals reporting.
- CDN log ingestion and external synthetic monitoring capabilities.
- CDN and Edge configuration as code.
- Production support runbooks and incident response playbooks.
- Structured vendor escalation and production support processes.
- Git-based release management and deployment workflows.
- Disaster recovery and restoration procedures.
- Tested and measurable rollback processes.
- Service Level Objectives and reliability reporting.
MUST-HAVE SKILLS AND EXPERIENCE
- Strong hands-on experience supporting high-traffic public websites using enterprise CDN platforms.
- Experience with Adobe Experience Manager (AEM), ideally Edge Delivery Services and Assets as a Cloud Service.
- Experience configuring CDN behavior, origin settings, caching, cache invalidation, Edge logic, TLS, and DNS.
- Practical experience configuring and troubleshooting Web Application Firewalls.
- Experience tuning WAF rules, investigating false positives, and implementing bot protection.
- Knowledge of observability, Service Level Objectives, error budgets, log monitoring, Real User Monitoring, and synthetic monitoring.
- Strong understanding of web performance, Core Web Vitals, page rendering, browser behavior, and network waterfall analysis.
- Ability to understand and troubleshoot JavaScript and Cascading Style Sheets (CSS) related to website behavior and performance.
- Experience with Git-based release engineering and CI/CD pipelines.
- Familiarity with Infrastructure as Code and configuration management.
- Strong production incident troubleshooting, coordination, and resolution skills.
- Excellent analytical, problem-solving, communication, and documentation skills.
- Ability to work effectively with engineering teams, vendors, and production support stakeholders.
PREFERRED TECHNICAL EXPERIENCE
- Akamai and/or Cloudflare CDN platforms.
- Adobe Experience Manager Edge Delivery Services.
- Adobe Experience Manager Assets as a Cloud Service.
- Enterprise website delivery and vendor-managed or Software as a Service (SaaS) platforms.
- Python, Java, or JavaScript scripting for operational automation.
- Automation of operational procedures and development of reusable troubleshooting tools.
- Web accessibility and Search Engine Optimization (SEO).
- Experience supporting bilingual English/French websites.
- Experience in financial services or large enterprise environments.
ROLE CLARIFICATION
This is a CDN/Edge-focused Site Reliability Engineering role with an emphasis on Adobe Experience Manager, website delivery, web performance, and production reliability.
The primary focus is CDN and Edge configuration, Web Application Firewall, performance engineering, observability, CI/CD, production incident response, and operational resilience.
This is not primarily a traditional infrastructure administration role. The position does not focus on operating system patching, container orchestration, cluster management, application server administration, database administration, or traditional compute capacity planning.
IDEAL CANDIDATE
The ideal candidate is a hands-on Site Reliability Engineer with experience supporting Adobe Experience Manager-powered websites and strong expertise in enterprise CDN and Edge delivery.
They understand website performance across network, content delivery, and browser layers and can troubleshoot complex production issues, improve monitoring, automate operational workflows, and coordinate incident resolution across engineering and vendor teams.
Candidates with practical experience in Adobe Experience Manager, CDN configuration, Web Application Firewall, web performance, observability, CI/CD, and production incident management are encouraged to apply.