Hiring: Site Reliability Engineer – Adobe Experience Manager (AEM)
Location: Toronto, ON - 4 days/week onsite (non-negotiable)
Interview: In-person final round
About the Role:
We are looking for a Site Reliability Engineer (SRE) to support a net-new enterprise Content Management & Digital Asset Management platform delivering public web experiences through Adobe Experience Manager (AEM) – Edge Delivery Services (EDS).
This is a modern, vendor-operated architecture. Adobe operates the delivery tier, including the underlying platform and 24/7 platform support. The client owns the CDN, Edge, WAF, DNS, TLS, caching, integrations, release pipeline, observability, and operational reliability.
For this platform, web performance is the product. Availability and performance will be managed through measurable SLOs, error budgets, real-user telemetry, CDN logs, and external synthetic monitoring.
What You’ll Own:
🔹 Application Support & Incident Management:
- Own monitoring across CDN/Edge, DNS, certificates, caching and third-party integrations.
- Build external synthetic monitoring by template and language.
- Develop automated smoke tests for dependent interfaces.
- Participate in on-call and lead customer-facing incident management.
- Own escalation management with Adobe.
- Develop diagnostics for content publishing, invalidation and authoring-source failures.
🔹 Change & Release Reliability:
- Establish change management around Git-based deployments.
- Own CI/CD controls including branch protection, testing, performance and security gates.
- Design and rehearse end-to-end rollback procedures.
- Govern content publishing and maintain audit trails.
- Support CAB, release notes and freeze-calendar activities.
🔹 Business Continuity & Resilience:
- Own recovery of content source, Git repositories and CDN configuration.
- Manage CDN/Edge configuration as code.
- Define RTO/RPO and execute DR testing.
- Test failure scenarios including certificate expiry, WAF misconfiguration, invalidation failures and repository compromise.
- Maintain operational resilience and third-party technology risk evidence.
🔹 Reliability & Performance Engineering:
- Establish availability and performance SLOs and error budgets.
- Define and monitor Core Web Vitals by template.
- Build observability using RUM, CDN logs, SIEM and external synthetics.
- Govern third-party scripts and tags as reliability/performance controls.
- Manage CDN egress, asset storage, media delivery and client-hosted API capacity/costs.
- Deliver reliability and performance reporting to business and technology leadership.
🔹 Compliance & Control:
- Support SIEM log ingestion, retention and access controls.
- Maintain audit and control evidence.
- Support privacy, operational risk and third-party risk assessments.
- Maintain CMDB, support models and assignment groups.
What You Need:
Must Have:
- Strong hands-on experience operating high-traffic public websites behind enterprise CDNs.
- Deep expertise in CDN configuration, caching, invalidation, edge logic, TLS and DNS.
- Experience with Akamai and/or Cloudflare is highly desirable.
- Practical WAF tuning and bot management experience.
- Strong observability experience using logs, RUM and synthetic monitoring.
- Expertise in Core Web Vitals and web performance engineering.
- Comfortable understanding JavaScript, CSS and browser-side performance.
- Strong Git, CI/CD and infrastructure/configuration-as-code experience.
- Experience leading customer-facing incidents.
- Experience working in regulated environments, including change control, audit, access management and third-party risk.
Nice to Have:
- Adobe Experience Manager / AEM Edge Delivery Services experience.
- AEM Assets as a Cloud Service.
- Experience with vendor-operated/SaaS platforms.
- Financial services experience.
- English/French bilingual delivery experience.
- Accessibility and SEO awareness.
- Python, Java or JavaScript automation experience.
⭐ Ideal Candidate:
You’re someone who can build and run a reliability practice from the ground up - not just monitor existing infrastructure.
You understand that in a modern edge-delivered architecture, reliability means instrumenting the platform, controlling the edge, measuring real-user performance, automating operational processes, managing incidents and holding vendors accountable.
📩 Interested? Apply now or DM me with your updated resume.
kashish.kulbhaje@aptino.com / kashish.kulbhaje@aptino.biz
Contact No - 817 330 7081
#Hiring #SRE #SiteReliabilityEngineer #AEM #AdobeExperienceManager #EdgeDeliveryServices #CDN #Cloudflare #Akamai #DevOps #Observability #WebPerformance #CoreWebVitals #CICD