About the Role
We are seeking a Senior Site Reliability Engineer (SRE) to help build and operate a modern, public-facing enterprise web platform delivered through Adobe Experience Manager (AEM) Edge Delivery Services (EDS).
This is a unique reliability engineering opportunity focused on CDN and edge reliability, web performance, observability, release engineering, incident management, resilience, and third-party integrations rather than traditional server administration.
You will help establish the reliability practice from the ground up and work closely with platform vendors, cybersecurity, network/CDN, privacy, production support, engineering, and enterprise technology teams in a highly regulated environment.
In this role, web performance is treated as a core product requirement. You will define measurable performance objectives, monitor real-user experience, manage performance budgets, and drive improvements across the entire web delivery stack.
Key Responsibilities :
Application Support & Incident Management
- Own the operational health and monitoring of the public-facing web platform across CDN, edge configuration, DNS, TLS certificates, caching, cache invalidation, WAF, bot management, and third-party integrations.
- Monitor critical customer-facing journeys and identify reliability or performance issues before they significantly impact users.
- Build and maintain external synthetic monitoring across key templates, languages, regions, and customer journeys.
- Develop automated smoke tests and dependency checks to validate critical interfaces following production changes.
- Participate in an on-call rotation and serve as a senior technical escalation point for customer-facing incidents.
- Lead incident response, troubleshooting, stakeholder communication, root-cause analysis, and post-incident reviews.
- Establish and maintain vendor escalation procedures, including severity mapping, escalation contacts, evidence collection, SLA/OLA tracking, and incident follow-up.
- Create diagnostics and operational runbooks for edge-specific incidents such as content publishing failures, cache invalidation issues, DNS/TLS problems, WAF issues, and content-source outages.
Change & Release Reliability
- Design and operate change-management processes for a Git-based production environment.
- Establish reliable controls around Git-based deployments while maintaining enterprise change-management and audit requirements.
- Own CI/CD quality and security controls, including branch protection, required checks, linting, automated testing, performance validation, and secret scanning.
- Establish repeatable rollback, revert, republish, and cache-purge procedures.
- Rehearse recovery procedures and measure recovery time to ensure rollback processes are operationally proven.
- Support high-volume content publishing with appropriate approval, attribution, traceability, and audit evidence.
- Represent the platform through change governance and release-management processes.
- Maintain release documentation, freeze calendars, deployment procedures, and operational readiness requirements.
Business Continuity & Resilience
- Define and test recovery strategies for the content source, Git repository, CDN configuration, and edge platform configuration.
- Manage CDN and edge configuration using Infrastructure as Code (IaC) / Configuration as Code principles wherever practical.
- Define Recovery Time Objectives (RTO) and Recovery Point Objectives (RPO) in partnership with business and technology stakeholders.
- Plan and execute disaster recovery, resilience, and operational continuity exercises.
- Test realistic failure scenarios, including:
- Certificate expiration
- DNS issues
- Cache invalidation failures
- WAF misconfiguration
- Content-source unavailability
- Repository issues or compromise
- CDN/edge configuration failures
- Maintain operational resilience documentation and evidence required for a regulated enterprise technology environment.
- Identify platform dependencies and maintain appropriate recovery and exit/portability strategies for critical third-party services.
Reliability & Web Performance Engineering
- Define and maintain Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets for availability, reliability, and web performance.
- Establish Core Web Vitals targets and performance thresholds across key page templates and customer journeys.
- Build an observability practice using Real User Monitoring (RUM), CDN/edge logs, enterprise SIEM, metrics, and external synthetic monitoring.
- Analyze web performance regressions using browser/network telemetry and identify the specific source of degradation.
- Review browser waterfalls, rendering behavior, network requests, JavaScript execution, CSS, third-party scripts, APIs, and other dependencies.
- Establish governance around third-party scripts and tags, including performance measurement, approval processes, and performance budgets.
- Monitor CDN usage, asset storage, media delivery, egress, and other platform-related capacity and cost considerations.
- Identify trends and recurring reliability issues and drive preventative engineering improvements.
- Produce reliability, availability, and web-performance reporting for technology, risk, and business leadership.
Compliance & Operational Controls
- Establish operational controls and evidence for a vendor-operated / SaaS-based technology platform.
- Support SIEM log ingestion, retention, access reviews, repository controls, configuration management, and audit evidence.
- Partner with cybersecurity, privacy, compliance, risk, and third-party risk teams.
- Maintain accurate configuration management records, support models, ownership, assignment groups, and operational documentation.
- Ensure reliability and operational practices align with enterprise security, regulatory, and third-party risk requirements.
- Support technology control assessments and provide operational evidence for audits and regulatory reviews.
What You'll Build :
This is a build-and-run SRE role. During the initial phase, you will help establish and mature the reliability engineering practice, including:
- Real User Monitoring (RUM) and Core Web Vitals dashboards
- CDN and edge log ingestion and alerting
- External synthetic monitoring
- CDN and edge configuration as code
- Reliability and operational runbooks
- Incident response and escalation playbooks
- Vendor escalation and operational support processes
- Git-based change and release controls
- Automated smoke and dependency testing
- Disaster recovery and resilience procedures
- Rehearsed rollback and recovery processes
- SLOs, SLIs, error budgets, and reliability reporting
- Operational readiness and compliance evidence
Required Qualifications :
- Strong hands-on experience operating high-traffic, public-facing websites behind an enterprise CDN.
- Deep expertise in CDN configuration and edge delivery, including:
- Origin behavior
- Caching strategies
- Cache invalidation
- Edge logic
- TLS/SSL
- DNS
- Routing and traffic behavior
- Strong experience with Akamai and/or Cloudflare is highly desirable.
- Practical experience with Web Application Firewalls (WAF), including rule management, tuning, false-positive analysis, and production enforcement.
- Experience with bot management and protecting web applications while allowing legitimate search-engine and business-critical crawlers.
- Strong observability experience using logs, metrics, Real User Monitoring, synthetic monitoring, SLIs, SLOs, and error budgets.
- Strong understanding of web performance and Core Web Vitals, including browser rendering, network behavior, page-load performance, and waterfall analysis.
- Comfortable reviewing and troubleshooting JavaScript and CSS in modern web applications.
- Strong experience with Git-based development and release engineering.
- Hands-on experience with CI/CD pipelines and production deployment controls.
- Experience with Infrastructure as Code / Configuration as Code.
- Proven experience supporting or leading customer-facing incident management and incident response.
- Strong understanding of root-cause analysis, post-incident reviews, operational readiness, and continuous reliability improvement.
- Experience working within enterprise change management, audit, access-control, compliance, and third-party risk processes.
- Ability to work effectively across engineering, security, infrastructure, vendors, production support, and business stakeholders.
Preferred Qualifications :
- Experience operating vendor-managed or SaaS platforms, where reliability requires monitoring, escalation, governance, and vendor accountability.
- Experience with Adobe Experience Manager (AEM), particularly AEM Edge Delivery Services or Assets as a Cloud Service.
- Experience in financial services, banking, healthcare, or another highly regulated industry.
- Experience supporting multilingual or bilingual public-facing websites.
- Knowledge of web accessibility and SEO from a reliability, performance, compliance, and discoverability perspective.
- Automation experience using Python, Java, or JavaScript.
- Experience building automated operational tooling, diagnostics, synthetic tests, or runbooks.
- Experience with enterprise SIEM and log-management platforms.
Experience defining and reporting reliability and performance metrics to senior technology stakeholders.
What This Role Is Not :
This role is specifically focused on edge reliability, CDN, web performance, observability, release engineering, incident management, and modern platform reliability.
It does not primarily involve:
- Operating-system patching
- Kubernetes or cluster administration
- Container orchestration
- Traditional application-server administration
- Traditional database administration
- Server-based infrastructure operations
- Traditional compute capacity planning
The underlying delivery infrastructure is largely operated by the platform provider. Your focus will instead be on the reliability surface owned by the enterprise, including CDN/edge configuration, content delivery, integrations, release processes, observability, performance, resilience, and operational controls.
Why Join This Team?
This is an opportunity to help define reliability engineering for a modern, serverless/edge-based enterprise web platform rather than simply operating an existing infrastructure environment.
You will have the opportunity to:
- Build a reliability practice from the ground up
- Work with modern CDN and edge technologies
- Own enterprise-grade web performance and Core Web Vitals
- Establish SRE principles including SLOs, SLIs, and error budgets
- Build observability around RUM, synthetics, and edge telemetry
- Lead customer-facing incident response
- Design resilience and disaster-recovery practices
- Work directly with technology vendors and enterprise stakeholders
- Influence how reliability, performance, and operational excellence are implemented across a high-visibility public platform
If you are an experienced SRE / Reliability Engineer with deep CDN, edge, web-performance, observability, and incident-management experience, we would like to hear from you.