Seeking a Senior SRE to drive
system reliability, observability, automation and AI-powered operations across enterprise applications.
Key Responsibilities
- Lead SRE practices covering monitoring, alerting, logging, self-healing and reliability testing.
- Implement observability, automation and AIOps to improve incident detection, RCA and reduce MTTR.
- Use Generative AI / AI-powered tools for troubleshooting, runbook automation, knowledge management and production support.
- Support cloud-native applications, Kubernetes/OpenShift and application deployments.
- Lead Incident & Problem Management, production troubleshooting and RCA.
- Partner with development teams to ensure releases meet reliability and performance standards.
- Automate SRE processes and improve operational efficiency.