We are looking for an experienced Observability Engineer to own enterprise observability across
50+ production Kubernetes clusters in a financial services environment.
- Strong hands-on experience with Kubernetes, Prometheus, Grafana, Thanos, Loki and modern observability/collection agents.
- Build and manage metrics, logging, tracing, alerting, ServiceMonitors/PodMonitors, recording rules, and SLI/SLO monitoring.
- Experience with GitOps, Infrastructure as Code, multi-cluster observability, and cloud object storage.
- Develop Grafana dashboards, dashboard-as-code, RBAC/multi-tenancy, and performance/business KPI monitoring.
- Strong centralized logging experience with Loki, ELK, Splunk, Fluent Bit/Fluentd/Promtail/Vector and LogQL/Lucene.
- Exposure to AI/ML-driven monitoring, predictive alerting, anomaly detection, and self-healing infrastructure is highly preferred.