Original job description
We are seeking an experienced Observability Engineer to join our Enterprise Kubernetes Platform team at a leading financial services organization.
You ll own the complete observability stack across 50+ production Kubernetes clusters, providing metrics, logging, tracing, and alerting capabilities that ensure exceptional reliability and performance for mission-critical applications.
This role combines deep technical expertise in modern observability tools with emerging AI/ML capabilities to build intelligent monitoring solutions, predictive alerting, and self-healing infrastructure.
What Youll Do
Design, deploy, and maintain enterprise-scale observability infrastructure including Prometheus, Grafana, Thanos, Loki, and modern collection agents
Manage observability deployments using GitOps principles and infrastructures code
Implement long-term metrics storage solutions with cloud object storage
Maintain and upgrade observability components across development, QA, UAT, production, and DR environments
Configure distributed observability architecture spanning multiple datacenters and cloud providers
METRICS & MONITORING
Design and implement Prometheus monitoring strategies for Kubernetes infrastructure and containerized applications
Create Service Monitors, Pod Monitors for automated metrics collection
Develop rules for intelligent alerting with minimal false positives
Configure multi-cluster metrics federation and aggregation
Optimize metrics cardinality, storage deficiency, and query performance.
More about this job
Responsibilities
Design, deploy, and maintain an enterprise observability stack across more than 50 production Kubernetes clusters, covering metrics, logging, tracing, and alerting. Manage multi-environment and multi-cloud observability infrastructure, including long-term metrics storage, monitoring strategies, intelligent alerting, and performance optimization.
Requirements
The role calls for mid-to-senior-level experience and deep technical expertise in modern observability tools and Kubernetes environments. The posting emphasizes experience with Prometheus-based monitoring, GitOps and infrastructure as code, multi-cluster metrics architecture, and optimizing alerting, storage, and query performance.
Skills
- Observability
- Kubernetes
- Prometheus
- Grafana
- Thanos
- Loki
- Metrics Monitoring
- Logging
- Distributed Tracing
- Alerting
- GitOps
- Infrastructure as Code
- Cloud Object Storage
- Metrics Federation
- Metrics Cardinality Optimization
- Multi-Cluster Architecture
Visa sponsorship
Not detected in the job text
Categories
- Technology
- Software
- Engineering
- Data & Analytics
Keywords
- Enterprise Kubernetes Platform
- Kubernetes
- Observability
- Prometheus
- Grafana
- Thanos
- Loki
- Metrics
- Logging
- Distributed Tracing
- Alerting
- GitOps
- Infrastructure as Code
- Cloud Object Storage
- Service Monitors
- Pod Monitors
- Metrics Federation
- Multi-Cluster Architecture
- Metrics Cardinality
- Query Performance
- Intelligent Monitoring
- Predictive Alerting
- Self-Healing Infrastructure
- Production Environments
- Disaster Recovery
- Financial Services