We are looking for an experienced Observability Engineer to design, implement, and maintain enterprise observability solutions across cloud, application, infrastructure, and distributed environments. The ideal candidate will have strong hands-on experience with monitoring, logging, metrics, distributed tracing, alerting, and observability platforms.
Key Responsibilities
- Design and implement end-to-end observability solutions for applications, infrastructure, APIs, and cloud environments.
- Develop and maintain dashboards, alerts, metrics, logs, and distributed tracing solutions.
- Work with Grafana, Prometheus, OpenTelemetry, Dynatrace, and OpenSearch or similar observability platforms.
- Implement and manage telemetry pipelines for metrics, logs, and traces.
- Configure and optimize Grafana dashboards, Prometheus monitoring, and alerting rules.
- Implement application instrumentation using OpenTelemetry and related frameworks.
- Monitor Kubernetes clusters, containers, microservices, APIs, and cloud infrastructure.
- Troubleshoot production issues and perform detailed Root Cause Analysis (RCA).
- Define and monitor SLIs, SLOs, SLAs, availability, latency, and MTTR.
- Collaborate with DevOps, SRE, development, infrastructure, and application teams to improve system reliability.
- Automate monitoring, alerting, deployment, and operational processes using scripting and Infrastructure as Code.
- Support incident management, production deployments, post-incident reviews, and preventive actions.
- Identify performance bottlenecks and recommend improvements to application and infrastructure architecture.
- Create and maintain technical documentation, runbooks, monitoring standards, and operational procedures.
Required Technical Skills
- 7+ years of experience in Observability, SRE, DevOps, Platform Engineering, or related roles.
- Strong experience with Grafana and Prometheus.
- Hands-on experience with OpenTelemetry and telemetry pipelines.
- Strong understanding of Metrics, Logs, and Distributed Tracing.
- Experience with Kubernetes and Docker.
- Strong Linux administration and troubleshooting skills.
- Experience with AWS, Azure, or GCP.
- Knowledge of CI/CD tools such as Jenkins, GitLab CI/CD, or GitHub Actions.
- Experience with Infrastructure as Code such as Terraform or Ansible.
- Experience with monitoring/logging platforms such as Dynatrace, OpenSearch, ELK, Loki, or similar tools.
- Strong understanding of microservices, APIs, distributed systems, and cloud-native architectures.
- Good scripting/programming experience with Python, Bash, Go, or similar languages