Observability Engineer (KUBERNETES & CONTAINERS, GITOPS & AUTOMATION, QUERY LANGUAGES, CLOUD & INFRASTRUCTURE)
Toronto, On
Hybrid - 4 days
ABOUT THE ROLE
We are seeking an experienced Observability Engineer to join our Enterprise Kubernetes Platform team at a leading financial services organization. You'll own the complete observability stack across 50+ production Kubernetes clusters, providing metrics, logging, tracing, and alerting capabilities that ensure exceptional reliability and performance for mission-critical applications.
This role combines deep technical expertise in modern observability tools with emerging AI/ML capabilities to build intelligent monitoring solutions, predictive alerting, and self-healing infrastructure.
========================================================================
WHAT YOU'LL DO
========================================================================
OBSERVABILITY STACK OWNERSHIP
--------------------------------------------------------------------------------
- Design, deploy, and maintain enterprise-scale observability infrastructure including Prometheus, Grafana, Thanos, Loki, and modern collection agents • Manage observability deployments using GitOps principles and infrastructure as code • Implement long-term metrics storage solutions with cloud object storage • Maintain and upgrade observability components across development, QA, UAT, production, and DR environments • Configure distributed observability architecture spanning multiple data centers and cloud providers METRICS & MONITORING
--------------------------------------------------------------------------------
- Design and implement Prometheus monitoring strategies for Kubernetes infrastructure and containerized applications • Create ServiceMonitors, PodMonitors for automated metrics collection • Develop rules for intelligent alerting with minimal false positives • Configure multi-cluster metrics federation and aggregation • Optimize metrics cardinality, storage e_iciency, and query performance • Implement recording rules for pre-aggregated metrics and SLI calculations DASHBOARDS & VISUALIZATION
--------------------------------------------------------------------------------
- Build comprehensive Grafana dashboards for infrastructure health, application performance, and business metrics • Create reusable dashboard templates and libraries for development teams • Implement dashboard-as-code • Configure multiple datasources • Design executive dashboards with SLO/SLI tracking and business KPIs • Implement role-based access control and multi-tenancy in Grafana LOGGING INFRASTRUCTURE
--------------------------------------------------------------------------------
- Deploy and manage centralized logging solutions (Loki, ELK, Splunk, or
similar)
- Configure log collection agents (Promtail, Fluentd, FluentBit, Vector, etc.) • Design log retention policies balancing cost, compliance, and operational needs • Create LogQL/Lucene queries and log-based alerts • Implement log correlation with metrics and traces for unified troubleshooting • Build log aggregation pipelines with parsing, filtering, and enrichment ALERTING & INCIDENT MANAGEMENT
--------------------------------------------------------------------------------
- Configure intelligent alerting with Alertmanager or equivalent platforms • Design alert rules with appropriate severity levels, thresholds, and SLOs • Implement alert routing to Slack, PagerDuty, ServiceNow, email, and webhooks • Create automated runbooks and remediation workflows • Develop alert inhibition, silencing, and grouping strategies • Tune alerting to achieve signal-to-noise ratio improvements • Integrate with incident management and on-call rotation systems DISTRIBUTED TRACING
--------------------------------------------------------------------------------
- Implement distributed tracing solutions using OpenTelemetry, Jaeger, or Tempo • Instrument applications for trace collection and correlation • Configure trace sampling strategies for cost and performance optimization • Build trace-based dashboards for latency analysis and dependency mapping • Integrate tracing with metrics and logs for comprehensive observability AI/ML FOR OBSERVABILITY
--------------------------------------------------------------------------------
- Implement AI-powered anomaly detection for metrics and logs • Build predictive alerting using machine learning models to forecast issues before they occur • Develop intelligent alert correlation and root cause analysis systems • Integrate LLM-based tools for log analysis and troubleshooting assistance • Implement AIOps capabilities for automated incident triage and resolution • Use AI to optimize alert thresholds and reduce false positives • Build natural language query interfaces for observability data • Implement Model Context Protocol (MCP) for AI agent integration with observability platforms AUTOMATION & PLATFORM INTEGRATION
--------------------------------------------------------------------------------
- Automate observability deployment using GitOps workflows (FluxCD, ArgoCD) • Integrate with CI/CD pipelines for automated testing and validation • Build self-service portals for teams to create dashboards and alerts • Develop APIs and CLIs for observability automation • Integrate with secrets management solutions (Vault, AWS Secrets Manager) • Configure LDAP/AD/SSO authentication for observability platforms • Automate compliance reporting and audit logging ENABLEMENT & COLLABORATION
--------------------------------------------------------------------------------
- Onboard application teams to observability platforms • Provide guidance on instrumentation best practices • Create documentation, training materials, and self-service guides • Conduct workshops • Support development teams during incidents with observability insights • Collaborate with SRE, DevOps, and platform engineering teams PERFORMANCE & COST OPTIMIZATION
--------------------------------------------------------------------------------
- Monitor and optimize observability stack resource consumption • Implement autoscaling for stateless observability components • Tune data retention, compaction, and downsampling strategies • Conduct capacity planning for metrics and log storage growth • Optimize query performance and dashboard response times • Implement cost allocation and chargeback for multi-tenant environments =======================================================================
REQUIRED QUALIFICATIONS
=======================================================================
EXPERIENCE
--------------------------------------------------------------------------------
- 4-6 years of experience in observability, monitoring, SRE, or platform engineering roles • 3+ years hands-on production experience with Prometheus and Grafana • 2+ years working with Kubernetes and containerized environments • Strong expertise with PromQL for metrics querying and alerting • Experience deploying and managing observability stacks at scale (1000+ nodes) • Proven track record of reducing MTTR through e_ective observability CORE TECHNICAL SKILLS
--------------------------------------------------------------------------------
OBSERVABILITY PLATFORMS:
- Prometheus (including Prometheus Operator) • Grafana (dashboards, alerting, plugins) • Thanos, Cortex, or Mimir for long-term storage • Alertmanager or equivalent alerting platforms
LOGGING:
- Loki, Elasticsearch/ELK Stack, Splunk, or CloudWatch Logs • Log collection agents (Promtail, Fluentd, FluentBit, Vector) • LogQL, Lucene, or equivalent query languages
TRACING:
- OpenTelemetry (OTEL Collector, instrumentation) • Jaeger, Zipkin, Tempo, or AWS X-Ray • Trace sampling and correlation strategies KUBERNETES & CONTAINERS:
- Kubernetes architecture and operations (1.24+) • Custom Resource Definitions (CRDs) and Operators • ServiceMonitor, PodMonitor, PrometheusRule resources • Kubernetes metrics (kube-state-metrics, node-exporter, cAdvisor) GITOPS & AUTOMATION:
- GitOps workflows (FluxCD, ArgoCD, or similar) • Infrastructure as Code (Kustomize, Helm, Terraform) • CI/CD platforms (GitHub Actions, GitLab CI, Jenkins) • Scripting (Python, Bash, Go) QUERY LANGUAGES:
- PromQL (expert level required)
- Basic SQL for data analysis
CLOUD & INFRASTRUCTURE:
- AWS, Azure, or GCP cloud platforms
- Object storage (S3, GCS, Azure Blob)
- Multi-cloud and hybrid architectures
- vSphere or on-premises virtualization (nice to have) AI/ML & EMERGING TECHNOLOGIES
--------------------------------------------------------------------------------
- Experience with AI/ML frameworks for observability (Prophet, TensorFlow,
PyTorch)
- Anomaly detection algorithms and time-series forecasting • LLM integration for log analysis and troubleshooting (GPT, Claude, etc.) • Model Context Protocol (MCP) for AI agent integration • AIOps platforms (Moogsoft, BigPanda, Datadog Watchdog, etc.) • Natural language processing for log parsing and analysis • Familiarity with vector databases for semantic search (Pinecone, Weaviate) • Experience with AI-powered root cause analysis tools • Knowledge of prompt engineering for observability use cases PREFERRED QUALIFICATIONS
--------------------------------------------------------------------------------
- Kubernetes certifications (CKA, CKAD, or CKS) • Grafana certification or equivalent training • Knowledge of service mesh observability ( • Experience with SRE practices, SLIs, SLOs, and error budgets • Background in financial services or regulated industries • Familiarity with compliance requirements (SOX, PCI-DSS, etc.) • Contributions to open-source observability projects • Experience with APM tools