Pay Range: CAD 55-60/hr
We are seeking an experienced Observability Engineer to join our EnterpriseKubernetes Platform team at a leading financial services organization. Youllown the complete observability stack across 50+ production Kubernetes clusters|providing metrics| logging| tracing| and alerting capabilities that ensureexceptional reliability and performance for mission-critical applications.This role combines deep technical expertise in modern observability tools withemerging AI/ML capabilities to build intelligent monitoring solutions|predictive alerting| and self-healing infrastructure.========================================================================WHAT YOULL DO========================================================================OBSERVABILITY STACK OWNERSHIP-------------------------------------------------------------------------------- Design| deploy| and maintain enterprise-scale observability infrastructureincluding Prometheus| Grafana| Thanos| Loki| and modern collection agents Manage observability deployments using GitOps principles and infrastructureas code Implement long-term metrics storage solutions with cloud object storage Maintain and upgrade observability components across development| QA| UAT|production| and DR environments Configure distributed observability architecture spanning multiple datacenters and cloud providersMETRICS & MONITORING-------------------------------------------------------------------------------- Design and implement Prometheus monitoring strategies for Kubernetesinfrastructure and containerized applications Create ServiceMonitors| PodMonitors for automated metrics collection Develop rules for intelligent alerting with minimal false positives Configure multi-cluster metrics federation and aggregation Optimize metrics cardinality| storage e_iciency| and query performance Implement recording rules for pre-aggregated metrics and SLI calculationsDASHBOARDS & VISUALIZATION-------------------------------------------------------------------------------- Build comprehensive Grafana dashboards for infrastructure health| applicationperformance| and business metrics Create reusable dashboard templates and libraries for development teams Implement dashboard-as-code Configure multiple datasources Design executive dashboards with SLO/SLI tracking and business KPIs Implement role-based access control and multi-tenancy in GrafanaLOGGING INFRASTRUCTURE-------------------------------------------------------------------------------- Deploy and manage centralized logging solutions (Loki| ELK| Splunk| orsimilar) Configure log collection agents (Promtail| Fluentd| FluentBit| Vector| etc.) Design log retention policies balancing cost| compliance| and operationalneeds Create LogQL/Lucene queries and log-base