Senior Observability Engineer (KUBERNETES & CONTAINERS, GITOPS & AUTOMATION, QUERY LANGUAGES, CLOUD & INFRASTRUCTURE)
Toronto, ON - Hybrid (4 Days WFO)
12 months
ABOUT THE ROLE
We are seeking an experienced Observability Engineer to join our Enterprise
Kubernetes Platform team at a leading financial services organization. You'll
own the complete observability stack across 50+ production Kubernetes clusters,
providing metrics, logging, tracing, and alerting capabilities that ensure
exceptional reliability and performance for mission-critical applications.
This role combines deep technical expertise in modern observability tools with
emerging AI/ML capabilities to build intelligent monitoring solutions,
predictive alerting, and self-healing infrastructure.
========================================================================
WHAT YOU'LL DO
========================================================================
OBSERVABILITY STACK OWNERSHIP
--------------------------------------------------------------------------------
- Design, deploy, and maintain enterprise-scale observability infrastructure
including Prometheus, Grafana, Thanos, Loki, and modern collection agents
- Manage observability deployments using GitOps principles and infrastructure
as code
- Implement long-term metrics storage solutions with cloud object storage
- Maintain and upgrade observability components across development, QA, UAT,
production, and DR environments
- Configure distributed observability architecture spanning multiple data
centers and cloud providers
METRICS & MONITORING
--------------------------------------------------------------------------------
- Design and implement Prometheus monitoring strategies for Kubernetes
infrastructure and containerized applications
- Create ServiceMonitors, PodMonitors for automated metrics collection
- Develop rules for intelligent alerting with minimal false positives
- Configure multi-cluster metrics federation and aggregation
- Optimize metrics cardinality, storage e_iciency, and query performance
- Implement recording rules for pre-aggregated metrics and SLI calculations
DASHBOARDS & VISUALIZATION
--------------------------------------------------------------------------------
- Build comprehensive Grafana dashboards for infrastructure health, application
performance, and business metrics
- Create reusable dashboard templates and libraries for development teams
- Implement dashboard-as-code
- Configure multiple datasources
- Design executive dashboards with SLO/SLI tracking and business KPIs
- Implement role-based access control and multi-tenancy in Grafana
LOGGING INFRASTRUCTURE
--------------------------------------------------------------------------------
- Deploy and manage centralized logging solutions (Loki, ELK, Splunk, or
similar)
- Configure log collection agents (Promtail, Fluentd, FluentBit, Vector, etc.)
- Design log retention policies balancing cost, compliance, and operational
needs
- Create LogQL/Lucene queries and log-based alerts
- Implement log correlation with metrics and traces for unified troubleshooting
- Build log aggregation pipelines with parsing, filtering, and enrichment
ALERTING & INCIDENT MANAGEMENT
--------------------------------------------------------------------------------
- Configure intelligent alerting with Alertmanager or equivalent platforms
- Design alert rules with appropriate severity levels, thresholds, and SLOs
- Implement alert routing to Slack, PagerDuty, ServiceNow, email, and webhooks
- Create automated runbooks and remediation workflows
- Develop alert inhibition, silencing, and grouping strategies
- Tune alerting to achieve signal-to-noise ratio improvements
- Integrate with incident management and on-call rotation systems
DISTRIBUTED TRACING
--------------------------------------------------------------------------------
- Implement distributed tracing solutions using OpenTelemetry, Jaeger, or
Tempo
- Instrument applications for trace collection and correlation
- Configure trace sampling strategies for cost and performance optimization
- Build trace-based dashboards for latency analysis and dependency mapping
- Integrate tracing with metrics and logs for comprehensive observability
AI/ML FOR OBSERVABILITY
--------------------------------------------------------------------------------
- Implement AI-powered anomaly detection for metrics and logs
- Build predictive alerting using machine learning models to forecast issues
before they occur
- Develop intelligent alert correlation and root cause analysis systems
- Integrate LLM-based tools for log analysis and troubleshooting assistance
- Implement AIOps capabilities for automated incident triage and resolution
- Use AI to optimize alert thresholds and reduce false positives
- Build natural language query interfaces for observability data
- Implement Model Context Protocol (MCP) for AI agent integration with
observability platforms
AUTOMATION & PLATFORM INTEGRATION
--------------------------------------------------------------------------------
- Automate observability deployment using GitOps workflows (FluxCD, ArgoCD)
- Integrate with CI/CD pipelines for automated testing and validation
- Build self-service portals for teams to create dashboards and alerts
- Develop APIs and CLIs for observability automation
- Integrate with secrets management solutions (Vault, AWS Secrets Manager)
- Configure LDAP/AD/SSO authentication for observability platforms
- Automate compliance reporting and audit logging
ENABLEMENT & COLLABORATION
--------------------------------------------------------------------------------
- Onboard application teams to observability platforms
- Provide guidance on instrumentation best practices
- Create documentation, training materials, and self-service guides
- Support development teams during incidents with observability insights
- Collaborate with SRE, DevOps, and platform engineering teams
PERFORMANCE & COST OPTIMIZATION
--------------------------------------------------------------------------------
- Monitor and optimize observability stack resource consumption
- Implement autoscaling for stateless observability components
- Tune data retention, compaction, and downsampling strategies
- Conduct capacity planning for metrics and log storage growth
- Optimize query performance and dashboard response times
- Implement cost allocation and chargeback for multi-tenant environments
=======================================================================
REQUIRED QUALIFICATIONS
=======================================================================
EXPERIENCE
--------------------------------------------------------------------------------
- 4-6 years of experience in observability, monitoring, SRE, or platform
engineering roles
- 3+ years hands-on production experience with Prometheus and Grafana
- 2+ years working with Kubernetes and containerized environments
- Strong expertise with PromQL for metrics querying and alerting
- Experience deploying and managing observability stacks at scale (1000+ nodes)
- Proven track record of reducing MTTR through e_ective observability
CORE TECHNICAL SKILLS
--------------------------------------------------------------------------------
OBSERVABILITY PLATFORMS:
- Prometheus (including Prometheus Operator)
- Grafana (dashboards, alerting, plugins)
- Thanos, Cortex, or Mimir for long-term storage
- Alertmanager or equivalent alerting platforms
LOGGING:
- Loki, Elasticsearch/ELK Stack, Splunk, or CloudWatch Logs
- Log collection agents (Promtail, Fluentd, FluentBit, Vector)
- LogQL, Lucene, or equivalent query languages
TRACING:
- OpenTelemetry (OTEL Collector, instrumentation)
- Jaeger, Zipkin, Tempo, or AWS X-Ray
- Trace sampling and correlation strategies
KUBERNETES & CONTAINERS:
- Kubernetes architecture and operations (1.24+)
- Custom Resource Definitions (CRDs) and Operators
- ServiceMonitor, PodMonitor, PrometheusRule resources
- Kubernetes metrics (kube-state-metrics, node-exporter, cAdvisor)
GITOPS & AUTOMATION:
- GitOps workflows (FluxCD, ArgoCD, or similar)
- Infrastructure as Code (Kustomize, Helm, Terraform)
- CI/CD platforms (GitHub Actions, GitLab CI, Jenkins)
- Scripting (Python, Bash, Go)
QUERY LANGUAGES:
- PromQL (expert level required)
- Basic SQL for data analysis
CLOUD & INFRASTRUCTURE:
- AWS, Azure, or GCP cloud platforms
- Object storage (S3, GCS, Azure Blob)
- Multi-cloud and hybrid architectures
- vSphere or on-premises virtualization (nice to have)
AI/ML & EMERGING TECHNOLOGIES
--------------------------------------------------------------------------------
- Experience with AI/ML frameworks for observability (Prophet, TensorFlow,
PyTorch)
- Anomaly detection algorithms and time-series forecasting
- LLM integration for log analysis and troubleshooting (GPT, Claude, etc.)
- Model Context Protocol (MCP) for AI agent integration
- AIOps platforms (Moogsoft, BigPanda, Datadog Watchdog, etc.)
- Natural language processing for log parsing and analysis
- Familiarity with vector databases for semantic search (Pinecone, Weaviate)
- Experience with AI-powered root cause analysis tools
- Knowledge of prompt engineering for observability use cases
PREFERRED QUALIFICATIONS
--------------------------------------------------------------------------------
- Kubernetes certifications (CKA, CKAD, or CKS)
- Grafana certification or equivalent training
- Knowledge of service mesh observability (
- Experience with SRE practices, SLIs, SLOs, and error budgets
- Background in financial services or regulated industries
- Familiarity with compliance requirements (SOX, PCI-DSS, etc.)
- Contributions to open-source observability projects
- Experience with APM tools
Thanks & regards,
Lakshmi Bhavani
Apptoza Inc.
Phone: 437-291-3224 Ext 1105
Mail id : Lakshmi.bhavani@apptoza.com