Job Title: Chaos Engineer
Job Location :- Toronto ,ON
Experience: 5+ Years
Key Responsibilities
- Design and execute resiliency and chaos testing scenarios across cloud, Kubernetes, APIs, microservices, and distributed environments.
- AWS FIS knowledge, setup and Run Experiments
- AWS CDK knowledge.
- Identify resilience gaps, SPOFs, and operational risks and drive remediation.
- Validate High Availability (HA), Disaster Recovery (DR), failover, auto-healing, and recovery processes.
- Automate testing and reliability validation using scripting and cloud-native tools.
- Monitor application and infrastructure behavior using observability platforms such as AppDynamics, Prometheus, and Grafana.
- Collaborate with DevOps, SRE, Infrastructure, and Application teams to improve system reliability.
- Define and track reliability metrics including SLA, SLO, SLI, MTTR, RTO, and RPO.
Required Skills
- Strong experience in Java/Python, Microservices, APIs, Kubernetes, Docker, and Cloud Platforms (Azure/GCP/AWS).
- Hands-on experience with monitoring and observability tools such as Dynatrace metrics and observability, AppDynamics, Prometheus, and Grafana.
- Knowledge of Reliability Engineering, Chaos Testing, Incident Analysis, and Resiliency Validation.
- Failure-as-a-Service platforms to achieve resiliency in infrastructure failures, network and application failures.
- Chaos Testing mechanism and tools such as AWS-FIS, Lambda testing, On-premise , OpenShift testing.
- Monitoring tools test/scenario capture and report creation through standardized templates and hypothesis formation.
- Experience with automation, CI/CD, and cloud-native architectures.
- Excellent troubleshooting, analytical, and communication skills.