Position Name – SRE
Type of hiring – Subcon
Location – Brampton, ON or Toronto, ON (Hybrid)
Job Description:
Key Responsibilities & Skills:
- Monitor application and infrastructure health, performance, and availability.
- Troubleshoot production incidents, perform root cause analysis, and drive issue resolution.
- Implement and manage monitoring, alerting, logging, and observability solutions.
- Support cloud platforms, Kubernetes, Docker, and infrastructure operations.
- Automate operational processes and improve system reliability and efficiency.
- Support CI/CD pipelines, application deployments, and release activities.
- Perform capacity planning, performance optimization, and disaster recovery planning.
- Collaborate with engineering, cloud, and security teams to improve platform resiliency.
- Create operational dashboards, runbooks, and technical documentation.
Required Skills:
- 6+ Years of experience in Site Reliability Engineering, Production Support, DevOps, or Cloud Operations.
- Experience supporting production applications and cloud environments (Azure preferred).
- Strong knowledge of Kubernetes, Docker, Linux, networking, and system administration.
- Experience with monitoring tools such as Dynatrace, Grafana, or Azure Monitor.
- Scripting and automation experience using Python, PowerShell, or Bash.