Site Reliability Engineer – Production Reliability, Azure Operations & Databricks Support
Location: Toronto, ON
Work Model: Hybrid
Experience: 3+ Years
Job Overview
We are seeking an Intermediate Site Reliability Engineer (SRE) to support and continuously improve enterprise Microsoft Azure and Databricks environments. The role focuses on production reliability, monitoring, incident response, platform operations, availability, and operational readiness.
The successful candidate will work closely with Cloud, Platform Engineering, Data, Application, Network, Security, and Support teams to maintain stable, secure, scalable, and highly available production environments.
Key Responsibilities
- Monitor and support production Azure and Databricks environments, ensuring platform availability, performance, reliability, and operational readiness.
- Troubleshoot and resolve production incidents involving Azure infrastructure, Databricks, networking, storage, identity, access, and application integrations.
- Participate in on-call rotations, incident bridges, service requests, escalations, and emergency production support.
- Support Databricks workspaces, clusters, compute, cluster policies, jobs, workflows, user access, monitoring, and cost management.
- Support Unity Catalog, including catalogs, schemas, permissions, storage credentials, and external locations.
- Support integrations across Databricks, ADLS Gen2, Azure Data Factory, Azure SQL, Key Vault, and Managed Identities.
- Support Azure networking components including VNets, NSGs, route tables, private endpoints, DNS, VPN/ExpressRoute, and hub-and-spoke connectivity.
- Support Azure Storage, Blob Storage, and ADLS Gen2, including access controls, availability, lifecycle management, and troubleshooting.
- Develop and maintain alerts, dashboards, and monitoring using Azure Monitor, Log Analytics, Grafana, Prometheus, Dynatrace, Datadog, or New Relic.
- Perform production troubleshooting, service validation, maintenance activities, patching, upgrades, and planned changes.
- Participate in Root Cause Analysis (RCA), Problem Management, and remediation activities to eliminate recurring incidents.
- Support Business Continuity and Disaster Recovery exercises and platform recovery activities.
- Maintain operational runbooks, knowledge articles, support procedures, troubleshooting guides, and technical documentation.
- Manage incidents, requests, changes, and operational tasks using ServiceNow and JIRA.
- Work with engineering and platform teams to identify opportunities to improve reliability, automation, observability, performance, and operational efficiency.
- Contribute to daily standups, planning sessions, service reviews, and operational readiness activities.
Required Skills & Experience
- 3+ years of experience supporting Microsoft Azure cloud infrastructure in production environments.
- 1+ year of hands-on experience supporting Databricks environments.
- Hands-on experience with Azure Storage, ADLS Gen2, and Blob Storage.
- Strong understanding of Azure networking concepts including VNets, NSGs, routing, DNS, and private endpoints.
- Experience with Azure identity and security concepts including Microsoft Entra ID, RBAC, Managed Identities, and Azure Key Vault.
- Experience with cloud monitoring and observability tools such as Azure Monitor, Log Analytics, Grafana, Prometheus, Dynatrace, Datadog, or New Relic.
- Experience troubleshooting production incidents and participating in on-call support.
- Understanding of Incident Management, Change Management, Problem Management, escalation procedures, and RCA.
- Experience working with ServiceNow, JIRA, operational runbooks, and enterprise support processes.
- 1+ year of Windows Server administration experience.
- 1+ year of Linux administration experience.
- Basic understanding of network troubleshooting, connectivity, TCP/IP, DNS, routing, and firewall concepts.
- Strong troubleshooting, analytical, and problem-solving skills.
Databricks & Data Platform Experience
- Experience supporting Databricks workspaces, clusters, jobs, workflows, compute, cluster policies, and user access.
- Understanding of Unity Catalog, catalogs, schemas, permissions, storage credentials, and external locations.
- Experience supporting integrations between Databricks and ADLS Gen2, Azure Data Factory, Azure SQL, Key Vault, and Managed Identities.
- Ability to troubleshoot Databricks platform, connectivity, authentication, job, compute, and access-related issues.
Nice-to-Have Skills
- 1+ year of Azure SQL operational support experience.
- 1+ year of Azure Data Factory operational support, including linked services, integration runtimes, pipelines, and Databricks orchestration.
- Experience with Business Continuity and Disaster Recovery, including RTO and RPO concepts.
- Experience supporting enterprise data and analytics platforms.
- Experience supporting AI/GenAI platforms, Azure OpenAI, model endpoints, RAG services, or MLOps environments.
- Experience with capacity planning, platform health reporting, cost monitoring, and performance optimization.
- Exposure to automation and scripting using PowerShell, Python, Bash, or similar technologies.
- Experience improving operational processes through automation, monitoring, and self-healing capabilities.
Incident & Production Support
- Strong experience working in a production support/SRE environment.
- Ability to triage incidents, identify service impact, troubleshoot systematically, and restore services within established SLAs.
- Experience participating in incident bridges and major incident response.
- Ability to communicate technical status, impact, remediation steps, and recovery progress to technical and business stakeholders.
- Experience conducting RCA and post-incident reviews and implementing corrective and preventive actions.
- Willingness to participate in on-call and after-hours production support as required.
Soft Skills
- Strong analytical and troubleshooting skills with a structured approach to problem solving.
- Calm and effective under pressure during critical production incidents.
- Ability to work independently while knowing when and how to escalate issues.
- Strong collaboration skills across Cloud, Data, Application, Network, Security, and Platform Engineering teams.
- Strong communication skills with the ability to explain technical issues clearly to both technical and non-technical stakeholders.
- Strong documentation and knowledge-sharing skills.
- Organized and comfortable managing multiple incidents, requests, and operational priorities.
- Proactive mindset with a focus on reliability, automation, continuous improvement, and operational excellence.
- Willingness to learn new technologies and adapt to evolving cloud and data-platform operating models.