Position: Senior Azure Databricks Consultant
Location: Burnaby Canada Hybrid
Duration: Long Term Contract
Job Description:
About the Role
We’re seeking an intermediate-level Azure Databricks Engineer to build and optimize scalable data pipelines, manage Delta Lake-based data products, and support analytics/ML workloads. You’ll collaborate with data engineers, analysts, and platform teams to deliver reliable, well-documented, and secure solutions on Azure.
Key Responsibilities
- Design, build, and maintain PySpark/SQL pipelines in Azure Databricks for batch and streaming data.
- Develop robust ingestion from Azure Data Lake Storage (ADLS Gen2), Azure Synapse/SQL, Event Hub, Kafka, and REST/JSON sources.
- Optimize Spark jobs (partitioning, caching, broadcast joins, AQE) for performance and cost.
- Implement monitoring and alerting (cluster/job metrics, driver/executor logs).
- Use Databricks Repos, notebooks, and modular PySpark projects with unit tests (pytest).
- Build CI/CD pipelines (e.g., Azure DevOps, GitHub Actions) for jobs, notebooks, and infrastructure-as-code (Terraform/ARM/Bicep).
- Manage environments (dev/test/prod), secrets/Key Vault, and configuration promotion.
Required Qualifications (Intermediate Level)
- 6+ years in data engineering; 4+ years hands-on with Azure Databricks and Spark.
- Strong PySpark and SQL skills: DataFrames, joins, window functions, UDFs, incremental loads.
- Practical experience with Delta Lake, Unity Catalog, and Databricks Jobs/Workflows.
- Familiarity with Azure services: ADLS Gen2, Azure Key Vault, Event Hub, Azure SQL/Synapse.
- Version control (Git) and CI/CD experience; basic testing practices (pytest).
- Ability to optimize Spark jobs and troubleshoot: skew, shuffle, OOM, driver/executor tuning.
- Solid understanding of data modeling (star schema, medallion/lakehouse), partitioning, and file formats (Parquet/JSON).
- Airflow, Azure Data Factory orchestration.
- Terraform for Databricks & Azure resources.
- Basic Scala and/or SQL Warehouses (Databricks SQL) for BI.
Education
- Bachelor’s/Master’s in Computer Science, Engineering, or related field (or equivalent experience).
Certifications (Optional but Valued)
- Databricks: Data Engineer Associate/Professional
- Microsoft Azure: DP-203 (Data Engineering on Microsoft Azure), AZ-900 (Fundamentals)
Tools & Tech Stack (Typical)
- Languages: Python (PySpark), SQL
- Databricks: Notebooks, Jobs/Workflows, Repos, Unity Catalog, Delta Lake, MLflow
- Azure: ADLS Gen2, Key Vault, Event Hub, Synapse/SQL, Monitor/Log Analytics
- DevOps: Git, Azure DevOps/GitHub Actions, Terraform/Bicep