Title : Senior Data Engineer (Databricks)
Location: Toronto, ON Hybrid
Position Type: Full Time
Responsibilities
- Design, develop, and maintain scalable ETL/ELT pipelines for structured and unstructured data.
- Build and optimize data solutions using Databricks, PySpark, and Spark SQL.
- Implement data ingestion frameworks from various source systems including databases, APIs, files, and streaming platforms.
- Develop and maintain data lakehouse architectures leveraging Databricks best practices.
- Design and implement Medallion Architecture (Bronze, Silver, Gold layers) for enterprise data platforms.
- Collaborate with business stakeholders, data scientists, architects, and application teams to understand data requirements.
- Optimize Spark jobs and data pipelines for performance, scalability, and cost efficiency.
- Implement data quality, governance, security, and monitoring frameworks.
- Support AI/ML initiatives by preparing and engineering datasets for model development and deployment.
- Develop CI/CD pipelines and automate deployment processes for data engineering workloads.
- Participate in architecture reviews and establish data engineering standards and best practices.
- Mentor junior engineers and provide technical leadership across projects.
Required Qualifications
- Bachelor's or Master's degree in Computer Science, Information Technology, Data Science, or related field.
- 8+ years of experience in Data Engineering and Data Platform development.
- Strong hands-on experience with Databricks and Apache Spark.
- Proficiency in Python, PySpark, and SQL.
- Experience with cloud platforms such as:
- Microsoft Azure (preferred)
- AWS
- Google Cloud Platform
- Experience with Delta Lake, Unity Catalog, and Databricks workflows.
- Strong understanding of Data Lake, Data Warehouse, and Lakehouse architectures.
- Experience working with large-scale distributed data processing systems.
- Knowledge of data modeling, partitioning, indexing, and query optimization techniques.
- Experience with Git, CI/CD pipelines, and DevOps practices.
Preferred Qualifications
- Databricks Certified Data Engineer Associate or Professional Certification.
- Experience with Azure Data Factory, Synapse Analytics, or equivalent cloud-native services.
- Experience with streaming technologies such as Kafka, Event Hubs, or Spark Streaming.
- Exposure to AI/ML platforms, MLOps, and Generative AI solutions.
- Experience working in Insurance, Financial Services, or Enterprise Digital Transformation programs.
Technical Skills
- Category Skills
- Data Engineering Databricks, Apache Spark, PySpark, Spark SQL
- Programming Python, SQL, Scala (Preferred)
- Cloud Azure, AWS, GCP
- Data Platforms Delta Lake, Unity Catalog, Data Lakehouse
- ETL/ELT Azure Data Factory, Databricks Workflows, Airflow
- DevOps Git, Azure DevOps, Jenkins, CI/CD
- Streaming Kafka, Spark Streaming, Event Hubs
- Databases SQL Server, PostgreSQL, Oracle, Snowflake