Key Responsibilities
Design and develop scalable ETL/ELT pipelines using Azure Databricks and PySpark.
Migrate BigQuery datasets, schemas, tables, and views to Delta Lake.
Build and optimize batch and incremental data processing pipelines.
Implement Medallion Architecture (Bronze, Silver, Gold) and Delta Lake best practices.
Develop data validation, reconciliation, and quality-check frameworks.
Support historical data migration, performance tuning, and production deployment.
Mandatory Skills
Proven experience with Azure Databricks, including Unity Catalog and Data Lakehouse Architecture.
Strong proficiency in PySpark and Spark SQL.
Expertise in ETL/ELT Development.
Experience with Azure Data Lake Storage (ADLS).
Familiarity with Git and CI/CD practices.
SQL Optimization and Performance Tuning skills.
Experience in Data Migration and Reconciliation Frameworks.
GCP BigQuery Migration Experience is essential.
Preferred Skills
Experience with Airflow / Cloud Composer Dataflow.
Familiarity with DBT (Data Build Tool) for data transformation.
Qualifications
The ideal candidate will possess a bachelor’s degree in computer science, Information Technology, or a related field. A minimum of 7-10 years of experience in data engineering, with a strong focus on cloud-based data solutions, is required. Candidates should demonstrate a solid understanding of data architecture principles and best practices.