Job Title: Data Engineer – PySpark & SQL
Location: Mississauga, ON
Fulltime
We are looking for an experienced 3+ years of Data Engineer with strong hands-on expertise in PySpark, Python, SQL, and data engineering. The ideal candidate should have experience building scalable data pipelines and working with large datasets in cloud or distributed data environments.
Key Responsibilities
- Design, develop, and maintain scalable data pipelines using PySpark and Python.
- Develop complex and optimized SQL queries for data processing and transformation.
- Build ETL/ELT workflows to ingest, transform, and process large datasets.
- Perform data cleansing, validation, transformation, and performance optimization.
- Work with structured and semi-structured datasets.
- Troubleshoot data quality and pipeline performance issues.
- Collaborate with business, analytics, and engineering teams to understand data requirements.
- Follow data engineering best practices, coding standards, and documentation processes.
Required Skills
- Strong hands-on experience with PySpark / Apache Spark.
- Excellent proficiency in SQL, including complex queries, joins, CTEs, window functions, and query optimization.
- Good programming experience with Python.
- Experience developing ETL/ELT data pipelines.
- Good understanding of data warehousing and data modeling concepts.
- Experience working with large-scale/distributed data processing.
- Exposure to at least one major cloud platform such as AWS, Azure, or GCP.
- Strong problem-solving and debugging skills.
Certification Requirement
Candidate must hold at least one relevant, active Data Engineering or Cloud certification, preferably from:
- Databricks Certified Generative AI Engineer Associate
- Google Professional Machine Learning Engineer
- NVIDIA Certified Professional - Agentic AI (NCP-AAI)
- AWS Certified AI Practitioner