Archived listing. The details below describe a past opening.
Role :: Big Data Engineer
Location :: Mississauga, ON– Onsite
Fulltime_ Permanent Role
Job Description
- Design and develop Big Data pipelines using Spark and Scala to process large volumes of raw data stored in Hive/Hadoop.
- Analyze raw Hive tables, source schemas, data dictionaries and source-to-target mappings to understand business data.
- Develop Spark/Scala transformation frameworks to convert raw source data into standardized conformed data views/tables.
- Implement complex business transformation rules including joins, filters, aggregations, derivations and reference-data lookups.
- Standardize data across multiple source systems to create a consistent enterprise/conformed data model.
- Implement incremental processing, partition-based processing and historical data handling for high-volume datasets.
- Develop reusable Spark/Scala components for data enrichment, transformation and validation.
- Implement data-quality controls for completeness, uniqueness, validity, referential integrity and business-rule validation.