At a glance
Automatically prepared from the listing. Check the original description for the full requirements.
Jump to the original descriptionResponsibilities
Design, develop, and optimize scalable data pipelines and real-time processing solutions using PySpark, Spark, and Kafka. Collaborate with cross-functional teams to implement ETL/ELT workflows and ensure data quality and governance.
Requirements
Requires strong expertise in Big Data technologies including the Hadoop ecosystem, Apache NiFi, and Python. Candidates should be proficient in SQL and experienced in building event-driven architectures and data ingestion solutions.
Working hours
40 hours per week
Skills
- PySpark
- Apache Spark
- Apache Kafka
- Hadoop
- Apache NiFi
- SQL
- Python
- Spark Streaming
- Airflow
- Oozie
- Hive
- Scala
- Jenkins
- Bitbucket
- Git
- Cloud Platforms
Visa sponsorship
Not detected in the job text
Categories
- Data & Analytics
- Technology
- Software
- Engineering
- Consulting
Keywords
- PySpark
- Apache Spark
- Apache Kafka
- Hadoop
- HDFS
- Hive
- YARN
- Apache NiFi
- SQL
- Python
- Spark Streaming
- Airflow
- Oozie
- Scala
- Jenkins
- Bitbucket
- Git
- JIRA
- Confluence
- GCP
- AWS
- Azure
- Data Warehousing
- Dimensional Modeling
- ETL
- ELT
- Big Data
- Data Pipelines
- Event-Driven Architecture
- Real-time Processing
Original job description
An experienced Data Engineer with strong expertise in Big Data technologies to design, develop, and support enterprise-scale data platforms. The ideal candidate should possess hands-on experience in PySpark, Apache Spark, Kafka, Hadoop ecosystem components, and Apache NiFi, with a strong understanding of data ingestion, transformation, and real-time processing frameworks.
Key Responsibilities
Design, develop, and optimize scalable data pipelines using PySpark, Spark, Hadoop, and Apache NiFi.
Build and maintain batch and real-time data processing solutions.
Develop and support Kafka-based streaming applications and event-driven architectures.
Create and optimize ETL/ELT workflows for large-scale structured and unstructured datasets.
Develop complex SQL queries for data extraction, transformation, validation, and troubleshooting.
Implement data ingestion solutions from databases, APIs, files, and streaming sources.
Monitor, troubleshoot, and enhance the performance of Spark jobs and data pipelines.
Collaborate with architects, business analysts, and development teams to deliver high-quality data solutions.
Support platform upgrades, deployments, testing, certification, and production releases.
Ensure data quality, governance, security, and operational excellence across data platforms.
Mandatory Skills
PySpark
Apache Spark (Spark SQL, DataFrames)
Apache Kafka
Hadoop Ecosystem (HDFS, Hive, YARN)
Apache NiFi
SQL
Python
Preferred Skills
Spark Streaming
Airflow / Oozie
Hive
Scala
Jenkins, Bitbucket, Git
JIRA, Confluence
Cloud Platforms (GCP/AWS/Azure)
Data Warehousing concepts and Dimensional Modeling