Job Summary
Work with cutting-edge big data platforms (e.g., Databricks, Apache Spark) at large scale, pushing the boundaries of data processing and model enablement.
•
Build and maintain robust ETL/ELT pipelines for ingestion, transformation, and aggregation of large-scale datasets on Hadoop and enterprise data platforms.
•
Develop high-performance data processing jobs using PySpark/Spark, Python on data platforms such as cloudera and databricks.
•
Optimize pipeline performance and cost through partitioning, file formats, compute tuning, and efficient query patterns
•
Contribute to CI/CD for data workflows (testing, code reviews, deployment automation), promoting engineering best practices and maintainable codebases.
•
Partner with Product Managers to develop a deep understanding of users and use cases and apply that knowledge to scoping and building new modules and features
Ideal Candidate Qualifications:
•
Strong hands-on experience in data engineering building production-grade pipelines on big data platforms (Hadoop ecosystem and cloud data platforms - databricks).
•
High proficiency in using Python, Spark, Hadoop platforms & tools (Hive, Impala, Airflow, NiFi), SQL to build Big Data products.
•
Hands-on experience with cloud data platforms such as databricks, snowflake (databricks preferred)
•
Experience with orchestration/integration tools such as Apache Airflow, Apache NiFi, or Talend.
•
Working knowledge of DevOps/CI-CD practices: version control (Git), automated testing, release pipelines, and observability.
•
Strong problem-solving skills with the ability to debug complex data issues and communicate clearly with technical and non-technical stakeholders.
•
Experience developing Java based applications is an added advantage.
LONGDESCRIPTION section. 2 of 6.
Section Title: Key Responsibilities
Key Responsibilities
Work with cutting-edge big data platforms (e.g., Databricks, Apache Spark) at large scale, pushing the boundaries of data processing and model enablement.
•
Build and maintain robust ETL/ELT pipelines for ingestion, transformation, and aggregation of large-scale datasets on Hadoop and enterprise data platforms.
•
Develop high-performance data processing jobs using PySpark/Spark, Python on data platforms such as cloudera and databricks.
•
Optimize pipeline performance and cost through partitioning, file formats, compute tuning, and efficient query patterns
•
Contribute to CI/CD for data workflows (testing, code reviews, deployment automation), promoting engineering best practices and maintainable codebases.
•
Partner with Product Managers to develop a deep understanding of users and use cases and apply that knowledge to scoping and building new modules and features
Ideal Candidate Qualifications:
•
Strong hands-on experience in data engineering building production-grade pipelines on big data platforms (Hadoop ecosystem and cloud data platforms - databricks).
•
High proficiency in using Python, Spark, Hadoop platforms & tools (Hive, Impala, Airflow, NiFi), SQL to build Big Data products.
•
Hands-on experience with cloud data platforms such as databricks, snowflake (databricks preferred)
•
Experience with orchestration/integration tools such as Apache Airflow, Apache NiFi, or Talend.
•
Working knowledge of DevOps/CI-CD practices: version control (Git), automated testing, release pipelines, and observability.
•
Strong problem-solving skills with the ability to debug complex data issues and communicate clearly with technical and non-technical stakeholders.
•
Experience developing Java based applications is an added advantage.
LONGDESCRIPTION section. 3 of 6.
Section Title: Skill Requirements
Skill Requirements
Work with cutting-edge big data platforms (e.g., Databricks, Apache Spark) at large scale, pushing the boundaries of data processing and model enablement.
•
Build and maintain robust ETL/ELT pipelines for ingestion, transformation, and aggregation of large-scale datasets on Hadoop and enterprise data platforms.
•
Develop high-performance data processing jobs using PySpark/Spark, Python on data platforms such as cloudera and databricks.
•
Optimize pipeline performance and cost through partitioning, file formats, compute tuning, and efficient query patterns
•
Contribute to CI/CD for data workflows (testing, code reviews, deployment automation), promoting engineering best practices and maintainable codebases.
•
Partner with Product Managers to develop a deep understanding of users and use cases and apply that knowledge to scoping and building new modules and features
Ideal Candidate Qualifications:
•
Strong hands-on experience in data engineering building production-grade pipelines on big data platforms (Hadoop ecosystem and cloud data platforms - databricks).
•
High proficiency in using Python, Spark, Hadoop platforms & tools (Hive, Impala, Airflow, NiFi), SQL to build Big Data products.
•
Hands-on experience with cloud data platforms such as databricks, snowflake (databricks preferred)
•
Experience with orchestration/integration tools such as Apache Airflow, Apache NiFi, or Talend.
•
Working knowledge of DevOps/CI-CD practices: version control (Git), automated testing, release pipelines, and observability.
•
Strong problem-solving skills with the ability to debug complex data issues and communicate clearly with technical and non-technical stakeholders.
•
Experience developing Java based applications is an added advantage.
SKILL section. 4 of 6.
Section Title: Must Have Skills
Must Have Skills
- Click Enter to show the proficiency description of Data EngineeringData Engineering
- Click Enter to show the proficiency description of DatabricksDatabricks
- Click Enter to show the proficiency description of Apache SparkApache Spark
- Click Enter to show the proficiency description of PySparkPySpark
- Click Enter to show the proficiency description of PythonPython
- Click Enter to show the proficiency description of CI/CDCI/CD
SKILL section. 5 of 6.
Section Title: Good to have Skills
Good to have Skills
LONGDESCRIPTION section. 6 of 6.
Section Title: Other Requirements
Other Requirements
Work with cutting-edge big data platforms (e.g., Databricks, Apache Spark) at large scale, pushing the boundaries of data processing and model enablement.
•
Build and maintain robust ETL/ELT pipelines for ingestion, transformation, and aggregation of large-scale datasets on Hadoop and enterprise data platforms.
•
Develop high-performance data processing jobs using PySpark/Spark, Python on data platforms such as cloudera and databricks.
•
Optimize pipeline performance and cost through partitioning, file formats, compute tuning, and efficient query patterns
•
Contribute to CI/CD for data workflows (testing, code reviews, deployment automation), promoting engineering best practices and maintainable codebases.
•
Partner with Product Managers to develop a deep understanding of users and use cases and apply that knowledge to scoping and building new modules and features
Ideal Candidate Qualifications:
•
Strong hands-on experience in data engineering building production-grade pipelines on big data platforms (Hadoop ecosystem and cloud data platforms - databricks).
•
High proficiency in using Python, Spark, Hadoop platforms & tools (Hive, Impala, Airflow, NiFi), SQL to build Big Data products.
•
Hands-on experience with cloud data platforms such as databricks, snowflake (databricks preferred)
•
Experience with orchestration/integration tools such as Apache Airflow, Apache NiFi, or Talend.
•
Working knowledge of DevOps/CI-CD practices: version control (Git), automated testing, release pipelines, and observability.
•
Strong problem-solving skills with the ability to debug complex data issues and communicate clearly with technical and non-technical stakeholders.
•
Experience developing Java based applications is an added advantage