We are looking for an experienced Senior Big Data Engineer with strong expertise in Scala, Apache Spark, and Apache Kafka to design and build large-scale, high-performance data processing systems. The ideal candidate will have hands-on experience in real-time and batch data pipelines and a deep understanding of distributed computing.
- Provide consulting services on new initiatives (small to large of varying complexity).
- Explore new emerging technologies and how they best suit our applications
- Develop, code, document, and execute unit tests, system, integration, and acceptance tests using different languages and testing tools for functions of high complexity.
- Ensure adequate technical documentation and training.
- Optimize spark jobs and java applications.
- Architect, design, and implement solutions that meet the stakeholder’s needs
- Participate actively in requirements gathering, data modeling, and design sessions
- Prepare high-level and detailed technical specifications for the projects in accordance with PLC, security, and architecture documentation objectives
- Develop detailed plans and accurate estimates for the completion of build, system testing, and implementation phases of a project
- Develop, code, document, and execute unit tests, systems, integration and acceptance tests, and testing tools for functions of high complexity
- Write, test, and maintain detailed programs according to specifications given by computer software engineers and systems analysts
Knowledge and Experience
- Develop and maintain scalable data pipelines using PySpark and Spark SQL for processing large datasets efficiently.
- Write clean, reusable, and optimized code in Python for data manipulation, analysis, and automation tasks.
- Design and implement ETL workflows to extract, transform, and load data from various structured and unstructured sources.
- Collaborate with data engineers, analysts, and stakeholders to understand data requirements and deliver solutions.
- Optimize Spark jobs for performance tuning, resource utilization, and minimizing execution time.
- Work with distributed computing frameworks to process and analyze big data in a cloud or on-premises environment.
- Develop and maintain unit tests to ensure the accuracy and reliability of data pipelines and transformations.
- Utilize Spark SQL for querying and managing large datasets stored in distributed systems like Hadoop or cloud storage.
- Monitor and troubleshoot data pipeline issues, ensuring reliability and timely delivery of data.
- Stay updated with the latest advancements in PySpark, Spark SQL, and big data technologies to improve existing systems.
What do you need to succeed?
Must Have
- Experience in developing and optimizing Big Data applications using Scala and Spark on Cloudera/HDP.
- Experience in building data pipelines
- Experience in developing/designing micro-service architecture.
- Working knowledge of Jenkins CI, Git, JIRA
- Ability to seek improvements to all aspects of the development process