Key Requirements:We are seeking a skilled Data Engineer to design, deploy, and manage robust data pipelines across hybrid cloud environments. The ideal candidate will combine expertise in Python, AWS services, and enterprise cloud data platforms to build scalable, secure solutions for structured and unstructured data at enterprise scale.
Key Responsibilities:
Design and implement complex ETL/ELT workflows for diverse data sources both on-prem and Cloud platform (AWS)
Build and optimize data lake architectures on Cloudera Data Platform (CDP)
Develop data models in Snowflake cloud warehouse and CDP environments
Implement CI/CD pipelines using Helios for data infrastructure automation
Collaborate with analysts to enable advanced analytics on data lake assets
Ensure security/governance compliance for sensitive data at scale
Technical Expertise Required:
Python: Advanced scripting for data processing, API integration, and automation
AWS : Amazon S3, AWS Glue
Containerization: Docker image management; orchestration via Kubernetes/OCP4
Data Platforms:
Cloudera Data Platform (CDP) for Hadoop/Spark ecosystems
Snowflake cloud data warehouse architecture
Storage Solutions: Data lake design (S3/ADLS), structured/unstructured data optimization
Pipeline Tools: Airflow, Kafka, Helios CI/CD
Version Control: GitHub collaboration
Cloud Services: AWS/Azure storage integration with CDP/Snowflake
Soft Skills:
Problem-solv
ing for hybrid cloud data challenges
Cross-functional collaboration with data scientists and analysts
Clear communication of complex technical solutions
Nice-to-Have Qualifications:
Certifications: Cloudera CCP, SnowPro, Red Hat OpenShift
Experience with streaming platforms (Flink, Kafka Streams)
MLOps patterns for model deployment