Skills Required:
Must to Have:
- 14+ years of experience in Enterprise Data Architecture, Hadoop Ecosystem, Big Data Engineering, and Data Platform Modernization.
- Strong hands-on expertise in Hadoop, HDFS, YARN, Hive, Spark, PySpark, Kafka, SQL, Python, Java, and Scala.
- Experience designing and implementing large-scale Big Data and Data Lake architectures supporting analytics, reporting, and AI workloads.
- Strong understanding of Hadoop platform architecture, including HDFS, YARN, high availability, cluster management, workload optimization, and scalability.
- 14+ years of experience in Enterprise Data Architecture, Hadoop Ecosystem, Big Data Engineering, and Data Platform Modernization.
- Strong hands-on expertise in Hadoop, HDFS, YARN, Hive, Spark, PySpark, Kafka, SQL, Python, Java, and Scala.
- Experience designing and implementing large-scale Big Data and Data Lake architectures supporting analytics, reporting, and AI workloads.
- Strong understanding of Hadoop platform architecture, including HDFS, YARN, high availability, cluster management, workload optimization, and scalability.
- Expertise in batch and distributed data processing using Spark, Hive, HBase, and MapReduce frameworks.
- Experience building data ingestion and streaming solutions using Kafka and Hadoop ecosystem tools.
- Strong knowledge of Data Engineering practices, including ETL/ELT, data integration, data quality, metadata management, lineage, observability, and DataOps.
- Experience with Data Lake, Data Warehousing, Enterprise Data Models, Master Data, and Data Governance frameworks.
- Strong understanding of security and governance controls, including Kerberos, Ranger, encryption, privacy, auditability, compliance, and controlled data access.
- Proven experience in architecture definition, HLD/LLD creation, migration planning, technical governance, production readiness, and performance optimization.
- Experience supporting platform modernization, cloud adoption, solutioning, effort estimation, and client-facing technical discussions.
- Good understanding of BFSI regulatory, security, privacy, resilience, and audit requirements.
- Preferred certifications in Hadoop, Spark, Cloud Data Engineering, Enterprise Architecture, Data Management, Security, or Big Data Platforms.
Roles and Responsibilities:
- Engage with business and technology stakeholders to understand business objectives, data platform requirements, and modernization goals.
- Define end-to-end architecture for Hadoop and Big Data platforms supporting analytics, reporting, and AI workloads.
- Design scalable, secure, and high-performance Hadoop ecosystems using HDFS, YARN, Spark, Kafka, Hive, and related technologies.
- Architect data ingestion, batch processing, streaming, and data lake solutions for enterprise-scale environments.
- Define data models, data lake zones, governance frameworks, metadata management, lineage, and security standards.
- Lead Hadoop modernization, migration, and platform optimization initiatives to improve performance, scalability, and operational efficiency.
- Establish architecture standards, reusable frameworks, engineering best practices, and DataOps operating models.
- Provide technical leadership across solution design, implementation, testing, migration, deployment, and production support activities.
- Conduct architecture reviews, identify risks and dependencies, and guide teams on mitigation strategies and best practices.
- Drive performance tuning, capacity planning, reliability improvements, and infrastructure cost optimization.
- Collaborate with data engineers, platform teams, architects, vendors, and client stakeholders to deliver enterprise data solutions.
- Support solutioning, RFP responses, effort estimation, technical presentations, and client workshops.
- Mentor architects and engineers and contribute to hiring, capability development, and knowledge-sharing initiatives.
Develop reusable accelerators, reference architectures, implementation frameworks, and technical best practices for Big Data platforms