Design, develop and maintain scalable ETL pipelines using AWS glue and python, write PySpark scripts for extraction, transformation and loading, configure glue jobs, crawlers and workflows for automated data processing.
Cloud Data Engineering: Work with AWS S3 for data storage and management. Use AWS IAM for access control and security, Leverage AWS Lambda for serverless orchestration and event driven processing.
Monitor Glue jon metrics, using AWS cloud watch, Tune Spark configurations for Optimal performance.
Document ETL processes, data flows, and technical decisions
Experience with infrastructure as code (Terraform, CloudFormation)