Data Platform Engineering
•
Design, build, and support scalable data pipelines using Python, Spark, and Databricks.
•
Modernize legacy data processing workloads and migrate existing pipelines to secure cloud-native platforms.
•
Implement data quality, monitoring, observability, and operational controls.
•
Optimize performance, scalability, reliability, and cost efficiency of data workloads.
Data Ingestion & Integration
•
Build and maintain ingestion pipelines for structured and unstructured data sources.
•
Connect to APIs to ingest data from internal and external systems.
•
Build and maintain batch-processing pipelines, with occasional work on streaming pipelines.
•
Integrate data from SharePoint, document repositories, enterprise applications, and cloud platforms.
•
Follow established security and data-handling practices when connecting to sources, managing pipeline credentials, and writing data.
Document Intelligence & Automation
•
Develop document extraction, classification, and metadata enrichment pipelines.
•
Support AI-assisted processing of large-scale document repositories.
•
Automate data transformation and enrichment workflows.
Platform Modernization
•
Implement migration and refactoring of existing data assets and pipelines.
•
Apply software engineering practices including CI/CD, automated testing, and code reviews.
•
Understand the purpose and broader context of assigned work, identify gaps or assumptions, and suggest changes or alternative approaches when they would better achieve the intended outcome.
Engineering Excellence
•
Collaborate with architects, tax subject matter experts, developers, and platform teams.
•
Use AI-assisted development tools such as Codex to build data pipelines, and review and test the resulting code.
Required Qualifications
Technical Skills
•
3+ years of experience in Data Engineering.
•
Strong Python development experience.
•
Strong Apache Spark and PySpark experience.
•
Experience with Databricks.
•
SQL expertise and data modeling experience.
•
Experience building, testing, deploying, and troubleshooting production ETL/ELT pipelines.
Software Engineering
•
Experience with Git-based source control and CI/CD pipelines.
•
Experience consuming REST APIs for data ingestion, including authentication, pagination, and handling failures.
•
Experience writing automated tests for data pipelines.
Preferred Qualifications
•
Experience with Azure cloud services.
•
Experience with document-processing or intelligent document extraction solutions.
•
Experience using AI-assisted development tools such as Codex, GitHub Copilot, or similar developer productivity platforms.
Desired Characteristics
•
Strong problem-solving and troubleshooting skills.
•
Exercise judgment about when to proceed independently, ask for clarification, or challenge the proposed approach.
•
Comfortable analyzing unfamiliar codebases and modernizing legacy implementations.
•
Ability to identify and explain improvements to existing data pipelines.
•
Strong communication and collaboration skills.
•
Willingness to learn new tools and data sources.
Success Criteria
Success in this role means:
•
Modernize key Spark/Python workloads.
•
Improve security, maintainability, and reliability of existing pipelines.
•
Deliver scalable ingestion and document processing capabilities.
•
Use AI-assisted development practices to deliver maintainable, tested pipeline code.