Role: Python Developer – Data Engineering
Job Description
We are seeking a skilled Python Developer to join our data engineering team. You will
design, develop, and maintain high-performance data processing pipelines using modern
Python frameworks and tools. In this role, you'll work with large-scale datasets,
containerized systems, and distributed computing platforms to deliver robust data
solutions.
Key Responsibilities
Develop and optimize data manipulation workflows using pandas and polars to handle
large datasets efficiently. Design and implement containerized applications using Docker
and Kubernetes to ensure scalable, reliable deployments. Build and maintain data
pipelines integrating with ClickHouse columnar databases for analytical workloads.
Develop event-driven architectures using NATS messaging systems for asynchronous
data processing. Write comprehensive unit tests using pytest to ensure code quality and
reliability. Implement distributed computing solutions with Dask for processing data
beyond single-machine memory constraints. Manage version control using Git and
collaborate on code repositories following best practices.
Required Skills and Experience
Python & Data Processing: Advanced proficiency in pandas and polars for data
manipulation, transformation, and analysis. Experience optimizing code
performance for large datasets.
Containerization & Orchestration: Hands-on experience with Docker for building
container images and composing multi-container applications. Knowledge of
Kubernetes for container orchestration and deployment management.
Data Infrastructure: Working knowledge of ClickHouse or similar columnar
databases for OLAP workloads and analytical queries.
Messaging & Streaming: Familiarity with NATS.io for building message-driven
systems and asynchronous workflows.
Testing & Quality Assurance: Proficiency with pytest for writing unit tests,
integration tests, and maintaining code coverage standards.
Distributed Computing: Experience with Dask for parallel processing and handling
out-of-core computations.
Version Control: Strong command of Git workflows, branching strategies, and
collaborative development practices.
Preferred Qualifications
Experience with additional Python libraries for data science and machine learning.
Familiarity with CI/CD pipelines and DevOps practices. Background in financial services
or capital markets data systems.