Role: Python Developer – Data Engineering
Location: Toronto, ON (4 days hybrid)
Fulltime
Responsibilities:
- Develop and optimize data manipulation and transformation workflows using pandas and Polars.
- Design, build, and maintain efficient data processing pipelines for large-scale datasets.
- Optimize Python code and data workflows for performance, scalability, and memory efficiency.
- Develop containerized applications using Docker and deploy scalable services using Kubernetes.
- Build and maintain data pipelines integrated with ClickHouse or similar columnar databases for analytical workloads.
- Design and implement event-driven architectures using NATS for asynchronous data processing and messaging.
- Develop distributed computing solutions using Dask to process datasets that exceed single-machine memory limitations.
- Write and maintain comprehensive unit and integration tests using pytest.
- Maintain high code quality, test coverage, reliability, and performance standards.
- Use Git for version control and follow established branching, code review, and collaboration practices.
- Collaborate with data engineers, software developers, DevOps teams, and other stakeholders to deliver reliable data solutions.
- Troubleshoot production issues and support the continuous improvement of data engineering processes.
Requirements:
- Bachelor’s degree in computer science or Diploma in IT with equivalent experience.
Python and Data Processing
- Advanced proficiency in Python.
- Strong experience with pandas and Polars for data manipulation, transformation, and analysis.
- Experience working with large datasets and optimizing code for performance and memory usage.
- Strong understanding of Python development best practices, object-oriented programming, and clean code principles.
Containerization and Orchestration
- Hands-on experience building and managing container images using Docker.
- Experience creating and maintaining multi-container applications.
- Working knowledge of Kubernetes for container orchestration, deployment, scaling, and service management.
Data Infrastructure
- Experience with ClickHouse or a similar columnar database.
- Understanding of OLAP workloads, analytical queries, data modeling, and query optimization.
- Experience designing data pipelines for analytical and reporting use cases.
Messaging and Streaming
- Familiarity with NATS.io and message-driven architectures.
- Understanding of asynchronous processing, event-driven systems, messaging patterns, and distributed workflows.
Testing and Quality Assurance
- Proficiency with pytest.
- Experience writing unit tests, integration tests, and regression tests.
- Understanding of test coverage, mocking, test automation, and quality assurance practices.
Distributed Computing
- Experience with Dask for parallel processing and out-of-core computation.
- Understanding of distributed data processing and techniques for handling datasets that exceed available system memory.
Version Control
- Strong command of Git, including branching strategies, pull requests, code reviews, merging, and conflict resolution.
- Experience collaborating through shared code repositories and following software development best practices.