Role: Python Developer – Data Engineering
Location: Toronto, ON (4 days hybrid)
Fulltime
Requirements:
- Bachelor’s degree in computer science or Diploma in IT with equivalent experience.
Python and Data Processing
- Advanced proficiency in Python.
- Strong experience with pandas and Polars for data manipulation, transformation, and analysis.
- Experience working with large datasets and optimizing code for performance and memory usage.
- Strong understanding of Python development best practices, object-oriented programming, and clean code principles.
Containerization and Orchestration
- Hands-on experience building and managing container images using Docker.
- Experience creating and maintaining multi-container applications.
- Working knowledge of Kubernetes for container orchestration, deployment, scaling, and service management.
Data Infrastructure
- Experience with ClickHouse or a similar columnar database.
- Understanding of OLAP workloads, analytical queries, data modeling, and query optimization.
- Experience designing data pipelines for analytical and reporting use cases.
Messaging and Streaming
- Familiarity with NATS.io and message-driven architectures.
- Understanding of asynchronous processing, event-driven systems, messaging patterns, and distributed workflows.
Testing and Quality Assurance
- Proficiency with pytest.
- Experience writing unit tests, integration tests, and regression tests.
- Understanding of test coverage, mocking, test automation, and quality assurance practices.
Distributed Computing
- Experience with Dask for parallel processing and out-of-core computation.
- Understanding of distributed data processing and techniques for handling datasets that exceed available system memory.
Version Control
- Strong command of Git, including branching strategies, pull requests, code reviews, merging, and conflict resolution.
- Experience collaborating through shared code repositories and following software development best practices.
Preferred skills:
- Experience with additional Python libraries for data science, analytics, or machine learning.
- Familiarity with CI/CD pipelines and DevOps practices.
- Experience with cloud platforms and cloud-based data engineering solutions.
- Knowledge of data orchestration tools and workflow scheduling platforms.
- Experience working with financial services, capital markets, trading, or other high-volume data systems.