Role Overview
This is Precision AI's first dedicated data engineering role. Our AI team trains computer vision models on several hundred terabytes of drone imagery and flight telemetry collected across real growing seasons. That archive lives in S3 as Parquet and imagery, queried with Athena. Every training run depends on tribal knowledge about which prefixes hold which season.
You will build the platform that fixes this. You own how field data gets from the aircraft into a governed, queryable, reproducible asset that our ML engineers can pull training sets from with confidence, and that the business can eventually report on.
This is a hands-on senior individual contributor role reporting to the AI Team Lead. You will not manage people. You will set the technical direction for data at Precision AI and be measured on whether the AI team can move faster because of it.
This hybrid role is based in Calgary and will work from Precision AI’s headquarters 3 days a week.
Key Responsibilities
Data Platform Architecture
- Design the lakehouse layer over our existing S3 and Parquet footprint, including table format selection (Iceberg or Delta), partitioning strategy, schema evolution, and a searchable catalogue.
- Decide what stays as raw imagery and what becomes a managed table, and document why.
Pipeline Engineering and Orchestration
- Stand up orchestration from scratch.
- Evaluate and select the tooling (Airflow, Dagster, Prefect, or equivalent), then build batch pipelines that are idempotent, backfillable, and observable.
Field-to-Cloud Ingestion
- Build reliable ingestion for drone imagery and telemetry captured in the field, often over poor connectivity.
- Handle validation, deduplication, and integrity checks close to capture.
ML Data Enablement
- Work directly with AI engineers to turn raw captures into curated, versioned training and evaluation datasets.
- Own dataset lineage and versioning so that any model in MLflow can be traced back to the exact data that produced it, two years later. Support annotation workflows and our vector database.
Data Quality and Reliability
- Define data contracts, automated tests, and freshness and anomaly monitoring.
- Own lineage from raw capture through to the datasets and features that models consume. Establish the data quality standards the team codes against.
Cost and Performance Management
- Own S3 storage class lifecycle, file compaction, and Athena scan cost.
- At our scale, storage and query layout decisions are the primary cost lever, and we expect you to treat that as an engineering problem rather than a finance one.
Analytics Enablement
- As the platform matures, extend it to serve business and product analytics: a modelled warehouse layer, transformation tooling such as dbt, and a semantic layer for reporting.
- This is a later phase of the role, not a day-one responsibility, and you will help decide when it becomes the priority.
Engineering Practice
- Write production-grade Python.
- Use infrastructure as code, containerization, CI/CD, and version control as defaults rather than afterthoughts.
- Raise the bar across the AI team through code review and by setting the data engineering standards other engineers build on.
Technical Communication
- Communicate design decisions, tradeoffs, and progress clearly to engineers and to non-technical partners.
Relevant Experience
- 4+ years building and operating production data platforms, with clear ownership of systems you designed rather than only maintained.
- Strong Python, with solid software engineering fundamentals: testing, code structure, version control, and code review.
- Deep hands-on AWS experience, particularly S3, Athena or equivalent query engines, IAM, and cost management at scale.
- Production experience with a workflow orchestrator (Airflow, Dagster, Prefect, Step Functions, or similar).
- Advanced SQL and demonstrable data modeling judgment.
- Experience with columnar formats and open table formats: Parquet plus Iceberg, Delta, or Hudi.
- Track record of designing for reproducibility, including data versioning, lineage, and backfill correctness.
- Comfort operating without an existing platform to lean on and the judgment to sequence what gets built first.
Bonus
- Geospatial and raster data experience: GeoTIFF, cloud-optimized GeoTIFF, GDAL, tiling, and coordinate reference systems.
- Experience supporting computer vision or ML teams, including training dataset curation, annotation pipelines, or feature stores.
- Familiarity with MLflow, DVC, LakeFS, or comparable experiment and data versioning tooling.
- Distributed processing experience: Spark, Ray, or Dask.
- Analytics engineering exposure: dbt, dimensional modelling, BI tooling.
- Infrastructure as code: Terraform, CDK, or Pulumi.
- Kubernetes.
- Agriculture, remote sensing, robotics, or another domain with large sensor-derived datasets.
Education Requirements
- Bachelor's or master's degree in computer science, computer engineering, software engineering, and data science.