About Lasso Informatics
Research data is wild. It pours in from brain scans, genome sequencing, EEG and eye-tracking, biospecimens and wearables that track sleep and movement, across dozens of sites and years of study. Lasso Informatics ropes it in: our platform pulls it all into one place so researchers around the world can stop fighting with their data and get to discoveries faster.
That is where our name comes from. A lasso brings something wild under control, and that is the job. The people who do it are Wranglers, and we are quite choosy about who earns the title. Wranglers expect a lot of themselves, make everyone around them better, and never stop learning. We move fast, we have fun, and we keep the bar high.
What We Offer
- A flexible, hybrid or remote position
- Competitive compensation and benefits
- Work that matters, with real world scientific and clinical impact
- Variety that keeps it interesting: no two studies, pipelines, or releases look the same
- Room to grow with a scaling mission-driven company at the forefront of research data management
- A sharp, multidisciplinary team that is good at what it does and good to work with.
Remote/Hybrid (Remote: US, Hybrid: Minneapolis / Remote: Canada, Hybrid: Montreal)
Full-Time
Canada Range: $74,000 - $95,000 CAD
US Range: $81,000 - $105,000 USD
Want to Be a Wrangler? Here's What We Look For
- You hold yourself to a high standard and bring real technical rigor to your work
- You make the people around you better
- You stay curious and keep learning
- You see a problem and start solving it, without waiting to be asked
- You keep the whole pipeline in view and notice how a change in one place ripples through the data, the release, and the team
- You document your work clearly enough that someone else can pick it up
- You communicate well with data managers, stewards, and subject matter experts who don't share your technical background
- You take data quality and responsible stewardship of research data seriously
About the Role
This is a hands-on data engineering role working with our Study Operations team. The mission is simple to state and hard to do: make study data clean, well-documented, and release-ready, and run quality control and large-scale data sharing across the studies and consortia we support.
You take the pipelines that internal and external subject matter experts build for imaging, biospecimens, wearables, EEG, and other multi-modal study data, and turn them into standardized, production-ready workflows. You configure the containers that package the scientific software those pipelines run on, execute them efficiently across high-performance and cloud computing environments, and move data across the distributed storage platforms our studies use. You work closely with data managers, data stewards, the workgroups behind each study, and the internal and external pipeline developers whose work you standardize and scale.
What You'll Do
Data Engineering
- Take pipelines built by internal and external subject matter experts and standardize them into scalable, production-ready workflows
- Optimize pipeline performance and scalability as data volumes and study complexity grow
- Maintain and continuously improve pipelines once standardized
- Monitor the state of study data and report progress toward upcoming releases
- Coordinate with data managers and data stewards on release timelines and data quality
Process Automation & Technical Leadership
- Investigate agentic AI-assisted approaches to pipeline and process management
- Advise Scientific Data Operations teams on their data processing needs
- Contribute engineering effort to data processing needs on other Lasso-supported studies
Data Quality
- Work with subject matter experts and external study partners to resolve data errors and discrepancies
Qualifications
Required
- Bachelor's degree in Computer Science, Data Science, Neuroscience, Bioinformatics, or a related field plus 4+ years of relevant experience, or a master's degree in one of these fields plus 2+ years of experience
- High comfort level with Linux and the command line
- Experience with containers: Docker, Kubernetes, Apptainer/Singularity, or similar
- Experience with High Performance Computing (HPC)
- Experience with cloud computing: AWS, Azure, OpenStack, or similar
- Programming experience in Python, R, or shell scripting
- Experience with data analysis, ETL, or reporting work
- Experience with Git
Preferred
- Experience with workflow-management tools such as Apache Airflow, Datalad, or SQL
- Familiarity with object stores such as AWS S3 or Ceph
- Experience working with academic or clinical research investigators across multi-site studies
- Experience working with NIfTI and DICOM medical imaging data
- Familiarity with neuroimaging analysis tools such as FreeSurfer, FSL, or Connectome Workbench
- Familiarity with machine learning, deep learning, and annotation workflows