Data Engineer – Bioinformatics / Multi-Omics Pipelines (1 Year+ Contract)
Worldwide
*** YOU MUST UPLOAD YOUR RESUME IN MS WORD FORMAT TO BE CONSIDERED*** Data Engineer – Bioinformatics / Multi-Omics Pipelines Duration: 12 months (extension likely) Hours: Mon through Fri, 8hrs/day, 5 days/week, 40hrs/week Partial overlap 4 to 6 hours with US Eastern Time required Location: Remote ABOUT THE ROLE We're looking for a Data Engineer with a bioinformatics or data science background to help build pipelines that bring large-scale clinical and multi-omics research data (EMR, genomics, imaging, biospecimens) into a cloud-based (AWS) data platform. You'll turn messy, multi-source biomedical data into clean, model-ready datasets used for research and analytics in the pharma/life sciences space. WHAT YOU'LL DO - Build and maintain data ingestion pipelines for large, multi-terabyte clinical and omics datasets - Harmonize longitudinal clinical data (EMR, treatment timelines, lab results) across multiple sites and coding systems into unified, analysis-ready tables - Develop or adapt bioinformatics pipelines for RNA-seq, single-cell RNA-seq, microbiome WGS, and proteomics data - Build feature-engineering workflows linking clinical and molecular data at the patient level - Ensure pipelines are reliable, scalable, and well-documented (QC, monitoring, error handling) - Collaborate with data scientists and researchers to translate scientific questions into data pipeline requirements MUST-HAVE SKILLS Data Engineering: - Strong Python (pandas, PySpark) and SQL - Experience building production ETL/ELT pipelines - Hands-on experience with AWS (or similar cloud data stack) - Experience harmonizing large, heterogeneous healthcare/clinical datasets Bioinformatics: - Solid background in RNA-seq and single-cell RNA-seq analysis - Experience with microbiome WGS (taxonomic/functional profiling) - Experience with proteomics data processing - Comfortable with QC, normalization, and feature extraction methods used in translational research - Experience with ingestion of the datasets including - clinical, omics, imaging derived features, including file format assessment and schema definition. NICE TO HAVE - Experience with large multimodal research cohorts or biobank data - HPC (high-performance computing) experience for bioinformatics workloads - Familiarity with Nextflow, Docker, or similar pipeline/workflow tools - Exposure to multi-omics integration methods (e.g., MOFA, deconvolution, single-cell analysis) - Understanding of clinical trial or real-world evidence (RWE) data structures - Comfortable working with both technical and non-technical stakeholders *** YOU MUST UPLOAD YOUR RESUME IN MS WORD FORMAT TO BE CONSIDERED***
- More than 30 hrs/weekHourly
- 6+ monthsDuration
- ExpertExperience Level
$15.00
-
$20.00
Hourly- Remote Job
- Complex projectProject Type
Skills and Expertise
Activity on this job
- Proposals:5 to 10
- Last viewed by client:3 days ago
- Interviewing:7
- Invites sent:33
- Unanswered invites:13
About the client
- United StatesSugar Land6:50 AM
- $1.6K total spent5 hires, 0 active
- 88 hours
Explore similar jobs on Upwork
How it works
Create your free profileHighlight your skills and experience, show your portfolio, and set your ideal pay rate.
Work the way you wantApply for jobs, create easy-to-by projects, or access exclusive opportunities that come to you.
Get paid securelyFrom contract to payment, we help you work safely and get paid securely.
About Upwork
- 4.9/5(Average rating of clients by professionals)
- G2 2021#1 freelance platform
- 49,000+Signed contract every week
- $2.3BFreelancers earned on Upwork in 2020
Find the best freelance jobs
Growing your career is as easy as creating a free profile and finding work like this that fits your skills.
Trusted by