Home/Job List/Bioinformatics Data Engineer ( Python / Shell Scripting / Docker )
Sanskruti Solutions

Bioinformatics Data Engineer ( Python / Shell Scripting / Docker )

Sanskruti Solutions

India
Full-Time
Posted 1 month ago

Job Description & Responsibilities

  • (Important Notes: (1) It is a Desk Job (Onsite Job), operating for 6 days / week working pattern
  • 5 days working from an office environment and Saturday currently allowed with remote working and (2) Immediate Joiners Preferred)

Job Summary

  • We are looking for a highly skilled Data Engineer with strong programming and workflow automation skills to support and enhance our production bioinformatics pipelines.
  • You do not need a biology background
  • but you must be comfortable working with biological data formats, large datasets, and cloud-based workflows.
  • In this role, you will develop, maintain, and optimize our data processing pipelines used for NGS (Next-Generation Sequencing) analysis, clinical workflows, and research innovation.
  • You will work closely with bioinformaticians, software engineers, and data scientists to build scalable, reliable, and efficient systems.

Responsibilities

  • Pipeline & Data Engineering:
  • Develop and maintain scalable data pipelines for genomic and clinical datasets.
  • Build workflow automation using Python, Shell, Docker, and workflow managers (Nextflow, Snakemake, Airflow, etc.).
  • Optimize existing pipelines for performance, resource usage, and reliability.
  • Handle large biological datasets (FASTQ, BAM, VCF, CSV/TSV, metadata).
  • Software Engineering:
  • Write clean, modular, production-level code in Python and Shell.
  • Implement CI/CD processes for pipeline deployment.
  • Maintain code repositories (Git) and ensure high-quality documentation.
  • Cloud & Infrastructure:
  • Work with AWS/GCP/Azure services for scalable pipeline execution.
  • Manage container-based deployments using Docker.
  • Monitor job performance, logs, and system behavior.
  • Data Management:
  • Maintain data integrity, versioning, and audit trails.
  • Develop automated QC checks and validation workflows.
  • Support data ingestion, transformation, and ETL processes.
  • Collaboration:
  • Work with bioinformaticians to translate analysis logic into scalable workflows.
  • Collaborate with clinical and operations teams to ensure pipeline readiness.
  • Troubleshoot pipeline failures and optimize workflows in production.

Required Skills: Core Technical Skills

  • Strong proficiency in Python
  • Hands-on experience with Shell scripting (bash)
  • Strong understanding of Docker / containerization
  • Knowledge of Git, CI/CD, and software development best practices
  • Experience with workflow orchestration: Nextflow, Snakemake, Airflow, Cromwell, Prefect (any one) Data Engineering Skills:
  • Experience with large datasets, ETL pipelines, log processing
  • Strong understanding of file formats, data transformation, and automation
  • Familiarity with Linux and HPC or distributed systems Bioinformatics Data Handling (Training can be provided):
  • Understanding of at least one biological data type: FASTQ, BAM/CRAM, VCF, BED, GTF
  • Ability to process unstructured or semi-structured scientific data Good to Have Skills (Will Prefer If Available):
  • Experience with R
  • Exposure to ML/AI workflows
  • Experience with cloud (AWS Batch, Lambda, S3, EC2)
  • Knowledge of Next Generation Sequencing (NGS) pipelines
  • Familiarity with database systems (SQL/NoSQL)

Required Skills

PythonAWSAzureGCPDockerSQLGitCI/CDMachine Learning

Job Details

Employment TypeFull-Time
Work ModeRemote
Experience00 years
Positions1

Posted by

N/A

Posted on:

25 Jul 2026

About Sanskruti Solutions

More open roles

Browse all jobs →