Home/Job List/Data Engineer Pyspark and Mongo DB
Aligned Automation

Data Engineer Pyspark and Mongo DB

Aligned Automation

India
Full-Time
Posted 1 month ago

Job Description & Responsibilities

Job Title: Data Engineer (MongoDB, PySpark & Python)

Experience

5–8 Years

Location

As per business requirement

Job Summary

We are looking for an experienced Data Engineer with strong expertise in MongoDB, PySpark, and Python to design, develop, and optimize scalable data pipelines. The ideal candidate should have experience working with large datasets, NoSQL databases, distributed data processing, and cloud-based data platforms.

Key Responsibilities

  • Design, build, and maintain scalable ETL/ELT data pipelines using PySpark and Python.
  • Develop data ingestion frameworks to process structured, semi-structured, and unstructured data.
  • Work extensively with MongoDB for data modeling, querying, indexing, aggregation, and performance optimization.
  • Optimize Spark jobs for high-performance processing of large datasets.
  • Build reusable data transformation and validation frameworks.
  • Develop REST API integrations and automate data ingestion using Python.
  • Monitor, troubleshoot, and optimize data pipelines for reliability and performance.
  • Collaborate with business analysts, data scientists, and application teams to deliver data solutions.
  • Implement data quality, governance, and security best practices.
  • Participate in code reviews and follow CI/CD and Agile development practices.

Required Skills

  • Strong experience in Python programming.
  • Hands-on experience with PySpark and Spark SQL.
  • Strong knowledge of MongoDB, including: • CRUD Operations
  • Aggregation Framework
  • Indexing
  • Replication
  • Sharding
  • Performance Tuning
  • Good understanding of data structures and algorithms.
  • Experience in developing ETL/ELT pipelines.
  • Strong SQL skills.
  • Experience with Git version control.
  • Knowledge of Linux/Unix commands.
  • Experience working with JSON, XML, and Parquet data formats.

Preferred Skills

  • Experience with Databricks.
  • Experience with cloud platforms such as Azure, AWS, or GCP.
  • Knowledge of Apache Kafka or other streaming technologies.
  • Experience with orchestration tools such as Apache Airflow.
  • Understanding of Delta Lake and Lakehouse architecture.
  • Familiarity with CI/CD pipelines.

Required Skills

PythonAWSAzureGCPSQLMongoDBKafkaGitCI/CDAgile

Job Details

Employment TypeFull-Time
Work ModeOn-Site
Experience58 years
Positions1

Posted by

N/A

Posted on:

24 Jul 2026

About Aligned Automation

More open roles

Browse all jobs →