Home/Job List/AI Data Engineer (LLM Training Data)
ZenteiQ.ai

AI Data Engineer (LLM Training Data)

ZenteiQ.ai

Bengaluru, Karnataka, India
Full-Time
Posted 11 days ago

Job Description & Responsibilities

As a Data Engineer on the corpus team, you'll build scalable pipelines to source, process, curate, filter and prepare large-scale corpora for training and evaluating foundation models. You'll work closely with AI Researchers, ML Engineers and Platform Engineers to solve challenging data problems across web-scale datasets, unstructured data, multilingual corpora and benchmark datasets.

Responsibilities

  • Source and ingest large-scale datasets from Common Crawl, public web sources, open-source repositories and other data sources.
  • Build scalable data processing pipelines for large unstructured corpora.
  • Develop and optimise distributed data pipelines using tools such as Apache Spark, Ray and Dask.
  • Extract and process content from web pages and documents using Trafilatura and related web/content extraction libraries.
  • Design data cleaning, filtering, quality assessment and curation pipelines for LLM pretraining corpora.
  • Develop heuristic-based filters to remove low-quality, boilerplate, spam, duplicated or otherwise unsuitable content.
  • Implement document-level and corpus-level deduplication strategies.
  • Identify and mitigate benchmark contamination and evaluation-data leakage in training corpora.
  • Prepare and curate benchmark and evaluation datasets.
  • Work with large-scale datasets stored and processed through GCS/GCP.
  • Monitor data quality, pipeline performance, storage and processing efficiency.
  • Collaborate with AI Researchers and ML Engineers to continuously improve corpus quality and training-data pipelines.

Requirements

  • 2-5 years of experience in Data Engineering, AI Data Engineering, ML Data Engineering or a related field.
  • Strong Python programming skills.
  • Experience working with large-scale unstructured datasets.
  • Experience building data processing / ETL pipelines.
  • Strong understanding of data cleaning, filtering, validation and deduplication.
  • Hands-on experience with exact and near-duplicate detection and deduplication.
  • Familiarity with at least one distributed processing framework: Apache Spark, Ray or Dask.
  • Understanding of web-scale data ingestion, such as Common Crawl or other large public datasets.
  • Experience with HTML parsing and web/document extraction libraries such as Trafilatura, BeautifulSoup, lxml or Resiliparse.
  • Familiarity with GCS/GCP and cloud-based data processing.
  • Experience working with large-scale formats such as JSONL, Parquet or Arrow.
  • Strong problem-solving and ownership mindset.

Good to Have

  • Experience with LLMs, Foundation Models or Generative AI.
  • Experience building or maintaining LLM pretraining corpora.
  • Familiarity with Hugging Face Datasets, DataTrove or NVIDIA NeMo Curator.
  • Experience with WARC/WET/WAT and large-scale Common Crawl processing.
  • Knowledge of MinHash, fuzzy or semantic deduplication.
  • Research-level understanding of synthetic data generation techniques for model training and evaluation.
  • Familiarity with Rust for high-performance data processing.
  • Experience with multilingual or domain-specific corpora.

Required Skills

PythonRustGCPMachine Learning

Job Details

Employment TypeFull-Time
Work ModeOn-Site
Experience25 years
Positions1

Posted by

N/A

Posted on:

18 Aug 2026

About ZenteiQ.ai

More open roles

Browse all jobs →