AI Data Engineer (LLM Training Data)
ZenteiQ.ai
Bengaluru, Karnataka, India
Full-Time
Posted 11 days ago
Bengaluru, Karnataka, India
Full-Time
Posted 11 days ago
Job Description & Responsibilities
As a Data Engineer on the corpus team, you'll build scalable pipelines to source, process, curate, filter and prepare large-scale corpora for training and evaluating foundation models. You'll work closely with AI Researchers, ML Engineers and Platform Engineers to solve challenging data problems across web-scale datasets, unstructured data, multilingual corpora and benchmark datasets.
Responsibilities
- Source and ingest large-scale datasets from Common Crawl, public web sources, open-source repositories and other data sources.
- Build scalable data processing pipelines for large unstructured corpora.
- Develop and optimise distributed data pipelines using tools such as Apache Spark, Ray and Dask.
- Extract and process content from web pages and documents using Trafilatura and related web/content extraction libraries.
- Design data cleaning, filtering, quality assessment and curation pipelines for LLM pretraining corpora.
- Develop heuristic-based filters to remove low-quality, boilerplate, spam, duplicated or otherwise unsuitable content.
- Implement document-level and corpus-level deduplication strategies.
- Identify and mitigate benchmark contamination and evaluation-data leakage in training corpora.
- Prepare and curate benchmark and evaluation datasets.
- Work with large-scale datasets stored and processed through GCS/GCP.
- Monitor data quality, pipeline performance, storage and processing efficiency.
- Collaborate with AI Researchers and ML Engineers to continuously improve corpus quality and training-data pipelines.
Requirements
- 2-5 years of experience in Data Engineering, AI Data Engineering, ML Data Engineering or a related field.
- Strong Python programming skills.
- Experience working with large-scale unstructured datasets.
- Experience building data processing / ETL pipelines.
- Strong understanding of data cleaning, filtering, validation and deduplication.
- Hands-on experience with exact and near-duplicate detection and deduplication.
- Familiarity with at least one distributed processing framework: Apache Spark, Ray or Dask.
- Understanding of web-scale data ingestion, such as Common Crawl or other large public datasets.
- Experience with HTML parsing and web/document extraction libraries such as Trafilatura, BeautifulSoup, lxml or Resiliparse.
- Familiarity with GCS/GCP and cloud-based data processing.
- Experience working with large-scale formats such as JSONL, Parquet or Arrow.
- Strong problem-solving and ownership mindset.
Good to Have
- Experience with LLMs, Foundation Models or Generative AI.
- Experience building or maintaining LLM pretraining corpora.
- Familiarity with Hugging Face Datasets, DataTrove or NVIDIA NeMo Curator.
- Experience with WARC/WET/WAT and large-scale Common Crawl processing.
- Knowledge of MinHash, fuzzy or semantic deduplication.
- Research-level understanding of synthetic data generation techniques for model training and evaluation.
- Familiarity with Rust for high-performance data processing.
- Experience with multilingual or domain-specific corpora.
About ZenteiQ.ai
Required Skills
PythonRustGCPMachine Learning
Job Details
Employment TypeFull-Time
Work ModeOn-Site
Experience2 – 5 years
Positions1
Posted by
N/A
Posted on:
18 Aug 2026
About ZenteiQ.ai
More open roles
- Backend Developer-Python Django-Immediate joinersPhigital Care · Secunderabad, Telangana, India
- Python Backend Developer - AI AutomationWorkfall India · India
- Python (Full stack & Backend) Engineer - Multiple PositionsColaberry · Hyderabad, Telangana, India
- Lead Python DeveloperiCapital · Jaipur, Rajasthan, India
- Python (Full stack & Backend) Engineer - Multiple PositionsColaberry · Hyderabad, Telangana, India
- Lead Python DeveloperiCapital · Jaipur, Rajasthan, India