Home/Job List/Senior MLOps & DevOps Engineer

Senior MLOps & DevOps Engineer

Bellcom Technologies

Bhopal, Madhya Pradesh, India
Full-Time
Posted 11 days ago

Job Description & Responsibilities

Job Description

Senior MLOps & DevOps Engineer

LocationBhopal, MP (On-site)Experience610 YearsReports ToLead AI EngineerOpenings1Role Overview

Own and operate the end-to-end MLOps platform for enterprise AI/ML solutions on secure on-premise and air-gapped infrastructure. This is not a traditional DevOps role hands-on production experience with GPU clusters (H100/A100), Kubernetes, LLM inference platforms (vLLM/Triton), and air-gapped deployments is mandatory. The ideal candidate can stand up the full MLOps stack from bare metal without internet access and has operated AI/ML systems in defence, government, or classified environments.

Key Responsibilities

  • Design and operate the full ML pipeline: ingestion training validation deployment monitoring drift detection automated retraining.
  • Build and maintain ML pipelines with MLflow, Kubeflow, and/or Airflow; manage experiment tracking, model registry, and versioning.
  • Deploy and manage LLM/ML inference with vLLM, Triton, or TGI on local GPU hardware.
  • Administer Kubernetes on bare-metal on-premise infrastructure; manage NVIDIA GPU Operator, Helm, RBAC, and storage integration.
  • Configure and maintain NVIDIA GPU infrastructure: CUDA, cuDNN, TensorRT, NCCL, GPU drivers, and multi-GPU scheduling.
  • Build CI/CD pipelines for model code; automate testing, validation, and deployment via IaC (Terraform, Ansible, Helm).
  • Set up model monitoring dashboards: accuracy tracking, data drift detection, and performance alerting (Prometheus, Grafana).
  • Manage air-gapped environments: offline package repos, private container registries, RBAC, audit logging, and security controls.
  • Implement HA, disaster recovery, backup, and incident management for the ML platform.

Required Skills & Experience

  • 610 years in DevOps/MLOps; 3+ yrs DevOps, 2+ yrs dedicated MLOps with production GPU infrastructure.
  • Hands-on NVIDIA GPU ops: H100, H200, A100, L40S or RTX Ada; CUDA, cuDNN, TensorRT, NCCL, multi-GPU clusters.
  • Production experience: MLflow, Kubeflow/Airflow, Docker, Kubernetes (bare metal), CI/CD, and IaC (Terraform/Ansible/Helm).
  • LLM inference platforms: vLLM, NVIDIA Triton Inference Server, Ollama, or TGI on local GPU hardware.
  • Proven experience operating AI/ML in air-gapped, classified, defence, or government environments with offline registries and repos.
  • Ability to provision the complete local ML stack from bare metal (GPU K8s registry inference monitoring) without internet access.
  • Python automation, Linux administration, FastAPI/Flask model serving, Prometheus/Grafana monitoring.
  • Security controls: RBAC, secrets management, network policies, audit logging on on-premise infrastructure.

Preferred / Good to Have

  • DGX-class systems (DGX A100/H100); MinIO/Ceph distributed object storage; DVC for data versioning.
  • HA/DR configuration for ML serving; RPA platform integration; model governance frameworks.

Qualifications

  • B.Tech / M.Tech in Computer Science, Software Engineering, or related field.
  • Kubernetes (CKA) or MLOps certifications are a plus.

Required Skills

PythonDockerKubernetesCI/CDTerraformMachine Learning

Job Details

Employment TypeFull-Time
Work ModeOn-Site
Experience1015 years
Positions1

Posted by

N/A

Posted on:

18 Aug 2026

About Bellcom Technologies

More open roles

Browse all jobs →