Home/Job List/Senior Engineer, MLOps
Code and Theory

Senior Engineer, MLOps

Code and Theory

India
Full-Time
Posted 1 month ago

Job Description & Responsibilities

Senior Engineer MLOps

Location: Bangalore (Hybrid)

Experience: 6–10 Years

Job Summary

We are looking for a highly skilled Senior Engineer – MLOps to build, deploy, and operate scalable, reliable AI/ML infrastructure that powers our next-generation AI platforms. This role sits at the intersection of Machine Learning, DevOps, and Cloud Engineering, with a strong focus on supporting LLM agent systems, data pipelines, cloud infrastructure, and production AI workloads.

The ideal candidate will have hands-on experience deploying and managing machine learning systems in cloud environments, automating infrastructure, building CI/CD pipelines, and ensuring the performance, scalability, and reliability of AI/ML platforms.

Key Responsibilities

  • Design, deploy, and manage scalable AI/ML infrastructure for production environments.
  • Build and maintain infrastructure supporting LLM agent systems, machine learning models, and data pipelines.
  • Deploy and operate ML/AI workloads across major cloud platforms such as AWS, GCP, or Azure.
  • Develop and maintain containerized applications using Docker and serverless container platforms such as Cloud Run, ECS Fargate, or Azure Container Apps.
  • Design and manage cloud databases, data warehouses, and storage solutions for AI applications.
  • Build and optimize ETL/ELT pipelines to support machine learning workflows and analytics.
  • Implement Infrastructure as Code (IaC) using Terraform or similar tools.
  • Design and maintain CI/CD pipelines for automated model deployment and infrastructure provisioning.
  • Implement monitoring, logging, alerting, and performance optimization for AI/ML systems.
  • Manage event-driven architectures using messaging platforms such as Kafka, Pub/Sub, SNS/SQS, or similar technologies.
  • Collaborate closely with AI/ML Engineers, Data Scientists, and Software Engineers to support model deployment and production operations.
  • Optimize infrastructure for reliability, scalability, latency, and cost efficiency.
  • Ensure security, governance, and operational best practices across cloud environments.

Required Qualifications

  • Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field.
  • 6–10 years of experience in MLOps, DevOps, Cloud Engineering, or Infrastructure Engineering.
  • Proven experience deploying and operating Machine Learning or AI systems in production.
  • Strong expertise in cloud platforms including Google Cloud Platform (GCP), Amazon Web Services (AWS), or Microsoft Azure.
  • Experience with Infrastructure as Code (Terraform preferred).
  • Strong Python programming skills.
  • Excellent analytical, troubleshooting, and problem-solving skills.
  • Strong communication skills with the ability to collaborate across technical and non-technical teams.

Required Technical Skills

  • Expert-level Python or Typescript
  • Very deep experience with at least one major cloud platform

Including microservice creation and maintenance

Including databases and warehousing

Including messaging/streaming infrastructure

  • Strong experience with containerization
  • Strong experience with building-out telemetry and monitoring platforms
  • Strong experience with ML model serving
  • Expertise with infrastructure as code (IaC) such as Terraform
  • Experience building and maintaining CI/CD pipelines
  • Experience both implementing and advocating for MLOps best practices, model lifecycle management, and AI platform operations.
  • Knowledge of:

Vector Databases

Agent Orchestration

Cost Optimization

Preferred Skills

  • Experience with LLM Agent Frameworks such as Google ADK, LangChain, LangGraph, or similar.
  • Experience operating LLM/Generative AI workloads in production.

Preferred Candidate Profile

  • Experience supporting enterprise-scale AI/ML platforms.
  • Passion for building reliable, scalable, and automated infrastructure.
  • Ability to troubleshoot complex distributed systems.
  • Strong collaboration and stakeholder management skills.
  • Self-driven with a continuous learning mindset and enthusiasm for emerging AI technologies.

Required Skills

TypeScriptPythonAWSAzureGCPDockerKafkaCI/CDTerraformMachine Learning

Job Details

Employment TypeFull-Time
Work ModeRemote
Experience610 years
Positions1

Posted by

N/A

Posted on:

11 Jul 2026

About Code and Theory

More open roles

Browse all jobs →