Home/Job List/DevOps Specialist Engineer – SRE, Cloud & Applied AI
Clarus Advisers

DevOps Specialist Engineer – SRE, Cloud & Applied AI

Clarus Advisers

India
Full-Time
Posted 22 days ago

Job Description & Responsibilities

Company Overview

Our client is a leading technology organization focused on building scalable, cloud-native platforms and intelligent software solutions. The organization combines software engineering, cloud technologies, Site Reliability Engineering, and Applied AI to deliver highly resilient and production-ready products at scale.

Position Overview

We are looking for a DevOps Specialist Engineer with strong experience in Site Reliability Engineering, cloud platform engineering, software development, and production operations. The ideal candidate will have hands-on experience operating large-scale distributed systems across Azure, AWS, or GCP, along with exposure to AI/ML, GenAI, LLMOps/MLOps, observability, performance engineering, and cloud cost optimization. The role requires a strong engineering mindset and the ability to build reliable, secure, scalable, and highly automated production platforms.

Responsibilities

  • Design, build, operate, and continuously improve large-scale cloud-native production systems.
  • Define and own SLIs, SLOs, SLAs, error budgets, and reliability objectives.
  • Lead production incident response, on-call operations, root-cause analysis, and reliability improvements.
  • Build and manage CI/CD pipelines, Kubernetes platforms, infrastructure automation, and multi-environment deployments.
  • Implement Infrastructure as Code using Terraform and deployment automation using tools such as ArgoCD.
  • Develop production-grade observability across metrics, logging, and distributed tracing.
  • Work with technologies such as Prometheus, Grafana, OpenTelemetry, Datadog, Dynatrace, Splunk, Azure Monitor, and AWS CloudWatch.
  • Operate AI/ML and GenAI workloads in production, addressing reliability, performance, model drift, output variance, and train/serve skew.
  • Support MLOps/LLMOps platforms and AI control-plane capabilities such as model gateways and guardrails.
  • Implement Kubernetes/Docker-based solutions and cloud-native networking across multiple environments.
  • Conduct load and performance testing using tools such as LoadRunner, k6, or JMeter.
  • Implement chaos engineering, capacity planning, autoscaling, and resilience testing.
  • Drive cloud and AI FinOps, including GPU, inference, and token-cost attribution and optimization.
  • Implement security controls including RBAC, least privilege, secrets management, deployment approvals, and segregation of duties.
  • Collaborate with software engineering, security, data/AI, and product teams to improve platform reliability and operational excellence.

Skills & Experience

  • 6–9 years of experience in DevOps, SRE, Site Reliability Engineering, Platform Engineering, or Software Engineering.
  • Bachelor's degree in Computer Science, Software Engineering, Data Science, Machine Learning, or a related discipline.
  • Strong programming experience in one or more of Python, C#/.NET, Go, Java, or Bash.
  • Strong hands-on experience with Azure, AWS, or GCP; Azure/AWS preferred.
  • Experience with Kubernetes, Docker, Terraform, and cloud-native architectures.
  • Strong understanding of CI/CD, GitHub, Azure DevOps (ADO), and ArgoCD.
  • Experience with production observability and monitoring using tools such as Prometheus, Grafana, OpenTelemetry, Datadog, Dynatrace, Splunk, CloudWatch, or Azure Monitor.
  • Strong understanding of SRE principles, SLIs, SLOs, SLAs, error budgets, incident management, and production on-call operations.
  • 3+ years of experience operating or supporting large-scale production systems.
  • Experience with AI/ML or GenAI workloads in production and familiarity with Azure OpenAI, AWS Bedrock, or Vertex AI.
  • Knowledge of MLOps/LLMOps, MLflow, LangFuse, LangSmith, or equivalent AI/agent orchestration platforms.
  • Experience with load/performance testing, capacity planning, autoscaling, and chaos engineering.
  • Understanding of cloud cost optimization/FinOps, including AI/GPU/inference and token-cost management.
  • Knowledge of DevSecOps, RBAC, secrets management, security controls, and environment integrity.
  • Strong understanding of OOP/OOD, data structures, algorithms, code instrumentation, and software engineering practices.
  • Ability to understand and work with business context, sequence, activity, state, entity-relationship, and data-flow diagrams.
  • Exposure to XP, Lean, SRE, and AI-augmented/spec-driven software development is an advantage.

Required Skills

PythonJavaC#GoAWSAzureGCPDockerKubernetesCI/CD

Job Details

Employment TypeFull-Time
Work ModeOn-Site
Experience69 years
Positions1

Posted by

N/A

Posted on:

27 Aug 2026

About Clarus Advisers

More open roles

Browse all jobs →