Home/Job List/Senior DevOps Engineer
Simplex Services

Senior DevOps Engineer

Simplex Services

New Delhi, Delhi, India
Full-Time
Posted 23 days ago

Job Description & Responsibilities

Job Title: DevOps Platform Engineer – Gen AI & Agentic AI

Contract Duration: 6 months (high possibility of extension)

Work Mode: Hybrid (2–3 days per week in the London office)

Location: London, UK

Start Date: ASAP

About the Role

We are looking for a DevOps Platform Engineer to design, size, and manage the cloud and infrastructure foundation for Generative AI and Agentic AI solutions, preferably within the Consumer Packaged Goods (CPG), Food & Beverage industry. This role owns the technical backbone that GenAI/Agentic applications run on — from infrastructure planning and BOM creation to GPU/compute sizing, LLMOps pipelines, and continuous optimization of AI workloads.

The ideal candidate understands both classic cloud/DevOps fundamentals and the unique infrastructure demands of LLMs, RAG pipelines, and multi-agent systems — and can independently translate a customer's Gen AI use case into a right-sized, production-ready, cost-optimized platform.

Key Responsibilities

Infrastructure & Requirements Assessment for Gen AI Workloads

  • Engage with client stakeholders, data scientists, and solution architects to understand Gen AI/Agentic AI use cases (copilots, RAG applications, autonomous agents, document intelligence, etc.) and their infrastructure implications.
  • Assess current-state client environments and identify gaps for hosting LLM-based and agentic applications (compute, GPU access, networking, data pipelines, security).
  • Translate model requirements (model size, context length, throughput/latency targets, concurrency) into concrete infrastructure specifications.

BOM & Cloud Architecture for AI Platforms

  • Prepare and maintain detailed Bill of Materials (BOM) covering compute (CPU/GPU), storage, networking, vector databases, orchestration tooling, and LLM API/licensing costs.
  • Design cloud architecture (AWS/Azure/GCP) for Gen AI workloads — including model hosting/inference endpoints, vector databases, RAG pipelines, agent orchestration layers, and API gateways.
  • Evaluate build-vs-buy decisions: managed LLM APIs (OpenAI, Anthropic, Azure OpenAI, Bedrock) vs. self-hosted/open-source models (Llama, Mistral, etc.) based on cost, data privacy, and performance needs.
  • Support proposal and pre-sales efforts with accurate sizing, GPU costing, and architecture inputs for Gen AI engagements.

Sizing & Capacity Planning for LLM/Agentic Workloads

  • Perform workload analysis and capacity planning specific to Gen AI systems — token throughput, concurrent users, embedding/indexing volumes, and agent execution loads.
  • Size GPU/compute infrastructure for model inference and (where applicable) fine-tuning, balancing latency, throughput, and cost.
  • Plan for vector database scale (embedding volume, query load) and retrieval pipeline performance for RAG-based solutions.
  • Account for CPG-specific patterns — seasonal spikes (promotions, demand planning cycles), batch document/data processing volumes, and multi-brand/multi-market scaling needs.

Workload & Cost Optimization

  • Monitor Gen AI infrastructure spend and performance — GPU utilization, API token consumption, inference latency, and vector DB query costs.
  • Implement autoscaling and dynamic resource allocation for inference endpoints and agent workloads to manage cost-to-performance trade-offs.
  • Apply FinOps practices tailored to AI workloads — cost attribution by use case/agent, model routing to lower-cost models where appropriate, caching, and prompt/token optimization strategies.
  • Identify opportunities to right-size GPU instances, batch inference jobs, and reduce idle compute costs.

LLMOps / Platform Engineering & Operations

  • Build and maintain infrastructure-as-code (Terraform, CloudFormation, ARM/Bicep) for repeatable provisioning of AI platform components.
  • Set up CI/CD pipelines for Gen AI applications, including model/prompt versioning, evaluation gates, and safe rollout of agentic workflows.
  • Deploy and manage container orchestration (Kubernetes/Docker) for model serving, RAG pipelines, and agent runtimes.
  • Implement monitoring, logging, and observability for AI-specific metrics (latency, token usage, hallucination/error rates, agent task success rates) using tools like Prometheus, Grafana, Datadog, or LLM-specific observability platforms (e.g., LangSmith, Arize).
  • Ensure infrastructure security, data privacy, and compliance for AI systems handling sensitive or proprietary client data.

Client & Stakeholder Engagement

  • Act as the technical point of contact for infrastructure and cloud discussions on Gen AI/Agentic AI engagements.
  • Present sizing, BOM, and architecture recommendations for AI platforms in clear, business-friendly terms to technical and non-technical stakeholders.
  • Collaborate closely with Gen AI solution leads, data scientists, and application teams to ensure infrastructure choices align with use case goals and budget.

Pay: ₹6,483.30 - ₹20,952.10 per month

Work Location: Hybrid remote in Delhi, Delhi (Delhi)

Required Skills

AWSAzureGCPDockerKubernetesCI/CDTerraform

Job Details

Employment TypeFull-Time
Work ModeHybrid
Experience0 – 0 years
Positions1

Posted by

N/A

Posted on:

16 Sept 2026

About Simplex Services

More open roles

Browse all jobs →