Home/Job List/Senior Kubernetes Engineer (SME)
Neysa

Senior Kubernetes Engineer (SME)

Neysa

India
Full-Time
Posted 10 days ago

Job Description & Responsibilities

About the Role

We are building and running mission-critical production infrastructure on Kubernetes. As a Senior Kubernetes Engineer, you will own the full stack - from the underlying OS and container runtime through networking, storage, and the cluster control plane itself. You will engage across the full lifecycle: architecture, deployment, hardening, and day-2 operations for multi-cluster, multi-tenant environments. This is hands-on infrastructure work with direct ownership of production reliability.

What you will be doing

  • Design, deploy, and operate production-grade Kubernetes clusters across bare-metal, cloud, and hybrid environments.
  • Manage cluster lifecycle end-to-end: provisioning, version upgrades, patching, scaling, and capacity planning
  • Architect and run multi-cluster/multi-tenant setups using Kamaji, Rancher (hosted control planes), and vCluster (virtual clusters)
  • Configure and troubleshoot CNI plugins (Calico, Cilium) — pod networking, network policies, BGP, and eBPF dataplanes.
  • Own cluster DNS (CoreDNS) configuration, service discovery, and resolution troubleshooting.
  • Manage container runtime (containerd, CRI-O) and underlying OS: Linux tuning, kernel/sysctl parameters, systemd, cgroups
  • Deploy and operate service mesh (Istio, Cilium mesh) for traffic management
  • Manage persistent storage: PV/PVC, StorageClasses, CSI drivers (Rook/Ceph, Longhorn, cloud-native CSI)
  • Own the network stack: ingress controllers, load balancing, MetalLB/BGP, firewalling, and network troubleshooting
  • Implement GitOps and Infrastructure-as-Code (ArgoCD/FluxCD, Terraform, Helm) for cluster and workload delivery
  • Harden clusters: RBAC, Pod Security Standards, network policies, secrets management (Vault, Sealed Secrets), image scanning
  • Own backup and disaster recovery (Velero, etcd snapshotting) and run DR drills
  • Provide on-call production support: monitor cluster health, troubleshoot incidents, and drive root-cause resolution
  • Create comprehensive documentation, runbooks, and knowledge bases for operational continuity and knowledge transfer

What we need to see

  • Core Kubernetes & Infrastructure (5+ years)
  • Deep expertise in Kubernetes architecture: control plane, etcd, kube-apiserver, scheduler, controller-manager, kubelet
  • Proven experience designing, deploying, and troubleshooting production clusters at scale
  • Hands-on with multi-cluster/multi-tenant tooling — Rancher, Kamaji, and/or vCluster (strongly preferred)
  • CNI expertise: Calico, Cilium — network policy design, BGP, VXLAN/IPIP encapsulation, eBPF
  • Container runtime internals: containerd, CRI-O, runc — configuration and troubleshooting
  • Strong Linux systems administration: kernel tuning, systemd, cgroups/namespaces, sysctl, package/OS lifecycle management
  • Storage: PV/PVC, StorageClasses, CSI drivers, Ceph/Rook, Longhorn, NFS
  • Service mesh experience: Istio, Linkerd, or Cilium service mesh
  • CoreDNS configuration, custom resolvers, and DNS troubleshooting in cluster environments
  • Networking depth: ingress controllers (NGINX, Envoy, Traefik), load balancing, MetalLB, and diagnostic tooling (tcpdump, iptables/nftables, conntrack)

Automation & Infrastructure-as-Code

  • Helm chart authoring and lifecycle management
  • GitOps workflows: ArgoCD or FluxCD
  • IaC and configuration management: Terraform, Ansible
  • CI/CD pipeline integration for cluster and application delivery
  • Scripting proficiency: Bash and Python (Go a plus)

Ways to stand out from the rest

  • Kubernetes certifications: CKA, CKAD, CKS
  • Production experience with Rancher, Kamaji, and vCluster together (fleet/hosted-control-plane management)
  • Multi-cloud Kubernetes: EKS, AKS, GKE, and bare-metal
  • Experience with GPU-enabled clusters and AI/ML workload scheduling '
  • Familiarity with AI/LLM serving stacks on Kubernetes — vLLM, KServe, Triton Inference Server, GPU operator/device plugin, MIG partitioning
  • General AI infrastructure knowledge: model serving patterns, inference autoscaling, GPU scheduling constraints
  • Contributions to CNCF projects or an active open-source/GitHub presence

Minimum Qualifications

  • Bachelor’s degree in computer science, Electrical/Computer Engineering, or related field (or equivalent industry experience)
  • 5+ years of hands-on production Kubernetes experience
  • Demonstrated ownership of CNI, DNS, storage, and networking within Kubernetes environments
  • Solid grounding in containerd/OS-level troubleshooting
  • Production on-call and incident-response experience.

Soft Skills

  • Strong problem-solving and debugging abilities under production pressure
  • Ownership mindset with accountability for platform reliability and stability
  • Cross-functional collaboration with application, security, operations and platform teams
  • Proactive approach to continuous learning and staying current with the Kubernetes/CNCF ecosystem
  • Strong documentation and communication skills, with ability to defend design decisions in peer reviews

Required Skills

PythonGoKubernetesCI/CDTerraformMachine LearningCommunication

Job Details

Employment TypeFull-Time
Work ModeOn-Site
Experience510 years
Positions1

Posted by

N/A

Posted on:

19 Aug 2026

About Neysa

More open roles

Browse all jobs →