Home/Job List/Post-Training Optimization Engineer (LLM Inference & Efficiency)
Nava

Post-Training Optimization Engineer (LLM Inference & Efficiency)

Nava

Bengaluru, Karnataka, India
Full-Time
Posted 1 month ago

Job Description & Responsibilities

Role & Responsibilities

  • Optimize LLM inference pipelines for latency, throughput, and memory efficiency across GPU/TPU hardware—using quantization, pruning, kernel fusion, and runtime scheduling.
  • Implement and benchmark model compression techniques (INT4/INT8/FP8, LoRA, Sparse attention) to reduce model size while preserving performance.
  • Integrate and tune LLM serving frameworks (vLLM, TensorRT-LLM, HuggingFace TGI, ONNX Runtime) for high-throughput, low-latency inference in cloud and on-premise environments.
  • Collaborate with ML researchers and infrastructure engineers to profile bottlenecks and implement hardware-aware optimizations using CUDA, Triton, or custom kernels.
  • Design and automate model benchmarking suites to track performance regression, memory footprint, and cost-per-token across model versions and hardware targets.
  • Document optimization playbooks and contribute to internal tooling for repeatable, scalable model deployment workflows.

Skills & Qualifications

Must-Have

  • PyTorch
  • TensorRT
  • vLLM
  • Quantization (INT4/INT8/FP8)
  • CUDA
  • ONNX Runtime
  • Triton Inference Server
  • LLM Inference Optimization

Preferred

  • Experience with Mixture-of-Experts (MoE) models
  • Familiarity with HuggingFace Transformers and TGI
  • Knowledge of NPU/ASIC inference backends (e.g., Qualcomm, Groq, Cerebras)

Skills: cuda,ml,learning,compression,optimization,distillation,decoding,research,machine learning,training

Required Skills

Machine LearningPyTorch

Job Details

Employment TypeFull-Time
Work ModeOn-Site
Experience00 years
Positions1

Posted by

N/A

Posted on:

14 Jul 2026

About Nava

More open roles

Browse all jobs →