Home/Job List/AI20P Library Engineer, Machine Learning Acceleration

AI20P Library Engineer, Machine Learning Acceleration

Qpisemi

Bengaluru, Karnataka, India
Full-Time
Posted 1 month ago

Job Description & Responsibilities

AI20P Library Engineer, Machine Learning Acceleration

About the Role

We are looking for an experienced AI20P Library Engineer to bridge the gap between cutting-edge AI research and our hardware accelerators. In this role, you will design, develop, and optimize high-performance software libraries, kernels, and parallel compute runtimes for distributed AI20P environments. You will empower ML researchers and cloud developers to train and deploy frontier AI models at unprecedented scales.

Key Responsibilities

 Kernel Development & Optimization: Design, implement, and tune high- performance custom operators and mathematical kernels specifically for AI20P architectures (assembly or intrinsic level).

 Distributed Computing Runtimes: Build software abstractions and libraries that manage multi-host setups, sharding, and multi-dimensional parallelization (e.g., Megatron-LM style tensor/pipeline parallelism).

 Communication Primitives: Design and optimize custom collective communication algorithms (AllReduce, AllGather) to minimize latency and maximize throughput over high-speed interconnects.

 Performance Tuning: Profile distributed training workloads to identify and eliminate memory bottlenecks, network stalls, and suboptimal hardware utilization.

 Cross-Functional Collaboration: Partner with compiler developers (XLA), hardware architects, and research scientists to co-design the future software- hardware ecosystem.

Qualifications

 Education: B.S., M.S., or Ph.D. in Computer Science, Electrical Engineering, or a highly quantitative field.

 Programming: Exceptional proficiency in C++ and Python.

 ML Ecosystem: Hands-on experience with at least one major AI/ML framework (e.g., JAX, PyTorch).

 Systems Knowledge: Strong understanding of concurrent computations, memory hierarchies (HBM, cache), and hardware accelerators.

 Communication: Excellent cross-functional communication skills to translate complex research needs into production-ready software architecture.

Preferred Qualifications (Nice-to-Have)

 Experience working with compiler construction or optimizing code generation for hardware.

 Familiarity with cloud-based cluster managers (e.g., Kubernetes, SLURM) for executing large-scale distributed ML workloads.

 Prior contributions to optimizing Large Language Models (LLMs) or multimodal architectures for low-precision inference and training.

Required Skills

PythonKubernetesMachine LearningPyTorchCommunication

Job Details

Employment TypeFull-Time
Work ModeOn-Site
Experience00 years
Positions1

Posted by

N/A

Posted on:

15 Jul 2026

About Qpisemi

More open roles

Browse all jobs →