Voice AI Foundation Model Engineer
TalixoHR
Job Description & Responsibilities
We are looking for a hands-on
Voice AI / Machine Learning Engineer
to work on Voice AI initiatives, with a strong focus on
training and fine-tuning foundation models
. The role will involve building speech and conversational AI systems across ASR, TTS, speaker identification, voice activity detection, and related use cases. The ideal candidate will bring
3–4 years of relevant domain experience
and practical experience training or pre-training foundation models.
Key Responsibilities
- Develop, train, fine-tune, and evaluate foundation models for Voice AI use cases.
- Build and improve solutions involving
ASR, TTS, speaker identification, voice activity detection, and conversational AI
.
- Prepare, clean, optimize, and manage large-scale speech and text datasets.
- Design scalable model-training pipelines and conduct structured model experiments.
- Improve model
accuracy, latency, robustness, and inference efficiency
.
- Work with GPU-based workloads and distributed training environments.
- Evaluate models and apply optimization techniques for production readiness.
- Collaborate across ML/AI workflows to translate Voice AI requirements into deployable solutions.
Must-Have Requirements
- 3–4 years of hands-on experience
in Machine Learning, Speech AI, NLP, or a closely related field.
- Demonstrated experience
training or pre-training foundation models
.
- Strong understanding of deep learning architectures, particularly
Transformers
.
- Practical experience working with
speech technologies such as ASR and TTS models
.
- Strong proficiency in
Python
and ML frameworks such as
PyTorch or TensorFlow
.
- Experience with
distributed training, GPU workloads, and large-scale data pipelines
.
- Familiarity with model evaluation, optimization, quantization, and production deployment.
- Strong analytical, debugging, and problem-solving capabilities.
Preferred Qualifications
- Experience building
multilingual or Indian-language speech models
.
- Familiarity with
Whisper, wav2vec 2.0, HuBERT, NeMo, SpeechBrain, or Hugging Face
.
- Experience with audio preprocessing, augmentation, annotation, and dataset quality improvement.
- Experience working on large-scale Voice AI or conversational AI systems.
What Success Looks Like
- Successfully trains, fine-tunes, and evaluates foundation models for Voice AI applications.
- Delivers measurable improvements in model accuracy, latency, robustness, or inference efficiency.
- Builds reliable training pipelines and effectively manages large-scale speech/text datasets.
- Develops production-relevant ASR/TTS and related Voice AI capabilities.
- Runs structured model experiments and translates findings into model improvements.
Experience
3–4 years
of relevant hands-on experience in Machine Learning, Speech AI, NLP, or a closely related field.
Why Join
- Opportunity to work on
Voice AI initiatives
involving foundation-model development.
- Hands-on ownership across speech AI, model training, data pipelines, and model optimization.
- Exposure to advanced speech technologies including ASR, TTS, and conversational AI.
- Opportunity to work with large-scale datasets, GPU workloads, and distributed training.
About TalixoHR
Required Skills
Job Details
Posted by
N/A
Posted on:
28 Sept 2026
About TalixoHR
More open roles
- Python Backend Developer - Django/FastAPICustomertimes · India
- Technology Trainers – Python Full Stack & GenAIShazVerse Technologies · Hyderabad, Telangana, India
- Python DevOps EngineerUST · Bengaluru, Karnataka, India
- SENIOR PYTHON ENGINEERV2 Retail Ltd · Gurugram, Haryana, India
- SENIOR PYTHON ENGINEERV2 Retail Ltd · Gurugram, Haryana, India
- SENIOR PYTHON ENGINEERV2 Retail Ltd · Gurugram, Haryana, India