Machine Learning Engineer - ASR & Speech Recognition
Ethical Den
Job Description & Responsibilities
JOB TITLE: Machine Learning Engineer - ASR & Speech Recognition
Company: Ethical Den
Location: Kolkata, West Bengal
Work Mode: On-site only
Employment Type: Full-time
Experience: 2 to 5+ years
Openings: 1
Joining: Immediate to 30 days preferred
ABOUT THE ROLE
Ethical Den is hiring a Machine Learning Engineer specialising in Automatic Speech Recognition and real-time speech processing.
You will work with our AI research team on a confidential speech-technology initiative.
The role involves speech recognition, streaming inference, endpoint detection, speech datasets and optimisation for real-world audio conditions.
Detailed project information will be shared with suitable candidates subject to confidentiality requirements.
KEY RESPONSIBILITIES
- Research and evaluate ASR architectures.
- Fine-tune speech-recognition models.
- Build streaming ASR pipelines.
- Generate and improve partial transcripts.
- Optimise recognition for noisy and real-world audio.
- Develop Voice Activity Detection systems.
- Work on endpoint detection.
- Improve end-of-turn detection.
- Handle hesitation, silence and interrupted speech.
- Improve recognition of names, dates, numbers and domain-specific terminology.
- Work on multilingual and code-mixed speech recognition.
- Measure WER and CER.
- Develop business-critical error metrics.
- Measure time to first partial transcript.
- Measure transcript-finalisation latency.
- Build ASR regression datasets.
- Work closely with speech-data specialists.
- Work with infrastructure engineers on production inference.
- Optimise GPU utilisation and inference speed.
REQUIRED SKILLS
- Python
- PyTorch
- Automatic Speech Recognition
- Digital audio fundamentals
- Speech-model training/fine-tuning
- Streaming inference
- Audio preprocessing
- Voice Activity Detection
- GPU inference
- Linux
- Git
RELEVANT TECHNOLOGIES MAY INCLUDE
- Whisper
- Faster-Whisper
- wav2vec 2.0
- Conformer
- NVIDIA NeMo
- CTC
- RNN-T
- Streaming Transformer architectures
GOOD TO HAVE
- Multilingual ASR
- Telephone or low-bandwidth speech
- Diarisation
- Noise suppression
- Semantic endpoint detection
- Real-time media processing
EDUCATIONAL QUALIFICATION
Preferred
B.Tech/B.E./M.Tech/M.Sc. in Computer Science, Electronics, Artificial Intelligence, Machine Learning, Signal Processing, Speech Technology, Computational Linguistics or a related field.
Relevant research experience may compensate for lower conventional industry experience.
WHAT WE WILL EVALUATE
Candidates should understand the distinction between
- Offline transcription
- Streaming transcription
- Partial transcripts
- Endpointing
- Accuracy
- Real-time latency
APPLICATION DETAILS
Please include
- Updated CV
- Current location
- Current CTC
- Expected CTC
- Notice period
- GitHub / Hugging Face links
- Relevant publications
- Details of ASR or speech projects
CONFIDENTIALITY
This role involves confidential research and development. Detailed project requirements and architecture will be disclosed only during the appropriate stage of the recruitment process and may require an NDA.
About Ethical Den
Required Skills
Job Details
Posted by
N/A
Posted on:
13 Sept 2026
About Ethical Den
More open roles
- IN_Senior Associate_Python AI Developer_GCC_Advisory_HyderabadPwC India · Hyderabad, TG, India
- Software Engineer - Python FullstackHARMAN India · Bengaluru, KA, India
- Full Stack Developer - PythonWSP in India · Bengaluru, KA, India
- Principal Full Stack Developer - PythonWSP in India · Bengaluru, KA, India
- Full Stack Developer - PythonWSP in India · Bengaluru, KA, India
- Full Stack Developer - PythonWSP in India · Noida, UP, India