Home/Job List/Machine Learning Engineer - ASR & Speech Recognition
Ethical Den

Machine Learning Engineer - ASR & Speech Recognition

Ethical Den

Kolkata metropolitan area, WB, India
Full-Time
Posted 1 day ago

Job Description & Responsibilities

JOB TITLE: Machine Learning Engineer - ASR & Speech Recognition

Company: Ethical Den

Location: Kolkata, West Bengal

Work Mode: On-site only

Employment Type: Full-time

Experience: 2 to 5+ years

Openings: 1

Joining: Immediate to 30 days preferred

ABOUT THE ROLE

Ethical Den is hiring a Machine Learning Engineer specialising in Automatic Speech Recognition and real-time speech processing.

You will work with our AI research team on a confidential speech-technology initiative.

The role involves speech recognition, streaming inference, endpoint detection, speech datasets and optimisation for real-world audio conditions.

Detailed project information will be shared with suitable candidates subject to confidentiality requirements.

KEY RESPONSIBILITIES

  • Research and evaluate ASR architectures.
  • Fine-tune speech-recognition models.
  • Build streaming ASR pipelines.
  • Generate and improve partial transcripts.
  • Optimise recognition for noisy and real-world audio.
  • Develop Voice Activity Detection systems.
  • Work on endpoint detection.
  • Improve end-of-turn detection.
  • Handle hesitation, silence and interrupted speech.
  • Improve recognition of names, dates, numbers and domain-specific terminology.
  • Work on multilingual and code-mixed speech recognition.
  • Measure WER and CER.
  • Develop business-critical error metrics.
  • Measure time to first partial transcript.
  • Measure transcript-finalisation latency.
  • Build ASR regression datasets.
  • Work closely with speech-data specialists.
  • Work with infrastructure engineers on production inference.
  • Optimise GPU utilisation and inference speed.

REQUIRED SKILLS

  • Python
  • PyTorch
  • Automatic Speech Recognition
  • Digital audio fundamentals
  • Speech-model training/fine-tuning
  • Streaming inference
  • Audio preprocessing
  • Voice Activity Detection
  • GPU inference
  • Linux
  • Git

RELEVANT TECHNOLOGIES MAY INCLUDE

  • Whisper
  • Faster-Whisper
  • wav2vec 2.0
  • Conformer
  • NVIDIA NeMo
  • CTC
  • RNN-T
  • Streaming Transformer architectures

GOOD TO HAVE

  • Multilingual ASR
  • Telephone or low-bandwidth speech
  • Diarisation
  • Noise suppression
  • Semantic endpoint detection
  • Real-time media processing

EDUCATIONAL QUALIFICATION

Preferred

B.Tech/B.E./M.Tech/M.Sc. in Computer Science, Electronics, Artificial Intelligence, Machine Learning, Signal Processing, Speech Technology, Computational Linguistics or a related field.

Relevant research experience may compensate for lower conventional industry experience.

WHAT WE WILL EVALUATE

Candidates should understand the distinction between

  • Offline transcription
  • Streaming transcription
  • Partial transcripts
  • Endpointing
  • Accuracy
  • Real-time latency

APPLICATION DETAILS

Please include

  • Updated CV
  • Current location
  • Current CTC
  • Expected CTC
  • Notice period
  • GitHub / Hugging Face links
  • Relevant publications
  • Details of ASR or speech projects

CONFIDENTIALITY

This role involves confidential research and development. Detailed project requirements and architecture will be disclosed only during the appropriate stage of the recruitment process and may require an NDA.

Required Skills

PythonGitMachine LearningPyTorch

Job Details

Employment TypeFull-Time
Work ModeOn-Site
Experience25 years
Positions1

Posted by

N/A

Posted on:

13 Sept 2026

About Ethical Den

More open roles

Browse all jobs →