About the role
Focus on topics such as generative AI model development, speech generation, speech recognition, conversational AI systems, and machine learning data pipelines for voice-based applications.
Responsibilities
- Design and develop Generative AI models for voice and speech applications.
- Build and fine-tune speech generation models (TTS) and speech recognition models (ASR).
- Develop voice bots and conversational AI systems using modern LLM frameworks.
- Train and optimize deep learning models for audio and speech processing.
- Build and maintain data pipelines for large-scale speech and audio datasets.
Full description
About the Role
We are looking for a GenAI Engineer specializing in Voice to design, build, and deploy generative AI models for voice-based applications. The ideal candidate will have hands-on experience with speech generation, speech recognition, conversational AI systems, and ML data pipelines for production-grade voice AI systems.
You will work closely with AI researchers, product teams, and backend engineers to develop scalable AI solutions for real-time voice interactions, monitoring model performance and managing audio data pipelines.
Key Responsibilities
· Design and develop Generative AI models for voice and speech applications.
· Build and fine-tune speech generation models (TTS) and speech recognition models (ASR).
· Develop voice bots and conversational AI systems using modern LLM frameworks.
· Train and optimize deep learning models for audio and speech processing.
· Build and maintain data pipelines for large-scale speech and audio datasets.
· Implement ML observability, monitoring, and evaluation pipelines for voice models using tools such as Arize AI.
· Integrate advanced voice synthesis and voice cloning APIs such as ElevenLabs into conversational systems.
· Develop real-time audio streaming and inference pipelines for low-latency voice interactions.
· Integrate AI models into production systems, APIs, and microservices.
· Optimize models for latency, scalability, and real-time performance.
· Evaluate model performance using speech quality metrics, NLP metrics, and monitoring tools.
· Stay updated with advancements in GenAI, speech synthesis, voice cloning, and conversational AI technologies.
Required Skills
· 3–5 years of experience in Machine Learning / AI / Deep Learning.
· Strong programming skills in Python.
· Experience with ML frameworks such as PyTorch or TensorFlow.
· Experience working with speech technologies:
· Text-to-Speech (TTS)
· Automatic Speech Recognition (ASR)
· Voice synthesis or voice cloning
· Familiarity with LLMs and Generative AI frameworks.
· Experience with audio processing libraries (Librosa, Torchaudio).
· Understanding of transformer architectures and deep learning models.
· Experience working with ML data pipelines and dataset versioning.
· Familiarity with model monitoring, observability, and evaluation platforms such as Arize AI.
Preferred Qualifications
· Experience with LLM frameworks such as LangChain, Hugging Face, and OpenAI APIs.
· Experience integrating voice AI platforms like ElevenLabs for speech generation and voice cloning.
· Experience building voice agents or conversational AI systems. Knowledge of RAG pipelines for voice assistants.
· Experience building ML data pipelines for speech datasets.
· Experience deploying models using Docker / Kubernetes / cloud platforms (AWS, GCP, Azure).
· Familiarity with real-time audio streaming systems and inference pipelines.
· Experience with model monitoring, observability, and production ML systems.