About the role
Focus on topics such as machine learning fundamentals, MLOps platforms, model lifecycle management, observability, deployment automation, and analytics enablement in diverse ML domains like computer vision and NLP.
Responsibilities
- Design, build, and maintain end‑to‑end MLOps pipelines for multiple ML products.
- Own the MLOps backbone that supports model deployment, monitoring, governance, and lifecycle management.
- Build monitoring solutions for model performance, reliability, and usage.
- Implement CI/CD pipelines for ML systems.
- Collaborate with data science and product teams to operationalize models reliably.
Full description
Job Description
Role Overview
We are looking for a Senior MLOps Engineer to build and own the backbone of our machine learning production ecosystem. This role is critical to enabling scalable, reliable, and observable ML systems across multiple products.
The ideal candidate has strong machine learning fundamentals combined with real‑world MLOps and DevOps experience, and can operate end‑to‑end from model validation and deployment to monitoring, drift detection, and governance. This is not a DevOps‑only role; strong ML understanding is mandatory.
Key Responsibilities
Experience- 4+ Years
MLOps Platform & Architecture
Design, build, and maintain end‑to‑end MLOps pipelines for multiple ML products.
Own the MLOps backbone that supports model deployment, monitoring, governance, and lifecycle management.
Enable a “plug‑and‑play” monitoring and dashboarding capability usable across teams.
Model Lifecycle & Governance
Support model validation, registration, versioning, and lifecycle management.
Implement best practices for model cards, model registries, and traceability.
Collaborate with data science and product teams to operationalize models reliably.
Monitoring, Observability & Drift Detection
Build monitoring solutions for model performance, reliability, and usage.
Implement model drift and data drift detection, leveraging statistical techniques (e.g., KS test, Fisher test).
Integrate monitoring pipelines using tools like Prometheus and Grafana.
Enable root‑cause analysis by linking model degradation to data distribution changes.
Deployment & Automation
Implement CI/CD pipelines for ML systems.
Handle containerized deployments using Docker and Kubernetes.
Support model packaging and deployment, including optimization and edge deployment considerations.
Manage orchestration workflows using Apache Airflow.
ML & Analytics Enablement
Work across ML domains including computer vision, voice, and NLP/GenAI.
Collaborate on upcoming LLM / VLM (Vision‑Language Model) deployments.
Support KPIs and metrics (accuracy, precision, recall, etc.) and enable analytics via dashboards.
Must‑Have Skills (Non‑Negotiable)
Core Foundations
Strong understanding of Machine Learning fundamentals and statistics
Hands‑on experience with model development, validation, and deployment
Ability to reason about data distributions, drift, and performance degradation
MLOps & Engineering
Apache Airflow (or similar orchestration tools)
MLflow (tracking, model registry, versioning)
CI/CD pipelines for ML systems
Docker & Kubernetes
Model monitoring and observability
Monitoring & Visualization
Prometheus & Grafana
Experience building operational dashboards
Familiarity with visualization tools such as Power BI or Tableau
Nice‑to‑Have Skills
Experience with LLMs, GenAI, or Vision‑Language Models (VLMs)
Knowledge of edge deployment, model optimization, quantization
Experience with GitHub / GitLab
Exposure to ML systems in regulated or high‑reliability environments
What Success Looks Like
MLOps becomes a central, reusable service across teams.
Models in production are observable, reliable, and governed.
Drift detection is actionable, not just cosmetic.
Product teams can understand model health at a glance through dashboards.
ML systems scale without fragile, one‑off solutions.