Senior ML Ops Engineer at KAYAK
Job Description
📋 Description
- Build and maintain ML infra end-to-end: CI/CD, orchestration, automated training.
- Own model deployment and serving with low latency and high availability.
- Develop core MLOps capabilities: feature stores, model registries, monitoring.
- Operationalize infra for ML: Kubernetes autoscaling, GPU provisioning, self-service tools.
- Improve reliability and observability; define SLOs and automate uptime.
- Empower Data Scientists with standardized workflows to speed ML lifecycle.
🎯 Requirements
- Experience building and operating ML platforms in production.
- Docker, Kubernetes, Linux, and model serving at scale.
- ML lifecycle tooling: feature stores, registries, drift monitoring.
- Own prod systems: SLOs, observability (Prometheus, Grafana, Datadog).
- Production-quality Python or similar language.
- Modernize infra with reliability, risk, and cost focus.
- Own outcomes; communicate clearly with data.
🎁 Benefits
- Work from almost anywhere up to 20 days/year
- Mental health support: therapy and HeadSpace
- No meeting Fridays
- 6 weeks paid vacation + a day off for your birthday
- Paid parental leave
- Paid volunteer time
More Current Jobs at KAYAK
Apply to other open positions at KAYAK
