Lead/Manager Site Reliability Engineering Team (Amsterdam) at Together AI
Job Description
📋 Description
- Be on an on-call (PagerDuty) rotation to respond to incidents that impact availability
- Manage, develop and coach the SRE Team.
- Build and run our infrastructure with Ansible, Terraform, and Kubernetes.
- Build monitoring systems to ensure the highest quality service for our customers
- Design and implement operational processes (deployments and upgrades)
- Debug production issues across all services and levels of the stack
🎯 Requirements
- 7+ years of professional SRE or related experience
- Ideally 2 years as a Lead SRE
- Bachelor's degree in Computer Science or a related field or equivalent work experience
- Expert knowledge of Ansible (roles, playbooks), Terraform, and Kubernetes
- Proficiency in programming/scripting languages
- Direct experience in monitoring and observability practices
More Current Jobs at Together AI
Apply to other open positions at Together AI
