Sr. Evaluation Engineer at LogicMonitor
Job Description
📋 Description
- Define quality metrics for incident diagnostics, root-cause analysis, and alert correlation.
- Build offline and online evaluation pipelines in Python; CI/CD integration.
- Lead golden datasets and regression suites across metrics, logs, and incidents.
- Create customer-specific scenarios across technologies and failure modes.
- Use human-authored and AI-assisted methods to generate regression and edge test cases.
- Monitor AI quality and drift; establish evaluation-driven development; mentor engineers.
🎯 Requirements
- 5+ years in software engineering, ML, AI, or related field.
- Strong Python engineering and production systems experience.
- Hands-on AI evaluation, experimentation, testing, and quality frameworks.
- Experience with LangSmith, Arize Phoenix, MLflow, or similar evaluation tools.
- Ability to select/integrate evaluation frameworks for offline testing, online monitoring, regression analysis, and release gating.
- Strong understanding of LLMs, agents, retrieval-augmented generation, prompt engineering, tool calling.
- Experience evaluating non-deterministic, multi-step, or multi-agent AI systems.
- Experience with regression testing, CI/CD, production monitoring, drift detection, and failure analysis.
🎁 Benefits
- Great Place To Work certification.
- BuiltIn's Best Places to Work for the seventh year.
- Equal opportunity employer with an inclusive culture.
More Current Jobs at LogicMonitor
Apply to other open positions at LogicMonitor

