The Anomaly Detection team’s mission is to build scalable, cost effective systems that can automatically detect anomalies in our customers’ system telemetry. The ultimate goal is to be able to intelligently find problems in system telemetry proactively without the need for our customers to set up or configure specific alarms or alerts. This is a chance to work on a team that sits at the intersection of distributed systems and machine learning and applying those techniques at massive scale to serve a real world customer need.
At the company, we place value in our office culture - the relationships that it builds, the creativity it brings to the table, and the collaboration of being together. We operate as a hybrid workplace to ensure our employees can create a work-life harmony that best fits them.
Responsibilities
Lead and develop an existing engineering team of 8; building trust, setting technical direction, and establishing a high bar for ownership and execution
Own the roadmap and execution for our key signal finding systems
Work closely with product management and applied scientists to build an effective detection system and the right evaluation mechanisms to ensure the quality of our signals
Build and operate both traditional machine learning and LLM based systems in production at scale
Drive the team's AI-native development practices, setting high standards for safety, validation, and increasing agent autonomy over time
Requirements
Experienced engineering manager with a track record of shipping robust and scalable distributed systems with AI and/or ML components. Distributed systems experience is a must-have
Experience with Java, Python or Go, and experience working in a platform-focused team
Strong cross-team collaboration and stakeholder management skills
Comfortable making significant contributions to product strategy in an ambiguous environment rather than executing a pre-defined roadmap
Direct experience building and training ML models in production environments is a plus
Familiarity utilizing streaming technology (e.g. Flink, Kafka) is a plus
The company values people from all walks of life. We understand not everyone will meet all the above qualifications on day one. That's okay. If you’re passionate about technology and want to grow your skills, we encourage you to apply.
Benefits
New hire stock equity (RSUs) and employee stock purchase plan (ESPP)
Continuous professional development, product training, and career pathing
Intradepartmental mentor and buddy program for in-house networking
An inclusive company culture, ability to join our Community Guilds (the company employee resource groups)
Free, global mental health benefits for employees and dependents age 6+
Competitive global benefits
Benefits and Growth listed above may vary based on the country of your employment and the nature of your employment with the company.
The company offers a competitive salary and equity package, and may include variable compensation. Actual compensation is based on factors such as the candidate's skills, qualifications, and experience. In addition, the company offers a wide range of best in class, comprehensive and inclusive employee benefits for this role including healthcare, dental, parental planning, and mental health benefits, a 401(k) plan and match, paid time off, fitness reimbursements, and a discounted employee stock purchase plan.
The reasonably estimated yearly salary for this role at the company is:
$192,000 - $240,000 USD
About the company:
The company is the leading observability and security platform for the AI era, providing businesses with unified visibility across the technology stack to manage complexity at scale. It brings applications, infrastructure, data, models, and security into one place, using AI to detect and resolve issues before they impact customers. Learn more about DatadogLife on Instagram , LinkedIn, and the company Learning Center.