Still hiring?
13 checks1 red flag
Browse
All Tech JobsThe whole board, newest first.Roles That Fit MeAnswer a few questions, see your matches.Early WindowFound before the big boards.Direct ApplyStraight to the manager, past the ATS.By specialty
Your materials
CV AnalyzerWhat an ATS sees, and what to fix.Tailor CVBrought in line with one posting.Cover LetterWritten from your CV and the role.You and the process
Hey, I’m Wayjo. I find roles before the big boards.
Free to browse. An account unlocks the rest.
Jobs
All Tech JobsThe whole board, newest first.Roles That Fit MeAnswer a few questions, see your matches.Early WindowFound before the big boards.Direct ApplyStraight to the manager, past the ATS.Not stated
The whole posting in a few lines. Sign up to read it here and on every role you open.
Sign Up to ReadAI Platform builds the foundations of Datadog's AI efforts. The org is 70+ people organised in three pillars: training and serving (GPU clusters, distributed training, low-level infrastructure), agents (agent harnesses, memory systems, the internal AI gateway that routes every LLM request at Datadog), and evaluation and experimentation. This role sits in the evaluation and experimentation pillar, which owns Datadog's shared annotation and evaluation infrastructure — including the evaluation scenario store and the telemetry archival systems used across the Bits org. Together they let an agent travel back in time and query what Datadog looked like at the exact moment an incident happened, so scenarios can be replayed and agent performance tracked over time.
Specifically, you'll be the first applied scientist on GenSim (Generative Simulations), the team that builds the environments Datadog's agents learn in. GenSim doesn't replay sampled telemetry — it stands up real, fully instrumented applications that talk to Datadog, drives them with representative traffic, injects controlled failures, and records what happens. Because GenSim injected the failure, it knows the ground truth. That corpus — hundreds of postmortem-derived scenarios and thousands of runnable applications — is today the primary source of post-training data for Datadog's own SRE model, and the substrate that Bits AI SRE and our other agents are trained and evaluated against.
The team has no applied science support today and is learning post-training data methodology on the fly. That's the gap this role fills, and the open questions are the interesting part. How do you tell whether a generated environment is actually representative of the messy, incomplete telemetry real customers run — rather than a suspiciously clean one where every monitor exists and every service emits complete logs? How do you make injected problems genuinely hard, and how do you even measure difficulty? How do you define and control the quality of post-training data when correctness, representativeness and difficulty pull in different directions? How do you evaluate an agent end to end when the trajectory is non-deterministic? Creating simulated agent environments for monitoring and SRE work is not well solved in open source or in published research, and Datadog is the leading company in this field. If those are the problems you want to spend your time on, come build this with us.
Part of the week in the office
Datadog
Office in Paris, France
Also hiring in New York, United States, Dublin, Ireland, Boston, United States and 6 more places
12 of their 63 open roles are remote
Worth a look before you spend an evening tailoring a CV for it.
Still hiring?
13 checks1 red flag
How crowded?
6 checks2 red flags
How well does this role fit you?
Answer a few questions or drop your CV, and every role gets a fit score with the reasons, this one first.
Something wrong with this vacancy?