You and the process

With a person
Anna, in-house recruiter
9 years hiring engineers
30 minutes with a real recruiter
They read your CV with you, on a call, and say where the offers are being lost.
Didn’t find what you were looking for? Tell us what to build
Be the first to open itNo views yet
Datadog

Senior Software Engineer, Chaos Engineering

  • Hybrid
  • 6+ years

Salary

Not stated

Similar roles pay €75K - 115K a year · our estimate

AI summary

For members

The whole posting in a few lines. Sign up to read it here and on every role you open.

Sign Up to Read

Description

Datadog’s Chaos Engineering team builds systems that surface reliability weaknesses before they become outages. As a Senior Software Engineer, you will initially focus on zonal resilience, building automation that helps services safely evacuate and recover from zonal failures, while also contributing to fault injection, incident replay, gameday orchestration, and reliability tooling. You will work across engineering teams to design systems that safely exercise production failure modes and turn findings into verified remediation. You will also help advance the use of AI and automation to identify, test, and close resilience gaps as Datadog’s software and infrastructure evolve.

At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them.

Responsibilities

  • Build zonal-resilience automation that coordinates safe workload evacuations, switchovers, and recovery in partnership with the teams that own affected services.
  • Design and build fault-injection systems for production environments, including infrastructure- and application-level testing, incident replay, and controlled resilience experiments.
  • Develop safeguards such as blast-radius controls, kill switches, validation mechanisms, and rollback paths that keep production experiments contained and reversible.
  • Build agents and automation that help propose failure scenarios, triage experiment results, and connect reliability findings to tracked remediation and verification.
  • Lead gamedays from hypothesis and scenario design through execution, documented findings, remediation tracking, and validation of completed fixes.
  • Design and implement reliable distributed systems, including gRPC services, Kubernetes controllers, and shared platform components, while contributing to technical design and mentoring other engineers.

Requirements

  • You have strong distributed systems fundamentals and can reason about consistency, failure modes, backpressure, idempotency, quorum, retries, and failure recovery.
  • You understand Kubernetes workload lifecycles, including how pods, controllers, scheduling, draining, and eviction interact with resilient system design.
  • You have experience designing, building, or operating production systems where safety, availability, and controlled failure handling are important.
  • You communicate complex technical decisions clearly through design documents, runbooks, postmortems, and cross-functional technical discussions.
  • You are comfortable collaborating across engineering teams to understand unfamiliar systems, identify failure modes, and drive resilience improvements.
  • Experience with reliability engineering, chaos engineering, zonal failover, AI-assisted operational workflows, traffic interception, or large-scale observability systems is beneficial but not required.
  • Datadog values people from all walks of life. We know not everyone will meet all the above qualifications on day one. That’s okay. If you’re passionate about technology and want to grow your experience, we encourage you to apply.

Benefits

  • Develop deep expertise in distributed systems, production resilience, Kubernetes, and large-scale infrastructure.
  • Work on reliability systems that operate across Datadog’s production environment and influence how engineering teams design for failure.
  • Grow your experience designing safe, automated approaches to fault injection, zonal resilience, and incident reproduction.
  • Explore practical applications of AI and automation to reliability engineering and operational workflows.
  • Collaborate with engineers across infrastructure, databases, observability, and service teams on complex systems challenges.
  • Mentor other engineers and contribute to technical designs, engineering practices, and platform strategy.
  • Benefits and Growth listed above may vary based on the country of your employment and the nature of your employment with Datadog.
  • LI-Hybrid
  • About Datadog:
  • Datadog is the leading observability and security platform for the AI era, providing businesses with unified visibility across the technology stack to manage complexity at scale. It brings applications, infrastructure, data, models, and security into one place, using AI to detect and resolve issues before they impact customers. Trusted globally by Fortune 500 companies and high-growth AI leaders, Datadog enables businesses to move faster with clarity and confidence. Learn more about DatadogLife on Instagram , LinkedIn, and Datadog Learning Center.

Where you’d work

Part of the week in the office

About the company

Datadog

  • Industry: SaaS

Office in Paris, France

Also hiring in New York, United States, Dublin, Ireland, Boston, United States and 6 more places

12 of their 63 open roles are remote

Your chances

Worth a look before you spend an evening tailoring a CV for it.

  • 19 checks run
  • 1 red flag

Still hiring?

13 checks

No red flags

How crowded?

6 checks

1 red flag

Fits Me

How well does this role fit you?

Answer a few questions or drop your CV, and every role gets a fit score with the reasons, this one first.

  • Your field
  • Level
  • Stack
  • Work model
  • Salary floor
  • Must-haves
Details11 facts · Role, Location, Compensation, Company
Tech stack
  • Kubernetes
  • gRPC
Seniority
Senior
Industry
SaaS
Specialty
Backend
Region
Europe
Pay period
Annual
Show 5 more factsShow less

Role

Category
Development
Specialty
Backend
Seniority
Senior
Experience
6+ years
Tech stack
  • Kubernetes
  • gRPC

Location

Work model
Hybrid
Region
Europe
Office
  • Paris, France

Compensation

Salary
Salary by agreement
Pay period
Annual

Company

Industry
SaaS

Something wrong with this vacancy?

Similar vacancies

  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Staff Software Engineer

    Salary by agreement

    • Remote · Europe, Portugal
    • Senior
  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Senior Software Engineer

    Salary by agreement

    • Remote
    • Senior
  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Software Engineer

    Salary by agreement

    • Hybrid · Toronto
    • Junior
  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Senior Software Engineer

    €85,000 - 140,000 / year

    • Office · Ireland
    • Senior
    Direct apply

Share this vacancy

What's wrong with it?

The employer never sees who reported.

Reason