You and the process

With a person
Anna, in-house recruiter
9 years hiring engineers
30 minutes with a real recruiter
They read your CV with you, on a call, and say where the offers are being lost.
Didn’t find what you were looking for? Tell us what to build
Early Window: Be the first to open itNo views yetCloses in
Company hidden

Senior Site Reliability Engineer

  • Office
  • 6+ years

Salary

$170,000 - 220,000/ year

AI summary

For members

The whole posting in a few lines. Sign up to read it here and on every role you open.

Sign Up to Read

Description

E-commerce got real-time data infrastructure decades ago. Physical stores still have not. The company is changing that.

The company is building the data infrastructure layer for the physical world, starting with retail. Our hardware-enabled SaaS platform uses proprietary overhead sensors, software, and AI-powered analytics to locate every product in a store, continuously, down to the fixture. We are deployed across 1,500+ stores with retailers including American Eagle Outfitters and Old Navy, processing tens of billions of real-world events every day, delivering 99%+ accuracy in complex, noisy environments - at fleet scale.

Inventory accuracy is only the beginning. We believe the company can become foundational infrastructure for the physical economy, powering new AI-driven commerce experiences across retail and beyond.

Join us if you want to work on a large, unsolved, technically challenging problem with an ambitious team building category-defining technology.

OUR VALUES

Mission-Driven: We're transforming retail with cutting-edge technology and building something that truly matters.

Collaborative Team: We thrive on curiosity, shared goals, and solving complex problems together.

High Impact: You’ll make meaningful contributions from day one and help shape the future of our product and company.

Clear Communication: We value honesty, humility, and respectful dialogue—everyone’s voice matters.

Balanced Lives: We work hard, but not at the expense of well-being. We respect time, boundaries, and life outside of work.

Diverse Perspectives: We believe better ideas come from diverse backgrounds, experiences, and viewpoints.

Empathy-Driven Design: We build with deep respect for our end users, listening closely to their feedback and needs.

The company runs data infrastructure across 1,600+ live retail stores, processing tens of billions of real-world events every day. We’re hiring a Site Reliability Engineer to own the reliability of that system end to end — leading incident response, running day-to-day NOC operations, and building the observability foundation that lets us catch issues before they hit a store floor. You’ll be the steady hand during a live incident, and the engineer making sure there are fewer of them to begin with.

Responsibilities

  • Own the incident management lifecycle end to end: detection, triage, escalation, communication, resolution, and postmortem for production incidents.
  • Act as Incident Commander for high severity incidents, coordinating across engineering, support, and leadership until resolution.
  • Run day-to-day NOC (Network Operations Center) operations, including 24/7 shift coverage, escalation matrices, and shift handover protocols.
  • Coach and mentor NOC analysts on triage discipline and escalation judgment, and own NOC KPIs like response time and escalation accuracy.
  • Design and maintain observability pipelines across metrics, logs, and traces, and define SLIs/SLOs with engineering and product.
  • Build dashboards and alert that surface true signal from our sensor and platform data, cutting down on noise and alert fatigue.
  • Facilitate blameless postmortems and root cause analysis, and track corrective actions through to closure.
  • Maintain on-call rotations, runbooks, and escalation policies, and report on MTTA/MTTR/MTBF trends to leadership.
  • In your first 30 days, you will:
  • Learn the company’s mission, technology stack and core values.
  • Complete onboarding and security compliance training.
  • Shadow the NOC across shifts and review recent incident history and open postmortem action items.
  • In your first 60 days, you will:
  • Take on-call as primary or secondary responder for at least one service area.
  • Tune or consolidate at least one high-volume, low-signal alert source, and audit existing observability coverage for major gaps.
  • Draft or update runbooks for the top recurring incident types, and instrument one under-monitored service.
  • In your first 90 days, you will:
  • Lead Incident Commander duties for high severity incidents, including full postmortem facilitation.
  • Deliver a reliability report on incident trends, NOC KPIs, and observability maturity gaps.
  • Present a roadmap for the next 2–3 quarters covering NOC process, Automation, SLO definitions, and observability investment.
  • At the company, your base pay is one part of your total compensation package. The expected base salary range for this position is $170,000 - $219,000 . Individual pay is determined by work location and additional factors, including job-related skills, experience and relevant education or training. You will also be eligible to receive other benefits including: equity, comprehensive medical and dental coverage, life and disability benefits, 401k plan, flexible time off, and paid parental leave. The pay range listed for this position is a good faith and reasonable estimate of the range of possible base compensation at the time of posting.
  • Research has shown that women & underrepresented minorities are more likely to read lists of requirements and consider themselves unqualified if they don't meet every single one. This list represents what we're ideally looking for, but everyone has unique strengths & weaknesses, and we hire for strength & potential, not lack of weakness.

Requirements

  • Required :
  • You have 5+ years of experience in Site Reliability Engineering, DevOps, Infrastructure, or Production Operations, with direct incident response and on-call experience.
  • You have experience running or actively contributing to a NOC, including shift scheduling, escalation processes, and performance metrics.
  • You have strong hands-on experience with observability tooling (Prometheus, Grafana, Datadog, New Relic, Splunk, ELK, OpenTelemetry, or similar).
  • You have a solid understanding of SLIs, SLOs, SLAs, and error budgets, and how to use them to drive prioritization.
  • You have hands-on release engineering experience, including CI/CD pipelines, deployment automation, and safe rollout practices like canary releases, feature flags, and automated rollbacks.
  • You are proficient in at least one scripting or programming language (Python, Go, Bash, etc.).
  • You have experience with infrastructure-as-code tools (Terraform, Ansible).
  • You have experience with cloud platforms (AWS, GCP, or Azure) and container orchestration (Kubernetes, Docker).
  • You are a clear, direct communicator who stays calm and organized under pressure during live incidents.
  • Preferred:
  • You have experience building or scaling a NOC from the ground up.
  • You have a background in distributed systems architecture and microservices troubleshooting.
  • You have familiarity with chaos engineering and resilience testing.
  • You have a certification such as AWS Certified SysOps Administrator, Google Professional Cloud DevOps Engineer, or ITIL.

Where you’d work

From the office

About the company

Company hidden

  • Industry: SaaS

Offices in Seattle, United States, Sunnyvale, United States, San Diego, United States

Your chances

Still hiring, not crowded yet, and you'd be among the first.

  • 18 checks run
  • 8 good signs
  • 0 red flags

Still hiring?

12 checks

Actively hiring

In its favour5

  • Still on the company's own careers site, checked 1 h agoModerate evidence
  • Found in the last 48 hours, before the big job boardsModerate evidence
  • The company posted 3 roles in the last 2 weeksSlight evidence
2 moreFewer
  • Specific about the basics: pay, place, level, stack and contract all statedSlight evidence
  • States its salarySlight evidence

How crowded?

6 checks

Low

In its favour3

  • In its Early Window: not on the big job boards yetStrong evidence
  • Senior level: far fewer people qualifySlight evidence
  • Asks for Splunk, which fewer than 1% of open roles doSlight evidence

Fits Me

How well does this role fit you?

Answer a few questions or drop your CV, and every role gets a fit score with the reasons, this one first.

  • Your field
  • Level
  • Stack
  • Work model
  • Salary floor
  • Must-haves
Details13 facts · Role, Location, Compensation, Employment, Company
Tech stack
  • Python
  • Go
  • AWS
  • Kubernetes
  • Docker
  • Terraform
  • Azure
  • CI/CD
  • Ansible
  • Prometheus
  • Grafana
  • Datadog
  • Splunk
Seniority
Senior
Type
Full-time
Equity
Equity offered
Industry
SaaS
Specialty
SRE
Show 7 more factsShow less

Role

Category
DevOps & Infrastructure
Specialty
SRE
Seniority
Senior
Experience
5+ years
Tech stack
  • Python
  • Go
  • AWS
  • Kubernetes
  • Docker
  • Terraform
  • Azure
  • CI/CD
  • Ansible
  • Prometheus
  • Grafana
  • Datadog
  • Splunk

Location

Work model
Office
Region
United States
Offices
  • Seattle, United States
  • Sunnyvale, United States
  • San Diego, United States

Compensation

Salary
$170,000 - 220,000 / year
Pay period
Annual
Equity
Equity offered

Employment

Type
Full-time

Company

Industry
SaaS

Something wrong with this vacancy?

Similar vacancies

  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Senior Site Reliability Engineer II

    Salary by agreement

    • Office · New York, Pittsburgh +1
    • Senior
  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Senior Production Engineer

    Salary by agreement

    • Office · Mountain View
    • Senior
  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Staff Site Reliability Engineer

    $220,000 - 330,000 / year

    • Remote · US
    • Senior
  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Site Reliability Engineer

    Salary by agreement

    • Remote

Share this vacancy

What's wrong with it?

The employer never sees who reported.

Reason