You and the process

With a person
Anna, in-house recruiter
9 years hiring engineers
30 minutes with a real recruiter
They read your CV with you, on a call, and say where the offers are being lost.
Didn’t find what you were looking for? Tell us what to build
Be the first to open itNo views yet
SpaceXAI

Site Reliability Engineer - Memphis

  • Office
  • 3-6 years

Salary

Not stated

Similar roles pay $170K - 235K a year · our estimate

AI summary

For members

The whole posting in a few lines. Sign up to read it here and on every role you open.

Sign Up to Read

Description

SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates.

ABOUT THE ROLE:

As a Site Reliability Engineer focused on campus reliability, you will design what the campus watches and trusts, technically command cross-discipline SEVs, and build the guardrails that make the next incident smaller. You are the connective tissue across compute, network, storage, power, and cooling. This role demands calm incident leadership, fleet-scale observability judgment, and the ability to drive reliability work across software and facility boundaries. We are looking for candidates from power plants, nuclear power plants, data centers, or people who currently work or have worked in facilities like SpaceXAI with power, compute, cooling, and related plant systems.

Responsibilities

  • Own monitoring architecture and signal quality: what we alert on, suppress, and trust. Consume NOC noise-disposition feedback to drive suppression and redesign. Treat alert noise as a design failure, not an operator failure.
  • Provide SEV command support: technical incident leadership, bridge coordination with the NOC, and timeline and severity hygiene.
  • Run blameless postmortems and drive corrective actions to closed, not filed.
  • Lead cross-functional reliability projects spanning compute, network, storage, and facility signal boundaries.
  • Build and maintain playbooks, run game days, and keep cross-discipline dependency maps current. Own runbook quality jointly with the NOC (SRE designs; NOC operates and corrects).
  • Define error budgets and availability objectives at campus and service boundaries as adopted by the business.
  • Participate in on-call rotations and incident response for SEV-class events in the Memphis / Southaven data center campus.
  • PROFILE SIGNALS:
  • Experience in power plants, nuclear power plants, data centers, or facilities like SpaceXAI — current or prior — spanning power, compute, cooling, or related plant systems.
  • Proven large-scale incident command experience and calm technical leadership on a bridge.
  • Demonstrated monitoring and observability design at fleet or campus scale, including alert hygiene, suppression, and signal quality. Treats alert noise as a design failure, not an operator failure.
  • Cross-team facilitation and systems engineering depth across software and facility boundaries.
  • SUCCESS MEASURED BY:
  • MTTD / MTTR trend for SEV-class events
  • % of SEVs with a blameless postmortem and closed action items
  • Alert actionable ratio
  • Recurrence rate of incident classes
  • Monitoring coverage of critical dependencies across compute, network, storage, power, and cooling
  • BASIC QUALIFICATIONS:
  • Bachelor's degree in Systems Engineering, Computer Science, Electrical Engineering, or a related field (or equivalent experience).
  • 5+ years of experience in site reliability, systems engineering, plant operations, or large-scale production operations, preferably in power plants, nuclear power plants, high-performance computing, or data center environments.
  • Proven large-scale incident command experience and calm technical leadership on a bridge.
  • Demonstrated monitoring and observability design at fleet or campus scale, including alert hygiene, suppression, and signal quality.
  • Experience working across at least two of: compute, network, storage, power, and cooling / facilities telemetry.
  • Experience writing and operating playbooks or runbooks with a 24/7 operations, control room, or NOC partner.
  • Proficiency in scripting (Python, Bash) for automation and analysis, plus general experience in at least one systems language (C, C++, Java, Go, Rust, or similar). Not required to be expert in all of them.
  • Excellent problem-solving skills with a data-driven approach to reliability engineering.
  • Ability to work collaboratively with cross-functional teams, including NOC, data center operations, and infrastructure engineering.
  • PREFERRED SKILLS AND EXPERIENCE:
  • Experience in power plants, nuclear power plants, or other high-consequence industrial control environments.
  • Current or prior work in data centers or facilities like SpaceXAI spanning power, compute, cooling, and related plant systems.
  • Experience in AI/ML infrastructure or supercomputing environments.
  • Hands-on definition and use of SLOs, SLIs, and error budgets at service or campus boundaries.
  • Experience running game days, dependency mapping, and closed-loop corrective action programs.
  • Familiarity with data center hardware and plant signals (servers, GPUs, networking, power, cooling) in addition to software telemetry.
  • Prior work in a fast-paced startup or tech company like SpaceXAI.

Where you’d work

From the office

About the company

SpaceXAI

  • Industry: AI

Office in TN, United States

Also hiring in Palo Alto, United States, Austin, United States, New York, United States and 3 more places

2 of their 18 open roles are remote

Your chances

Worth a look before you spend an evening tailoring a CV for it.

  • 19 checks run
  • 2 red flags

Still hiring?

13 checks

1 red flag

How crowded?

6 checks

1 red flag

Fits Me

How well does this role fit you?

Answer a few questions or drop your CV, and every role gets a fit score with the reasons, this one first.

  • Your field
  • Level
  • Stack
  • Work model
  • Salary floor
  • Must-haves
Details10 facts · Role, Location, Compensation, Company
Tech stack
  • Python
  • Go
Industry
AI
Specialty
SRE
Region
United States
Pay period
Annual
Show 5 more factsShow less

Role

Category
DevOps & Infrastructure
Specialty
SRE
Experience
5+ years
Tech stack
  • Python
  • Go

Location

Work model
Office
Region
United States
Office
  • TN, United States

Compensation

Salary
Salary by agreement
Pay period
Annual

Company

Industry
AI

Something wrong with this vacancy?

Similar vacancies

Share this vacancy

What's wrong with it?

The employer never sees who reported.

Reason