You and the process

With a person
Anna, in-house recruiter
9 years hiring engineers
30 minutes with a real recruiter
They read your CV with you, on a call, and say where the offers are being lost.
Didn’t find what you were looking for? Tell us what to build
Be the first to open itNo views yet
Modal Labs

Member of Technical Staff - Machines

  • Office
  • 6+ years

Salary

$250,000 - 300,000/ year

AI summary

For members

The whole posting in a few lines. Sign up to read it here and on every role you open.

Sign Up to Read

Description

AI needs a new infrastructure layer. We're building it at Modal.

Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now.

Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale.

We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September.

Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience.

Responsibilities

  • We are looking for strong engineers with experience and interest in designing, building, and maintaining the novel, high-performance systems that make up our serverless platform. Specifically, you'll be working on Modal's machines layer: the fleet of bare metal and cloud hosts that every Function, Sandbox, and training job runs on, and the control plane that provisions, images, monitors, and repairs them. You'll automate the integration of new capacity from a growing set of hardware providers; from auditing and benchmarking hosts and clusters, to maintaining our machine images, configuring GPUs, RDMA, networking, and storage, and getting machines into production. You'll build the automation that keeps the fleet healthy without human intervention: detecting bad GPUs, thermals, and disks. You'll dig into whatever is between the hardware and the software that runs on top of it, whether that is a kernel panic, a broadcast storm during boot, or getting our container runtime to run on new architectures and platforms.

Requirements

  • 5+ years of experience writing high-quality production code
  • Experience operating fleets of physical hardware (bare metal provisioning, BMC/IPMI, PXE or network boot, firmware) or building the control planes that manage them (the more challenges you've worked through, the better)
  • Strong cloud skills
  • Strong knowledge of low-level operating system foundations (Linux kernel, drivers, networking, file systems, containers, etc.)
  • Effective at debugging across layers, from BGP flapping, Linux RPS, and vBIOS bugs to a Python control-plane service
  • Willingness to step into the thick of it with our on-call rotation and respond to production incidents
  • Nice-to-Haves:
  • Experience with GPUs and the NVIDIA software stack in production (drivers, health monitoring, XIDs, RDMA/NVLink)
  • Prior experience with Go
  • Key Things the Team Is Working On:
  • Automatic remediation of unhealthy machines (power cycling, reimaging, GPU recovery) to maximize uptime and minimize operator toil.
  • Automatic integration of new CPU, GPU, and storage servers into the fleet while managing hardware and network heterogeneity.
  • Network health monitoring and reliability across many datacenters, and standardization of bare metal network configuration.
  • Automatic hardware acceptance testing and benchmarking (CPU, disk, GPU, interconnect, network).
  • Custom network bootloader, machine image pipeline, and kernel and firmware management across the fleet.

Where you’d work

From the office

About the company

Modal Labs

Offices in San Francisco, United States, New York, United States

Also hiring in Stockholm, Sweden

1 of their 8 open roles is remote

Your chances

Worth a look before you spend an evening tailoring a CV for it.

  • 21 checks run
  • 2 red flags

Still hiring?

14 checks

No red flags

How crowded?

7 checks

2 red flags

Fits Me

How well does this role fit you?

Answer a few questions or drop your CV, and every role gets a fit score with the reasons, this one first.

  • Your field
  • Level
  • Stack
  • Work model
  • Salary floor
  • Must-haves
Details10 facts · Role, Location, Compensation, Employment
Tech stack
  • Python
  • Go
  • Linux
Seniority
Staff
Type
Full-time
Region
United States
Pay period
Annual
Show 5 more factsShow less

Role

Category
DevOps & Infrastructure
Seniority
Staff
Experience
5+ years
Tech stack
  • Python
  • Go
  • Linux

Location

Work model
Office
Region
United States
Offices
  • San Francisco, United States
  • New York, United States

Compensation

Salary
$250,000 - 300,000 / year
Pay period
Annual

Employment

Type
Full-time

Something wrong with this vacancy?

Similar vacancies

  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Lead IT Engineer

    Salary by agreement

    • Remote
    • Lead & Manager
  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Security Engineer, New Grad

    Salary by agreement

    • Office · Dublin
    • Junior
  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Site Reliability Engineer

    Salary by agreement

    • Remote
  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Senior Security Operations Analyst

    Salary by agreement

    • Hybrid · Wellington, Auckland
    • Senior

Share this vacancy

What's wrong with it?

The employer never sees who reported.

Reason