You and the process

With a person
Anna, in-house recruiter
9 years hiring engineers
30 minutes with a real recruiter
They read your CV with you, on a call, and say where the offers are being lost.
Didn’t find what you were looking for? Tell us what to build
Be the first to open itNo views yet
Company hidden

GPU Cluster Infrastructure Engineer

  • Remote
  • 3-6 years

Salary

Not stated

Similar roles pay $150K - 190K a year · our estimate

AI summary

For members

The whole posting in a few lines. Sign up to read it here and on every role you open.

Sign Up to Read

Description

The company is an ultrafast AI inference platform. We built a serverless runtime that launches GPU-backed containers in less than 1 second and quickly scales out to thousands of GPUs. Developers use our platform to serve apps to millions of users around the globe.

About the Role

We're building out our own GPU capacity and we're looking for an experienced contractor to help us stand up high-performance GPU clusters. The work runs from design review through bring-in, and you'll leave behind the operational foundation our team needs to run them.

Review cluster designs and bills of materials across compute, networking, and storage, and catch gaps before hardware is ordered.

Lead acceptance testing: validate cabling and optics, bring up the InfiniBand fabric, run burn-in, and hold vendors to their deliverables.

Stand up and validate high-performance storage alongside vendor teams.

Build the out-of-band management layer and firmware baselines, and secure the management plane for customer-facing environments.

Integrate hardware, fabric, and storage telemetry into our observability stack, with alerting and automated health checks.

Write runbooks, as-builts, and remote-hands procedures.

Provide escalation support after go-live and help our team ramp up.

Requirements

  • You've built and operated NVIDIA HGX or DGX clusters in production at a GPU cloud, HPC center, or AI lab.
  • Hands-on experience with InfiniBand: subnet management and UFM, fabric bring-up, and diagnosing degraded links and optics. NDR or newer.
  • GPU node bring-up and burn-in: firmware, BMC/Redfish, DCGM, NCCL testing, PXE and imaging, and XID error triage.
  • Parallel storage experience: WEKA, VAST, GPFS, Lustre, or similar.
  • Equally effective on the data center floor and remotely, including directing colo remote hands.
  • You troubleshoot methodically across hardware, fabric, and software, document as you go, and communicate clearly with technical and non-technical people.
  • Bonus: recent-generation NVIDIA platforms, bare-metal cloud operations, Ansible or similar automation, Prometheus/Grafana, NVIDIA certifications.

Benefits

  • Competitive salary and meaningful equity
  • Health, dental, and vision benefits with 90% coverage for you and 50% for dependents
  • Opportunities to participate in events across the cloud native community
  • Fitness stipend, learning budget, and much, much more
  • Experience: 3+ years
  • Visa: US citizen/visa only

Where you’d work

Fully remote

You can work from

  • United States

No visa sponsorship

You must already be able to work in the United States

About the company

Company hidden

  • Industry: AI

Your chances

Still hiring, not crowded yet, and a person reads your message.

  • 17 checks run
  • 8 good signs
  • 0 red flags

Still hiring?

11 checks

Actively hiring

In its favour3

  • Still on the company's own careers site, checked 5 h agoModerate evidence
  • Posted 2 days ago: newer than 94% of open rolesModerate evidence
  • A hiring contact is attached to itSlight evidence

How crowded?

6 checks

Low

In its favour5

  • Open for 2 days: you'd be among the earlier applicantsModerate evidence
  • You can message the hiring contact and skip the queueModerate evidence
  • Remote within United States onlySlight evidence
2 moreFewer
  • Only for people already authorized to work in United StatesSlight evidence
  • Asks for Prometheus, which only 1% of open roles doSlight evidence

Fits Me

How well does this role fit you?

Answer a few questions or drop your CV, and every role gets a fit score with the reasons, this one first.

  • Your field
  • Level
  • Stack
  • Work model
  • Salary floor
  • Must-haves
Details14 facts · Role, Location, Compensation, Employment, Company
Tech stack
  • Ansible
  • Prometheus
  • Grafana
Type
Contract
Equity
Equity offered
Industry
AI
Specialty
DevOps
Region
United States
Show 8 more factsShow less

Role

Category
DevOps & Infrastructure
Specialty
DevOps
Experience
3+ years
Tech stack
  • Ansible
  • Prometheus
  • Grafana

Location

Work model
Remote
Region
United States
Remote from
  • United States
Visa sponsorship
Not sponsored
Must already work in
  • United States

Compensation

Salary
Salary by agreement
Pay period
Annual
Equity
Equity offered

Employment

Type
Contract

Company

Industry
AI

Something wrong with this vacancy?

Similar vacancies

  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Software Engineer (DevOps)

    Salary by agreement

    • Office · Krakow
    • Junior
  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Principal Platform Engineer

    $255,000 - 280,000 / year

    • Remote
    • Senior
  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Senior Staff Cloud Platform Engineer

    $230,000 / year

    • Remote · US, Canada
    • Senior

Share this vacancy

What's wrong with it?

The employer never sees who reported.

Reason