You and the process

With a person
Anna, in-house recruiter
9 years hiring engineers
30 minutes with a real recruiter
They read your CV with you, on a call, and say where the offers are being lost.
Didn’t find what you were looking for? Tell us what to build
Be the first to open itNo views yet
HUD

Senior Software Engineer, Infrastructure

  • Remote
  • 6+ years

Salary

$105,000 - 260,000/ year

AI summary

For members

The whole posting in a few lines. Sign up to read it here and on every role you open.

Sign Up to Read

Description

About HUD

HUD (https://www.hud.ai/)'s mission is to build reliable, fair and open infrastructure for AI data. We want data to be valuable for the people who create it and trustworthy for the labs that train on it. Our team is a quickly growing group of researchers, engineers and operators building the economy that shapes what AI will become. Backed by $16M from top VCs and YC (W25), our marketplace and platform are used by startups, Fortune 500 companies and frontier labs.

About the role

We’re looking for a Senior Software Engineer, Infrastructure who can own the reliability, scale, performance, and developer experience of HUD’s core infrastructure and backend systems.

This is not a pure infrastructure role. The right person has strong production infra experience, but also thinks like a backend engineer: they can reason about service architecture, queues, databases, APIs, deployment safety, performance bottlenecks, and how product requirements translate into resilient systems. You’ll work across AWS, Kubernetes, Terraform, CI/CD, observability, and backend services to make HUD faster, more reliable, cheaper to run, and easier for engineers to build on.

Responsibilities

  • Own production uptime, latency, provisioning speed, infrastructure cost, and incident response for core platform services
  • Build and maintain AWS infrastructure with Terraform, Kubernetes/EKS, Helm, Docker, EC2, CodeBuild, ECR, S3, IAM, networking, and secrets management
  • Design and improve backend and platform systems for scale, including capacity planning, autoscaling, queueing, backpressure, cleanup jobs, retries, and rollback paths
  • Define and improve dashboards, alerts, logs, traces, SLOs, runbooks, and on-call workflows so failures are detected, debugged, and resolved quickly
  • Build reliable CI/CD, release automation, environment management, and deployment workflows that improve developer productivity and reduce production risk
  • Write clean, maintainable code where needed to automate systems, improve backend services, and create internal tooling
  • Experience
  • You may be a good fit if you:
  • Have owned production cloud infrastructure for a high-availability, user-facing platform, with responsibility for uptime, performance, deployment safety, and cost
  • Have deep experience with AWS infrastructure and containerized systems; experience with tools like Terraform, Kubernetes/EKS, Docker, EC2, CodeBuild, ECR, S3, IAM, load balancers, networking, and secrets management is strongly preferred
  • Have built or operated CI/CD, environment management, release automation, observability, alerting, and incident response systems
  • Have strong backend engineering judgment and can reason about service architecture, APIs, databases, async systems, queues, scaling limits, and production failure modes
  • Can write clean, maintainable code and apply strong software engineering judgment across product architecture, infrastructure, backend systems, and developer workflows
  • Strong candidates may also have:
  • Experience operating infrastructure for data-heavy, ML/AI, workflow, marketplace, developer-tools, or enterprise platforms
  • Experience designing systems for bursty workloads, long-running jobs, sandboxed execution, distributed workers, or high-concurrency services
  • Experience reducing cloud spend through better architecture, autoscaling, workload placement, caching, cleanup systems, or observability
  • Experience building internal platforms or tools that make engineers faster without hiding too much complexity
  • We prioritize technical aptitude, ownership, and learning potential over years of experience.
  • Team & company details
  • Team Size: \~25 people currently, mostly full-time in-person, but some remote.
  • Our team: Our team includes 4 International Olympiad medalists (IOI, ILO, IPhO), serial AI startup founders, and researchers with publications at ICLR, NeurIPS, etc.
  • Company stage: We have 8 figures in funding and are scaling profitably and quickly to meet very strong demand.
  • Logistics
  • Employment: Full-time.
  • Location: We have offices in San Francisco or Singapore but are open to remote candidates who can work hours that 70-80% overlap with either San Francisco or Singapore time zones.
  • Visa Sponsorship: We provide support for relocation and visas for strong full-time candidates to the US or Singapore.
  • Timeline: Applications are rolling. The process is 2 technical interviews and a 2-3 day work trial.

Conditions

Actual offers are adjusted for experience and location, but our base salary bands are

San Francisco (and other major US cities): $175,000 - $260,000

Singapore: $130,000 - $195,000

Rest of world: $105,000 - $195,000

Experience: 3+ years

Visa: US citizen/visa only

Benefits

  • Competitive compensation
  • 100% covered top-of-the-line medical, dental, and vision from Blue Shield of CA (US employees)
  • Lunch and dinner when you’re in the office (in-office employees)
  • Company-wide holiday break (Christmas Eve to New Year’s Day) on top of PTO and paid holidays
  • Other perks including an Equinox membership, 401k, and commuter benefits (US employees)
  • Unlimited\ access to tokens for ChatGPT, Claude Code, Cursor, etc. \By unlimited, we mean no one on our token usage leaderboard has ever hit a limit. So we have no idea what the limit is.

Where you’d work

Fully remote

You can work from

  • United States
  • Canada

Relocation offered

No visa sponsorship

About the company

HUD

  • Industry: AI

11 of their 11 open roles are remote

Your chances

Worth a look before you spend an evening tailoring a CV for it.

  • 22 checks run
  • 3 red flags

Still hiring?

14 checks

1 red flag

How crowded?

8 checks

2 red flags

Fits Me

How well does this role fit you?

Answer a few questions or drop your CV, and every role gets a fit score with the reasons, this one first.

  • Your field
  • Level
  • Stack
  • Work model
  • Salary floor
  • Must-haves
Details13 facts · Role, Location, Compensation, Employment, Company
Tech stack
  • AWS
  • Kubernetes
  • Docker
  • Terraform
  • CI/CD
  • Helm
Seniority
Senior
Type
Full-time
Industry
AI
Region
United States
Pay period
Annual
Show 7 more factsShow less

Role

Category
DevOps & Infrastructure
Seniority
Senior
Experience
3+ years
Tech stack
  • AWS
  • Kubernetes
  • Docker
  • Terraform
  • CI/CD
  • Helm

Location

Work model
Remote
Region
United States
Remote from
  • United States
  • Canada
Relocation
Offered
Visa sponsorship
Not sponsored

Compensation

Salary
$105,000 - 260,000 / year
Pay period
Annual

Employment

Type
Full-time

Company

Industry
AI

Something wrong with this vacancy?

Similar vacancies

  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Lead IT Engineer

    Salary by agreement

    • Remote
    • Lead & Manager
  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Security Engineer, New Grad

    Salary by agreement

    • Office · Dublin
    • Junior
  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Site Reliability Engineer

    Salary by agreement

    • Remote
  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Senior Security Operations Analyst

    Salary by agreement

    • Hybrid · Wellington, Auckland
    • Senior

Share this vacancy

What's wrong with it?

The employer never sees who reported.

Reason