You and the process

With a person
Anna, in-house recruiter
9 years hiring engineers
30 minutes with a real recruiter
They read your CV with you, on a call, and say where the offers are being lost.
Didn’t find what you were looking for? Tell us what to build
Be the first to open itNo views yet
HUD

Research Engineer, Privacy and Anonymization

  • Remote
  • 3-6 years

Salary

$80,000 - 230,000/ year

AI summary

For members

The whole posting in a few lines. Sign up to read it here and on every role you open.

Sign Up to Read

Description

About HUD

HUD (https://www.hud.ai/)'s mission is to build reliable, fair and open infrastructure for AI data. We want data to be valuable for the people who create it and trustworthy for the labs that train on it. Our team is a quickly growing group of researchers, engineers and operators building the economy that shapes what AI will become. Backed by $16M from top VCs and YC (W25), our marketplace and platform are used by startups, Fortune 500 companies and frontier labs.

About the role

We’re looking for a Research Engineer to build the privacy and anonymization systems that make sensitive, real-world data safe and useful for AI training. You’ll develop methods to detect and remove PII, secrets, and other sensitive information from raw data before it enters our processing and synthetic data pipelines. You’ll own the full pipeline for protecting privacy without destroying the structure and signal that make data valuable for training agents.

Responsibilities

  • Build systems to detect PII, quasi-identifiers, credentials, and other sensitive information and design transformations based on the data type and downstream use case
  • Develop and benchmark detection approaches that combine rules, statistical models, classifiers, and LLM-based methods
  • Build production pipelines that anonymize raw data before it enters downstream processing, training, evaluation, or synthetic data generation workflows
  • Create evaluation frameworks that measure privacy risk and retained data utility, including recall-weighted metrics, leakage tests, and adversarial re-identification attempts
  • Design systems that remain robust to new data sources, schema drift, unusual formats, and sensitive information embedded in unexpected fields
  • Work with engineering, research, operations, and customers to translate privacy requirements into practical technical policies and safeguards
  • Experience
  • You may be a good fit if you have:
  • Strong proficiency in Python and experience building reliable production data or ML systems
  • Experience with information extraction, named-entity recognition, classification, or related methods for detecting rare or sensitive content
  • Strong experimental instincts and the ability to compare approaches across recall, precision, latency, cost, and downstream data utility
  • An understanding of the difference between redaction, masking, pseudonymization, anonymization, and synthetic data—and when each is appropriate
  • High attention to detail and the ability to reason about subtle leakage paths, edge cases, and adversarial failure modes
  • Built data processing pipelines end-to-end without a fully prescribed roadmap
  • Strong candidates may also have:
  • Hands-on experience with privacy-enhancing technologies such as differential privacy, k-anonymity, secure aggregation, format-preserving encryption, etc.
  • Worked with sensitive data in areas such as healthcare, finance, or security
  • Built low-latency or high-throughput ML inference and data-processing systems
  • Worked in unstructured problem spaces and take ownership from early research through production deployment
  • Early-stage startup experience and strong communication skills for collaboration across teams and time zones
  • We prioritize technical aptitude and learning potential over years of experience. Motivated candidates are encouraged to apply even if they don't meet all criteria.
  • Team & company details
  • Team Size: \~25 people currently, mostly full-time in-person, but some remote.
  • Our team: Our team includes 4 International Olympiad medalists (IOI, ILO, IPhO), serial AI startup founders, and researchers with publications at ICLR, NeurIPS, etc.
  • Company stage: We have 8 figures in funding and are scaling profitably and quickly to meet very strong demand.
  • Logistics
  • Employment: Full-time.
  • Location: We have offices in San Francisco or Singapore but are open to remote candidates who can work hours that 70-80% overlap with either San Francisco or Singapore time zones.
  • Visa Sponsorship: We provide support for relocation and visas for strong full-time candidates to the US or Singapore.
  • Timeline: Applications are rolling. The process is 2 technical interviews and a 2-3 day work trial.

Conditions

Actual offers are adjusted for experience and location, but our base salary bands are

San Francisco (and other major US cities): $135,000 - $230,000

Singapore: $100,000 - $175,000

Rest of world: $80,000 - $175,000

Experience: 3+ years

Visa: US citizen/visa only

Benefits

  • Competitive compensation
  • 100% covered top-of-the-line medical, dental, and vision from Blue Shield of CA (US employees)
  • Lunch and dinner when you’re in the office (in-office employees)
  • Company-wide holiday break (Christmas Eve to New Year’s Day) on top of PTO and paid holidays
  • Other perks including an Equinox membership, 401k, and commuter benefits (US employees)
  • Unlimited\ access to tokens for ChatGPT, Claude Code, Cursor, etc. \By unlimited, we mean no one on our token usage leaderboard has ever hit a limit. So we have no idea what the limit is.

Where you’d work

Fully remote

You can work from

  • United States
  • Canada

Relocation offered

No visa sponsorship

About the company

HUD

  • Industry: AI

11 of their 11 open roles are remote

Your chances

Worth a look before you spend an evening tailoring a CV for it.

  • 22 checks run
  • 3 red flags

Still hiring?

14 checks

1 red flag

How crowded?

8 checks

2 red flags

Fits Me

How well does this role fit you?

Answer a few questions or drop your CV, and every role gets a fit score with the reasons, this one first.

  • Your field
  • Level
  • Stack
  • Work model
  • Salary floor
  • Must-haves
Details13 facts · Role, Location, Compensation, Employment, Company
Tech stack
  • Python
  • LLMs
Type
Full-time
Industry
AI
Specialty
ML
Region
United States
Pay period
Annual
Show 7 more factsShow less

Role

Category
Data & Analytics
Specialty
ML
Experience
3+ years
Tech stack
  • Python
  • LLMs

Location

Work model
Remote
Region
United States
Remote from
  • United States
  • Canada
Relocation
Offered
Visa sponsorship
Not sponsored

Compensation

Salary
$80,000 - 230,000 / year
Pay period
Annual

Employment

Type
Full-time

Company

Industry
AI

Something wrong with this vacancy?

Similar vacancies

  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Data & AI Engineer Intern

    Salary by agreement

    • Hybrid · Paris
    • Junior
  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    AI Engineer

    $80,000 - 210,000 / year

    • Office · CA, Austin
    • Senior
  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Senior AI/ML Engineer

    $200,000 - 260,000 / year

    • Office · San Francisco
    • Senior
  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Member of Technical Staff, ML Infra

    Salary by agreement

    • Office · San Francisco
    • Senior

Share this vacancy

What's wrong with it?

The employer never sees who reported.

Reason