You and the process

With a person
Anna, in-house recruiter
9 years hiring engineers
30 minutes with a real recruiter
They read your CV with you, on a call, and say where the offers are being lost.
Didn’t find what you were looking for? Tell us what to build
Be the first to open itNo views yet
Company hidden

ML Research Engineer, Data

  • Office

Salary

$140,000 - 200,000/ year

AI summary

For members

The whole posting in a few lines. Sign up to read it here and on every role you open.

Sign Up to Read

Description

Join Us, and Ship Robots

Weave was founded to build the robots we’d want to have in our own home. We believe the next generation of robotics will transform everyday life by enabling people to do more and to reclaim time to spend on what’s important.

We also believe robots are in a sense like any other product: to matter, they have to ship. Our robots are already operating in real homes and businesses, giving us the opportunity to rapidly improve from real-world experience. With a growing team, strong customer demand, and capital for expansion, we’re entering an exciting stage of growth—and we’re looking for people with exceptional talent and standards to help bring home robotics to millions of households.

Responsibilities

  • Most robot learning research is graded on evals that don't survive contact with the field. Ours is graded by robots doing useful work in real homes and businesses, every day. Our fleet generates robot data at terabyte scale that’s ingested by training runs every week. Our models are only as good as what we feed them, and you’ll architect and build the pipeline that drives their behavior.
  • You'll turn raw fleet data (video, proprioception, actions, sensor streams) into the datasets our models train on, and follow that data into the training loop: sampling ratios, data mixes, and curricula are decisions you'll shape with the research team. The job is equal parts data engineering and data understanding: build the platform that processes millions of episodes and feeds them to training, and know the data well enough to say what correlates with good and bad model behavior.
  • Know the data: characterize coverage, redundancy, and drift: spectral analysis on time-series, distributional statistics, clustering over embeddings.
  • Find the needles: dig bad data out of terabytes of episodes: bad trajectories, bad annotations, dropped frames, desynced streams, then automate the catch so it never gets through again.
  • Build the data lifecycle: curation, preprocessing, annotation, augmentation, versioning, and the training-ingest formats and loaders that serve it at full throughput.
  • Use models as instruments: embedding search to mine scenarios, model loss and disagreement as quality signals, VLM-assisted filtering and labeling.
  • Improve models through data: partner with researchers on model failures, then build the datasets, processing steps and samplers that target specific capabilities and failure modes, and own the sampling and mixture decisions that go into each run.
  • Build the eval datasets and benchmarks that measure performance across tasks, environments, embodiments, and model versions.
  • What You'll Bring
  • Data engineering at scale: pipelines over terabytes, object stores (S3, GCS), distributed storage, and indexing.
  • Training ingest: high-throughput formats and dataloaders.
  • Working ML experience: you can launch a fine-tune, read a loss curve, and design an ablation to test a data hypothesis.
  • Analytical range: signal processing and statistics on real sensor data, and unsupervised structure-finding (PCA, UMAP, clustering) when the labels don't exist yet.
  • Data debugging: you can trace a problem from sensor drift through a corrupted episode to a pipeline failure.
  • Strong Python and software engineering fundamentals. C++ is a plus.

Requirements

  • Batch processing at scale: Ray, Spark, or Dask over terabytes, with cost and throughput judgment.
  • Workflow orchestration in production: Airflow, Kubeflow, or similar, with retries and monitoring.
  • Robotics data pitfalls: timestamps, clock domains, and sensor modality quirks.
  • Robot learning exposure: you’ve trained policies (VLAs, world models, RL) and can tell a data problem from a model problem.
  • Experience: Any (new grads ok)
  • Visa: US citizen/visa only

Where you’d work

From the office

The office

No visa sponsorship

About the company

Company hidden

Office in San Francisco, United States

Your chances

Still hiring, not crowded yet, and a person reads your message.

  • 19 checks run
  • 9 good signs
  • 0 red flags

Still hiring?

13 checks

Actively hiring

In its favour5

  • Still on the company's own careers site, checked 3 h agoModerate evidence
  • Posted 5 days ago: newer than 88% of open rolesModerate evidence
  • The company opened 16 roles and closed 3 in the last 2 weeks: hiring is movingModerate evidence
2 moreFewer
  • States its salarySlight evidence
  • A hiring contact is attached to itSlight evidence

How crowded?

6 checks

Low

In its favour4

  • You can message the hiring contact and skip the queueModerate evidence
  • In the office in San Francisco: only people nearby can take itSlight evidence
  • Asks for Airflow, which only 2% of open roles doSlight evidence
1 moreFewer
  • Pays below most similar roles, which thins the crowdSlight evidence

Fits Me

How well does this role fit you?

Answer a few questions or drop your CV, and every role gets a fit score with the reasons, this one first.

  • Your field
  • Level
  • Stack
  • Work model
  • Salary floor
  • Must-haves
Details10 facts · Role, Location, Compensation, Employment
Tech stack
  • Python
  • Spark
  • Airflow
Type
Full-time
Specialty
ML
Region
United States
Pay period
Annual
Show 5 more factsShow less

Role

Category
Data & Analytics
Specialty
ML
Tech stack
  • Python
  • Spark
  • Airflow

Location

Work model
Office
Region
United States
Office
  • San Francisco, United States
Visa sponsorship
Not sponsored

Compensation

Salary
$140,000 - 200,000 / year
Pay period
Annual

Employment

Type
Full-time

Something wrong with this vacancy?

Similar vacancies

  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Data & AI Engineer Intern

    Salary by agreement

    • Hybrid · Paris
    • Junior
  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    AI Engineer

    $80,000 - 210,000 / year

    • Office · CA, Austin
    • Senior
  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Senior AI/ML Engineer

    $200,000 - 260,000 / year

    • Office · San Francisco
    • Senior
  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Member of Technical Staff, ML Infra

    Salary by agreement

    • Office · San Francisco
    • Senior

Share this vacancy

What's wrong with it?

The employer never sees who reported.

Reason