You and the process

With a person
Anna, in-house recruiter
9 years hiring engineers
30 minutes with a real recruiter
They read your CV with you, on a call, and say where the offers are being lost.
Didn’t find what you were looking for? Tell us what to build
Be the first to open itNo views yet
River AI Inc.

Software Engineer, Distributed Training

  • Office

Salary

$200,000 - 420,000/ year

AI summary

For members

The whole posting in a few lines. Sign up to read it here and on every role you open.

Sign Up to Read

Description

At River AI, our mission is to create personal AI owned and shaped by each individual. To achieve this, we are rewriting the entire stack from scratch: personal hardware for local inference, bespoke training infrastructure, next-generation UIs, and frontier deep learning research.

Who we are

We are scientists, engineers, and builders from the industry's top tech companies and AI labs. We bring a proven track record of scaling consumer systems for hundreds of millions of users and architecting the pre-training infrastructure behind today's frontier models.

About the Role

We are looking for exceptional systems engineers to build the distributed training engines behind the River API. Your goal is to make fine-tuning and reinforcement learning fast, numerically correct, and reliable across large GPU clusters.

You will own the execution of training workloads, including gradient computation, optimizer updates, rollout coordination, and checkpoint recovery. Working closely with researchers and inference engineers, you will bring new learning methods into production and improve how efficiently models use compute.

Responsibilities

  • Build and optimize distributed training for large dense and mixture-of-experts models, including low-rank adapter training.
  • Improve reinforcement-learning pipelines by coordinating sampling, reward computation, training updates, and weight transfer.
  • Optimize GPU memory use, parallelism, and communication to increase training throughput.
  • Implement reliable checkpointing, resumption, and worker recovery while preserving consistent training state.
  • Validate losses, gradients, and optimizer behavior, and diagnose numerical or distributed execution failures.
  • Partner with researchers to implement new algorithms and make them accessible through the River API.

Requirements

  • Minimum Qualifications:
  • Bachelor’s degree in Computer Science, Computer Engineering, or equivalent practical industry experience.
  • Hands-on experience building or substantially improving distributed model-training systems.
  • Strong proficiency in Python and a modern deep-learning framework, such as PyTorch or JAX.
  • Solid understanding of backpropagation, optimizers, mixed-precision training, and GPU memory management.
  • Strong debugging skills across concurrent execution, collective communication, and distributed failure recovery.
  • A highly collaborative mindset and a bias for action to push boundaries across the stack.
  • Preferred Qualifications: (We encourage you to apply even if you don't meet all of these)
  • Experience with reinforcement-learning infrastructure, rollout generation, or asynchronous training.
  • Familiarity with tensor, pipeline, expert, or data parallelism and their performance tradeoffs.
  • Work on LoRA, mixture-of-experts training, distributed optimizers, or activation checkpointing.
  • Experience with NCCL, communication profiling, and overlapping computation with data transfers.
  • Proficiency in C++, Rust, or CUDA, with experience investigating performance below the framework layer.
  • Contributions to training frameworks or a track record of operating large training runs.
  • Logistics & Benefits
  • Location: Palo Alto, California.
  • Compensation: Depending on experience and skills the expected base pay is $200,000 - $420,000 USD per year, plus equity.
  • Benefits: Comprehensive health, dental, and vision insurance; unlimited PTO; and relocation assistance as needed.
  • Visa Sponsorship: We sponsor visas and are committed to supporting the process for the right candidate.

Where you’d work

From the office

Relocation offered

Visa sponsored

About the company

River AI Inc.

  • Industry: AI

Office in Palo Alto, United States

Your chances

Worth a look before you spend an evening tailoring a CV for it.

  • 23 checks run
  • 6 red flags

Still hiring?

14 checks

2 red flags

How crowded?

9 checks

4 red flags

Fits Me

How well does this role fit you?

Answer a few questions or drop your CV, and every role gets a fit score with the reasons, this one first.

  • Your field
  • Level
  • Stack
  • Work model
  • Salary floor
  • Must-haves
Details12 facts · Role, Location, Compensation, Company
Tech stack
  • Python
  • PyTorch
  • C++
  • Rust
  • CUDA
Equity
Equity offered
Industry
AI
Specialty
Backend
Region
United States
Pay period
Annual
Show 6 more factsShow less

Role

Category
Development
Specialty
Backend
Tech stack
  • Python
  • PyTorch
  • C++
  • Rust
  • CUDA

Location

Work model
Office
Region
United States
Office
  • Palo Alto, United States
Relocation
Offered
Visa sponsorship
Sponsored

Compensation

Salary
$200,000 - 420,000 / year
Pay period
Annual
Equity
Equity offered

Company

Industry
AI

Something wrong with this vacancy?

Similar vacancies

  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Staff Software Engineer

    Salary by agreement

    • Remote · Europe, Portugal
    • Senior
  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Senior Software Engineer

    Salary by agreement

    • Remote
    • Senior
  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Software Engineer

    Salary by agreement

    • Hybrid · Toronto
    • Junior
  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Senior Software Engineer

    €85,000 - 140,000 / year

    • Office · Ireland
    • Senior
    Direct apply

Share this vacancy

What's wrong with it?

The employer never sees who reported.

Reason