You and the process

With a person
Anna, in-house recruiter
9 years hiring engineers
30 minutes with a real recruiter
They read your CV with you, on a call, and say where the offers are being lost.
Didn’t find what you were looking for? Tell us what to build
Be the first to open itNo views yet
River AI Inc.

Software Engineer, Inference Systems

  • Office

Salary

$200,000 - 420,000/ year

AI summary

For members

The whole posting in a few lines. Sign up to read it here and on every role you open.

Sign Up to Read

Description

At River AI, our mission is to create personal AI owned and shaped by each individual. To achieve this, we are rewriting the entire stack from scratch: personal hardware for local inference, bespoke training infrastructure, next-generation UIs, and frontier deep learning research.

Who we are

We are scientists, engineers, and builders from the industry's top tech companies and AI labs. We bring a proven track record of scaling consumer systems for hundreds of millions of users and architecting the pre-training infrastructure behind today's frontier models.

About the Role

We are looking for exceptional inference systems engineers to build the engines that serve large models through the River API. Your goal is to deliver fast, reliable inference while making efficient use of GPU compute and memory.

You will take ownership of the serving runtime, from request scheduling and continuous batching to KV-cache management, distributed model execution, and checkpoint loading. Your work will support both customer-facing inference and the sampling workloads that power reinforcement learning.

Working closely with GPU kernel engineers, researchers, and infrastructure engineers, you will bring new models into production and improve their performance across realistic workloads. You will measure success through latency, throughput, reliability, and cost, with careful attention to numerical correctness and model behavior.

Responsibilities

  • Optimize inference for dense and mixture-of-experts models, including fine-tuned models and adapters.
  • Improve batching, caching, and admission control to balance throughput, latency, memory use, and fairness.
  • Accelerate multi-GPU execution, communication, and model loading while preserving model-version consistency.
  • Improve RL sampling throughput while keeping samples and log probabilities tied to the correct model version.
  • Build reliable streaming, cancellation, and recovery under failures and overload.
  • Profile bottlenecks and validate improvements through reproducible performance and correctness tests.

Requirements

  • Minimum Qualifications:
  • Bachelor’s degree in Computer Science, Computer Engineering, or equivalent practical experience.
  • Experience building inference engines or performance-sensitive distributed services.
  • Strong understanding of transformer inference, GPU memory, concurrency, and networking.
  • Proficiency in Python and C++ or Rust.
  • Strong debugging and profiling skills across models, runtimes, and services.
  • A collaborative mindset and strong ownership of engineering outcomes.
  • Preferred Qualifications: (We encourage you to apply even if you don't meet all of these)
  • Experience extending SGLang, vLLM, TensorRT-LLM, or similar frameworks.
  • Work on advanced serving techniques, such as speculative decoding or disaggregated prefill and decode.
  • Familiarity with expert parallelism, GPU collectives, and quantized inference.
  • Experience with multi-adapter serving, dynamic checkpoint loading, or RL sampling.
  • Familiarity with CUDA graphs, custom kernels, and NVIDIA profiling tools.
  • Experience operating model-serving systems under production traffic.
  • Logistics & Benefits
  • Location: Palo Alto, California.
  • Compensation: $200,000–$420,000 USD annual base pay, depending on experience and skills, plus equity.
  • Benefits: Comprehensive health, dental, and vision insurance; unlimited PTO; and relocation assistance as needed.
  • Visa Sponsorship: We sponsor visas and support the process for the right candidate.

Where you’d work

From the office

Relocation offered

Visa sponsored

About the company

River AI Inc.

  • Industry: AI

Office in Palo Alto, United States

Your chances

Worth a look before you spend an evening tailoring a CV for it.

  • 23 checks run
  • 6 red flags

Still hiring?

14 checks

2 red flags

How crowded?

9 checks

4 red flags

Fits Me

How well does this role fit you?

Answer a few questions or drop your CV, and every role gets a fit score with the reasons, this one first.

  • Your field
  • Level
  • Stack
  • Work model
  • Salary floor
  • Must-haves
Details12 facts · Role, Location, Compensation, Company
Tech stack
  • Python
  • C++
  • Rust
  • LLMs
  • CUDA
Equity
Equity offered
Industry
AI
Specialty
Backend
Region
United States
Pay period
Annual
Show 6 more factsShow less

Role

Category
Development
Specialty
Backend
Tech stack
  • Python
  • C++
  • Rust
  • LLMs
  • CUDA

Location

Work model
Office
Region
United States
Office
  • Palo Alto, United States
Relocation
Offered
Visa sponsorship
Sponsored

Compensation

Salary
$200,000 - 420,000 / year
Pay period
Annual
Equity
Equity offered

Company

Industry
AI

Something wrong with this vacancy?

Similar vacancies

  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Staff Software Engineer

    Salary by agreement

    • Remote · Europe, Portugal
    • Senior
  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Senior Software Engineer

    Salary by agreement

    • Remote
    • Senior
  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Software Engineer

    Salary by agreement

    • Hybrid · Toronto
    • Junior
  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Senior Software Engineer

    €85,000 - 140,000 / year

    • Office · Ireland
    • Senior
    Direct apply

Share this vacancy

What's wrong with it?

The employer never sees who reported.

Reason