Still hiring?
11 checksActively hiring
In its favour2
- Still on the company's own careers site, checked 1 h agoModerate evidence
- Found in the last 48 hours, before the big job boardsModerate evidence
Browse
All Tech JobsThe whole board, newest first.Roles That Fit MeAnswer a few questions, see your matches.Early WindowFound before the big boards.Direct ApplyStraight to the manager, past the ATS.By specialty
Your materials
CV AnalyzerWhat an ATS sees, and what to fix.Tailor CVBrought in line with one posting.Cover LetterWritten from your CV and the role.You and the process
Hey, I’m Wayjo. I find roles before the big boards.
Free to browse. An account unlocks the rest.
Jobs
All Tech JobsThe whole board, newest first.Roles That Fit MeAnswer a few questions, see your matches.Early WindowFound before the big boards.Direct ApplyStraight to the manager, past the ATS.Not stated
Similar roles pay $210K - 270K a year · our estimate
The whole posting in a few lines. Sign up to read it here and on every role you open.
Sign Up to ReadAs an ML Infrastructure Engineer at the company, you will build the infrastructure for AI systems that learn to improve their own training, inference, and kernels.
Our goal is to unlock a 10× improvement every month somewhere in the stack - from GPU scale and model size to throughput, memory efficiency, caching, and latency.
This role spans training, inference, and kernels. Bring deep expertise in at least one area and curiosity across the stack. We’ll shape your initial ownership around your strengths.
What you will work on
Training
Build and optimize distributed training and RL infrastructure, including rollout execution, GPU scheduling, checkpointing, and recovery. Improve training experimentation throughput and shorten research iteration cycles.
Inference
Optimize serving for production agents and training rollouts. Improve batching, scheduling, and KV-cache management while balancing latency, throughput, cost, and model quality.
Kernels and runtimes
Develop and optimize GPU kernels and runtimes using CUDA, Triton, or comparable tools. Improve memory use and execution efficiency, preserving numerical correctness and verifying gains in real workloads.
Across all three
Build reproducible benchmarks, observability, and automated research workflows that propose changes, run experiments, and validate improvements. Work with researchers to turn new algorithms into reliable systems across training, inference, and kernels.
Continual learning for ML infrastructure
We’re building one of the world’s best continual learning loops for ML infrastructure: agents propose optimizations, run experiments, measure gains, and learn from the results across training, inference, and kernels.
Required Qualifications
Strong fundamentals in distributed systems, networking, storage, and failure recovery, with experience shipping and operating demanding systems.
Required: deep specialization in training infrastructure, inference systems, or GPU kernels, supported by systems built or measurable optimizations delivered.
Strong Python skills and languages relevant to your specialty, such as C++, CUDA, or Triton.
Understanding of PyTorch, JAX, or comparable framework internals, with strong profiling and debugging skills.
Ownership and clear communication: work closely with researchers and deliver measurable performance gains while preserving correctness and reliability.
We value demonstrated capability over credentials.
Preferred
Hands-on experience with training and RL stacks such as Miles, SkyRL, Prime Intellect’s verifiers, or comparable systems. Depending on your specialty, experience with vLLM, SGLang, collective communication, or ML compilers is also valuable.
About the company
The company is a research and product lab creating the platform for continual learning.
AI is the most capable software ever built, and the least able to learn. Every valuable correction and edit that happens in a product evaporates at the next session. A few teams have closed this gap by hand-coupling their models to their products: Composer, Claude Code, Windsurf SWE-1.
The company is the first scalable approach for every company: our platform unlocks the signal already sitting in product use, so companies can continuously post-train large-scale agentic models that outperform the frontier.
Our research team comes from Deepmind, OpenAI, Meta Superintelligence, and product team from Figma, Apple, Stripe, and Windsurf.
From the office
Company hidden
Office in San Francisco, United States
Still hiring, not crowded yet, and you'd be among the first.
Still hiring?
11 checksActively hiring
In its favour2
How crowded?
5 checksLow
In its favour3
How well does this role fit you?
Answer a few questions or drop your CV, and every role gets a fit score with the reasons, this one first.
Something wrong with this vacancy?
$310,000 - 555,000 / year