You and the process

With a person
Anna, in-house recruiter
9 years hiring engineers
30 minutes with a real recruiter
They read your CV with you, on a call, and say where the offers are being lost.
Didn’t find what you were looking for? Tell us what to build
Be the first to open itNo views yet
Mistral AI

Research Engineer, ML Platform

  • Hybrid
  • 3-6 years

Salary

Not stated

Similar roles pay $175K - 245K a year · our estimate

AI summary

For members

The whole posting in a few lines. Sign up to read it here and on every role you open.

Sign Up to Read

Description

About Mistral

Mistral provides full-stack AI solutions: from frontier models to developer tools, applications, and compute. We partner with enterprises tackling the hardest problems across high-stakes industries like finance, manufacturing, defense, healthcare, and the public sector, co-creating customized AI systems that they can run on their terms.

We are a dynamic, collaborative team passionate about AI and its potential to transform society. Our diverse workforce thrives in competitive environments and is committed to driving innovation. Our teams are distributed between Europe, North America, Asia and the Middle East. We are creative, low-ego and team-spirited.

Responsibilities

  • This role focuses on building and operating the ML platform that powers large-scale training, evaluation, and batch inference at Mistral AI. You will develop the infrastructure that enables researchers and engineers to run distributed GPU workloads reliably across clusters, hardware types, and regions.
  • You will work across the full ML lifecycle, from workload scheduling and capacity management to platform APIs, observability, and production operations. You will take ownership of critical systems and help turn complex infrastructure into reliable, self-service capabilities.
  • Build the ML Platform: Develop services, APIs, controllers, and tooling for training, evaluation, fine-tuning, and batch inference.
  • Orchestrate GPU Workloads: Build systems for queueing, admission control, quotas, priorities, preemption, and topology-aware placement.
  • Manage Compute Capacity: Improve how heterogeneous GPU resources are provisioned, allocated, and utilized across clusters.
  • Enable Multi-Cluster Execution: Place workloads based on capacity, data locality, hardware requirements, and organizational priorities.
  • Improve Researcher Experience: Create self-service workflows that make distributed workloads easy to launch, observe, debug, and reproduce.
  • Optimize Performance: Improve GPU utilization, scheduling latency, workload startup time, throughput, and infrastructure efficiency.
  • Build for Reliability: Develop observability, failure recovery, capacity planning, and operational tooling for critical ML workloads.
  • Operate What You Build: Participate in on-call rotations and troubleshoot issues across applications, schedulers, networking, storage, and GPU infrastructure.

Requirements

  • Have 4+ years of experience in ML infrastructure, distributed systems, Kubernetes platform engineering, or a related field.
  • Are proficient in Python or Go and comfortable working with production-grade distributed systems.
  • Have strong Kubernetes knowledge, including controllers, operators, CRDs, scheduling, networking, storage, and resource management.
  • Understand technologies such as Kueue, Karpenter, Volcano, and Kyverno, and the problems they address in workload scheduling, provisioning, and policy enforcement.
  • Understand distributed ML workloads, including training, fine-tuning, evaluation, checkpointing, and batch inference.
  • Are familiar with GPU infrastructure and technologies such as PyTorch, CUDA, NCCL, and high-performance networking.
  • Understand concepts such as quotas, priorities, preemption, gang scheduling, topology awareness, and workload admission.
  • Can diagnose performance and reliability problems across software, orchestration, networking, storage, and hardware.
  • Care about developer experience and enjoy turning complex infrastructure into simple, reliable interfaces.
  • Thrive in an ambiguous, fast-moving environment shaped by frontier AI research.

Benefits

  • We offer a comprehensive benefits package designed to support your well-being, growth, and work-life balance. Benefits vary by country and may include healthcare coverage, parental leave, retirement plans, relocation support, wellness programs, meal and transportation allowances, and other location-specific perks.
  • For the most up-to-date details on benefits available in your location, please refer to our Benefits page .
  • Privacy Policy
  • Your privacy matters to us. You can learn more about how we handle your personal data in our Applicant Privacy Policy .

Where you’d work

Part of the week in the office

You can work from

  • United States

Relocation offered

About the company

Mistral AI

  • Industry: AI

Office in Palo Alto, United States

Also hiring in Paris, France, New York, United States, London, United Kingdom and 10 more places

8 of their 37 open roles are remote

Your chances

Worth a look before you spend an evening tailoring a CV for it.

  • 20 checks run
  • 4 red flags

Still hiring?

13 checks

1 red flag

How crowded?

7 checks

3 red flags

Fits Me

How well does this role fit you?

Answer a few questions or drop your CV, and every role gets a fit score with the reasons, this one first.

  • Your field
  • Level
  • Stack
  • Work model
  • Salary floor
  • Must-haves
Details13 facts · Role, Location, Compensation, Employment, Company
Tech stack
  • Python
  • PyTorch
  • Kubernetes
  • CUDA
Type
Full-time
Industry
AI
Specialty
ML
Region
United States
Pay period
Annual
Show 7 more factsShow less

Role

Category
Data & Analytics
Specialty
ML
Experience
4+ years
Tech stack
  • Python
  • PyTorch
  • Kubernetes
  • CUDA

Location

Work model
Hybrid
Region
United States
Office
  • Palo Alto, United States
Remote from
  • United States
Relocation
Offered

Compensation

Salary
Salary by agreement
Pay period
Annual

Employment

Type
Full-time

Company

Industry
AI

Something wrong with this vacancy?

Similar vacancies

  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Data & AI Engineer Intern

    Salary by agreement

    • Hybrid · Paris
    • Junior
  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    AI Engineer

    $80,000 - 210,000 / year

    • Office · CA, Austin
    • Senior
  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Senior AI/ML Engineer

    $200,000 - 260,000 / year

    • Office · San Francisco
    • Senior
  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Member of Technical Staff, ML Infra

    Salary by agreement

    • Office · San Francisco
    • Senior

Share this vacancy

What's wrong with it?

The employer never sees who reported.

Reason