You and the process

With a person
Anna, in-house recruiter
9 years hiring engineers
30 minutes with a real recruiter
They read your CV with you, on a call, and say where the offers are being lost.
Didn’t find what you were looking for? Tell us what to build
Early Window: Be the first to open itNo views yetCloses in
Company hidden

Research Scientist / Engineer

  • Remote

Salary

Not stated

AI summary

For members

The whole posting in a few lines. Sign up to read it here and on every role you open.

Sign Up to Read

Description

You'll make Luma's multimodal models fast — profiling and optimizing GPU, CPU, and accelerator code so they train efficiently and deploy at scale without sacrificing quality. You'll write the kernels and operations that get the most out of the hardware.

This is deep performance work: fused kernels, tensor cores, Triton and CUDA, distributed multi-node deployment. It fits someone with expert GPU-optimization skills and a deep understanding of transformer internals. If you're not at home in CUDA, Triton, and profilers, this is the wrong depth.

What You'll Own

Profile and optimize GPU/CPU/accelerator code for maximum utilization and minimal latency.

Write high-performance PyTorch, Triton, and CUDA, dropping to custom operations when needed.

Develop fused kernels and leverage tensor cores and modern hardware features across platforms.

Optimize model architectures and implementations for distributed multi-node production deployment.

Build performance monitoring and analysis tools and automation.

Research and implement cutting-edge optimization techniques for transformer models.

First 90 Days

One way the first 90 could unfold.

Days 1–30 — Immerse & Diagnose: Profile the current training and inference paths and find the biggest performance wins.

Days 30–60 — Ship & Validate: Land a kernel or architecture optimization that measurably improves utilization or latency.

Days 60–90 — Scale & Systemize: Build the monitoring and automation that keeps performance gains from regressing.

What You Bring

Expert-level Triton/CUDA programming and GPU optimization.

Strong PyTorch skills, including kernel development and custom operations.

Proficiency with profiling tools (NVIDIA Nsight, torch profiler, custom tooling).

Deep understanding of transformer architectures and attention mechanisms.

Requirements

  • Experience with compilers and exporters (torch.compile, TensorRT, ONNX, XLA).
  • Experience optimizing inference workloads for latency and throughput.
  • Triton compiler and kernel fusion techniques.
  • Knowledge of warp-level intrinsics and advanced CUDA optimization.
  • About Luma: Luma's mission is to build unified general intelligence that can generate, understand, and operate in the physical world. We believe multimodality is critical for intelligence — the next step beyond language models comes from vision. Luma is an equal opportunity employer.

Where you’d work

Fully remote

You can work from

  • Europe

About the company

Company hidden

  • Industry: AI

Your chances

Still hiring, not crowded yet, and you'd be among the first.

  • 16 checks run
  • 5 good signs
  • 0 red flags

Still hiring?

11 checks

Actively hiring

In its favour3

  • Still on the company's own careers site, checked 1 h agoModerate evidence
  • Found in the last 48 hours, before the big job boardsModerate evidence
  • The company opened 5 roles and closed 3 in the last 2 weeks: hiring is movingModerate evidence

How crowded?

5 checks

Low

In its favour2

  • In its Early Window: not on the big job boards yetStrong evidence
  • Senior level: far fewer people qualifySlight evidence

Fits Me

How well does this role fit you?

Answer a few questions or drop your CV, and every role gets a fit score with the reasons, this one first.

  • Your field
  • Level
  • Stack
  • Work model
  • Salary floor
  • Must-haves
Details10 facts · Role, Location, Compensation, Employment, Company
Tech stack
  • PyTorch
  • CUDA
Type
Full-time
Industry
AI
Specialty
Data Science
Region
Europe
Pay period
Annual
Show 4 more factsShow less

Role

Category
Data & Analytics
Specialty
Data Science
Tech stack
  • PyTorch
  • CUDA

Location

Work model
Remote
Region
Europe
Remote from
  • Europe

Compensation

Salary
Salary by agreement
Pay period
Annual

Employment

Type
Full-time

Company

Industry
AI

Something wrong with this vacancy?

Similar vacancies

  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Copy of Research Scientist / Engineer

    Salary by agreement

    • Hybrid · London
    • Senior
  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Copy of Research Scientist / Engineer

    Salary by agreement

    • Hybrid · London
    • Senior
  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Research Scientist / Engineer

    Salary by agreement

    • Remote · Europe
    • Senior
  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Data Science Lead

    Salary by agreement

    • Hybrid · London
    • Lead & Manager

Share this vacancy

What's wrong with it?

The employer never sees who reported.

Reason