You and the process

With a person
Anna, in-house recruiter
9 years hiring engineers
30 minutes with a real recruiter
They read your CV with you, on a call, and say where the offers are being lost.
Didn’t find what you were looking for? Tell us what to build
Be the first to open itNo views yet
Inferact

Member of Technical Staff, Inference

  • Remote
  • 6+ years

Salary

Not stated

Similar roles pay $190K - 245K a year · our estimate

AI summary

For members

The whole posting in a few lines. Sign up to read it here and on every role you open.

Sign Up to Read

Description

Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware—a position that took years to build.

About the Role

We're looking for an inference runtime engineer to push the boundaries of what's possible in LLM and diffusion model serving. Models grow larger. Architectures shift: mixture-of-experts, multimodal, agentic. Every breakthrough demands innovations on the inference engine itself. You'll work at the core of vLLM, optimizing how models execute across diverse hardware and architectures. Your work will directly impact how the world runs AI inference.

Requirements

  • Minimum qualifications:
  • Bachelor's degree or equivalent experience in computer science, engineering, or similar.
  • Deep understanding of transformer architectures and their variants.
  • Strong programming skills in Python with experience in PyTorch internals.
  • Experience with LLM inference systems (vLLM, TensorRT-LLM, SGLang, TGI).
  • Ability to read and implement model architectures and inference techniques from research papers.
  • Demonstrate the ability to contribute performant and maintainable code and debug in complex ML codebases.
  • Preferred qualifications:
  • Deep understanding of KV-cache memory management, prefix caching, and hybrid model serving.
  • Familiarity with RL frameworks and algorithms for LLMs.
  • Experience with multimodal inference (audio/image/video/text).
  • Contributions to open-source ML or system infrastructure projects.
  • Bonus points if you have:
  • Implemented core features in vLLM or other inference engine projects.
  • Contributed to vLLM integrations (verl, OpenRLHF, Unsloth, LlamaFactory, etc).
  • Written widely-shared technical blogs or side projects on vLLM or LLM inference.
  • Logistics
  • Location: Fully remote, worldwide. We're timezone-flexible but expect regular overlap with Pacific Time for critical syncs.
  • Compensation: We offer competitive compensations (salary + equity) compared to the local market conditions.
  • Visa sponsorship: We sponsor visas on a case-by-case basis.
  • Benefits: Inferact offers competitive benefits appropriate to your location, including health coverage where applicable.

Where you’d work

Fully remote

You can work from

  • United States

Visa sponsored

About the company

Inferact

  • Industry: AI

Also hiring in San Francisco, United States

2 of their 12 open roles are remote

Your chances

Worth a look before you spend an evening tailoring a CV for it.

  • 20 checks run
  • 3 red flags

Still hiring?

13 checks

1 red flag

How crowded?

7 checks

2 red flags

Fits Me

How well does this role fit you?

Answer a few questions or drop your CV, and every role gets a fit score with the reasons, this one first.

  • Your field
  • Level
  • Stack
  • Work model
  • Salary floor
  • Must-haves
Details13 facts · Role, Location, Compensation, Employment, Company
Tech stack
  • Python
  • PyTorch
  • LLMs
Seniority
Staff
Type
Full-time
Equity
Equity offered
Industry
AI
Region
United States
Show 7 more factsShow less

Role

Category
Development
Seniority
Staff
Experience
6+ years
Tech stack
  • Python
  • PyTorch
  • LLMs

Location

Work model
Remote
Region
United States
Remote from
  • United States
Visa sponsorship
Sponsored

Compensation

Salary
Salary by agreement
Pay period
Annual
Equity
Equity offered

Employment

Type
Full-time

Company

Industry
AI

Something wrong with this vacancy?

Similar vacancies

  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Staff Software Engineer

    Salary by agreement

    • Remote · Europe, Portugal
    • Senior
  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Software Engineer Intern

    Salary by agreement

    • Hybrid · Romania, Berlin
    • Junior
  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Senior Software Engineer

    Salary by agreement

    • Remote
    • Senior
  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Software Engineer

    Salary by agreement

    • Hybrid · Toronto
    • Junior

Share this vacancy

What's wrong with it?

The employer never sees who reported.

Reason