You and the process

With a person
Anna, in-house recruiter
9 years hiring engineers
30 minutes with a real recruiter
They read your CV with you, on a call, and say where the offers are being lost.
Didn’t find what you were looking for? Tell us what to build
Be the first to open itNo views yet
Inferact

Member of Technical Staff, Forward Deployed Engineer

  • Office
  • 6+ years

Salary

$200,000 - 400,000/ year

AI summary

For members

The whole posting in a few lines. Sign up to read it here and on every role you open.

Sign Up to Read

Description

Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware, a position that took years to build.

About the Role

We're looking for a Forward Deployed Engineer to make Inferact successful inside real customer environments. You'll work directly with customers to deploy, integrate, debug, and optimize vLLM-powered inference systems across cloud, Kubernetes, GPU, networking, and model-serving environments.

This is a hands-on engineering role, not a traditional pre-sales position. You'll move from architecture discussions to implementation, own difficult production problems end-to-end, and work closely with core product and engineering teams to turn what you learn in the field into reusable product capabilities. Your work will directly affect customer time-to-value and how Inferact's platform evolves.

Requirements

  • Minimum qualifications:
  • Bachelor's degree or equivalent experience in computer science, engineering, systems, machine learning, or similar.
  • Strong software engineering ability in Python, Go, TypeScript, or similar, with experience building production-quality integrations, tooling, services, automation, or prototypes.
  • Hands-on experience deploying or operating ML systems, model serving, AI infrastructure, cloud platforms, Kubernetes, or high-scale backend systems in production.
  • Ability to work directly with sophisticated customer engineering teams, understand ambiguous technical environments, and personally drive implementations and debugging to resolution.
  • Strong systems debugging skills across application, runtime, infrastructure, networking, identity, storage, observability, and distributed-system boundaries.
  • Ability to reason about latency, throughput, batching, model/runtime compatibility, scaling, reliability, and cost tradeoffs in production inference environments.
  • High ownership and strong technical communication, with the judgment to distinguish one-off customer work from problems that should become reusable product capabilities.
  • Preferred qualifications:
  • Experience with vLLM, SGLang, TensorRT-LLM, TGI, Ray Serve, BentoML, or other LLM inference and model-serving systems.
  • Experience with NVIDIA or AMD GPUs, CUDA / ROCm, GPU scheduling, multi-GPU serving, or accelerator-backed infrastructure.
  • Experience deploying infrastructure software into enterprise, regulated, security-sensitive, or bring-your-own-cloud environments.
  • Experience building APIs, SDKs, CLIs, developer tooling, deployment platforms, control planes, or infrastructure products used by technical teams.
  • Experience profiling latency, throughput, concurrency, GPU utilization, bottlenecks, and performance regressions.
  • Bonus points if you have:
  • Contributed to open-source ML systems, inference infrastructure, cloud infrastructure, Kubernetes, or developer tooling.
  • Worked in a forward-deployed, customer engineering, field engineering, or highly technical solutions role where you personally wrote and shipped code.
  • Built deployment playbooks, reference architectures, automation, or tooling that materially reduced customer time-to-production.
  • Resolved severe customer-facing production issues that crossed multiple technical layers and required close partnership with core engineering.
  • Turned repeated customer problems into reusable product features, abstractions, documentation, or platform improvements.
  • Logistics
  • Location: This role is based in San Francisco, California. Will consider remote in the US for exceptional candidates.
  • Compensation: Depending on background, skills, and experience, the expected annual salary range for this position is $200,000 - $400,000 USD + equity.
  • Visa sponsorship: We sponsor visas on a case-by-case basis.
  • Benefits: Inferact offers generous health, dental, and vision benefits as well as 401(k) company match.

Where you’d work

From the office

The office

Visa sponsored

About the company

Inferact

  • Industry: AI

Office in San Francisco, United States

2 of their 12 open roles are remote

Your chances

Worth a look before you spend an evening tailoring a CV for it.

  • 22 checks run
  • 2 red flags

Still hiring?

14 checks

No red flags

How crowded?

8 checks

2 red flags

Fits Me

How well does this role fit you?

Answer a few questions or drop your CV, and every role gets a fit score with the reasons, this one first.

  • Your field
  • Level
  • Stack
  • Work model
  • Salary floor
  • Must-haves
Details14 facts · Role, Location, Compensation, Employment, Company
Tech stack
  • TypeScript
  • Python
  • Go
  • Kubernetes
  • LLMs
  • CUDA
Seniority
Staff
Type
Full-time
Equity
Equity offered
Industry
AI
Specialty
Forward Deployed
Show 8 more factsShow less

Role

Category
Solutions & Support
Specialty
Forward Deployed
Seniority
Staff
Experience
6+ years
Tech stack
  • TypeScript
  • Python
  • Go
  • Kubernetes
  • LLMs
  • CUDA

Location

Work model
Office
Region
United States
Office
  • San Francisco, United States
Visa sponsorship
Sponsored

Compensation

Salary
$200,000 - 400,000 / year
Pay period
Annual
Equity
Equity offered

Employment

Type
Full-time

Company

Industry
AI

Something wrong with this vacancy?

Similar vacancies

  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Deployment Lead

    $230,000 - 295,000 / year

    • Hybrid · San Francisco, New York City
    • Lead & Manager
  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    ITAV Deployment Engineer

    $185,000 - 210,000 / year

    • Hybrid · San Francisco
  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Forward Deployed Engineer

    $190,000 - 210,000 / year

    • Remote · US
  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Forward Deployed Creative

    Salary by agreement

    • Hybrid · London

Share this vacancy

What's wrong with it?

The employer never sees who reported.

Reason