Still hiring?
13 checks1 red flag
Browse
All Tech JobsThe whole board, newest first.Roles That Fit MeAnswer a few questions, see your matches.Early WindowFound before the big boards.Direct ApplyStraight to the manager, past the ATS.By specialty
Your materials
CV AnalyzerWhat an ATS sees, and what to fix.Tailor CVBrought in line with one posting.Cover LetterWritten from your CV and the role.You and the process
Hey, I’m Wayjo. I find roles before the big boards.
Free to browse. An account unlocks the rest.
Jobs
All Tech JobsThe whole board, newest first.Roles That Fit MeAnswer a few questions, see your matches.Early WindowFound before the big boards.Direct ApplyStraight to the manager, past the ATS.Not stated
Similar roles pay $190K - 245K a year · our estimate
The whole posting in a few lines. Sign up to read it here and on every role you open.
Sign Up to ReadInferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware—a position that took years to build.
About the Role
We're looking for an inference runtime engineer to push the boundaries of what's possible in LLM and diffusion model serving. Models grow larger. Architectures shift: mixture-of-experts, multimodal, agentic. Every breakthrough demands innovations on the inference engine itself. You'll work at the core of vLLM, optimizing how models execute across diverse hardware and architectures. Your work will directly impact how the world runs AI inference.
Fully remote
Visa sponsored
Inferact
Also hiring in San Francisco, United States
2 of their 12 open roles are remote
Worth a look before you spend an evening tailoring a CV for it.
Still hiring?
13 checks1 red flag
How crowded?
7 checks2 red flags
How well does this role fit you?
Answer a few questions or drop your CV, and every role gets a fit score with the reasons, this one first.
Something wrong with this vacancy?