Still hiring?
14 checksNo red flags
Browse
All Tech JobsThe whole board, newest first.Roles That Fit MeAnswer a few questions, see your matches.Early WindowFound before the big boards.Direct ApplyStraight to the manager, past the ATS.By specialty
Your materials
CV AnalyzerWhat an ATS sees, and what to fix.Tailor CVBrought in line with one posting.Cover LetterWritten from your CV and the role.You and the process
Hey, I’m Wayjo. I find roles before the big boards.
Free to browse. An account unlocks the rest.
Jobs
All Tech JobsThe whole board, newest first.Roles That Fit MeAnswer a few questions, see your matches.Early WindowFound before the big boards.Direct ApplyStraight to the manager, past the ATS.$199,000 - 270,000/ year
The whole posting in a few lines. Sign up to read it here and on every role you open.
Sign Up to ReadThe Perception team is pioneering the development of a multi-modality foundation model to drive the next generation of autonomous system intelligence.
As a Perception Deployment Engineer, you will focus on bringing highly efficient, production-ready large-scale models to our on-vehicle stack. We are looking for experts with hands-on experience in compressing, accelerating, and deploying complex computer vision or foundation models for power- and thermal-constrained vehicle SOCs. You will optimize the ML models, write custom CUDA kernels, and build highly concurrent inference code to ensure real-time, deterministic execution on edge devices.
In this role, you will:
Design and develop production-level, low latency, and memory-safe C++ and CUDA code for real-time perception algorithms on vehicle systems.
Optimize large-scale models (Multi-Modal Sensor Fusion models, LLMs, VLMs) using advanced quantization (PTQ, QAT), pruning, mixed-precision inference frameworks.
Architect and implement model conversion and compilation pipelines using TensorRT for edge deployment.
Perform rigorous parity checking, accuracy recovery, and latency benchmarking between PyTorch frameworks and compiled edge binaries.
Develop and optimize custom ML OPs and TensorRT Plugins with efficient CUDA kernels to minimize latency and maximize memory bandwidth on AI accelerators.
Part of the week in the office
Zoox
Offices in United States, Boston, United States, Seattle, United States, San Diego, United States
We've checked whether it's still hiring and how crowded it is.
Still hiring?
14 checksNo red flags
How crowded?
7 checksNo red flags
How well does this role fit you?
Answer a few questions or drop your CV, and every role gets a fit score with the reasons, this one first.
Something wrong with this vacancy?