Description
Read This First
We work with unusual intensity. In-person in San Francisco, six days a week, long days, most weekends. This is not a phase we'll grow out of. It's how we've chosen to build, because we're in a market where speed decides who wins.
We're telling you this in the first paragraph, not the last, because we only want people who read that and feel pulled in, not talked into it. If you want a 9-to-5 (genuinely, no judgment), this isn't your role, and we'd rather you know now.
Here's what you get in exchange:
Top-of-market cash. We don't pay average salaries and ask for extraordinary hours. The comp reflects the commitment.
Meaningful equity you can believe in. Significant grants, employee-friendly terms, and our intent to create liquidity opportunities as we raise.
Zero commute tax. We support you living close to the office, with dinner at the office every night. Your hours go into building, not commuting.
Founders in the trenches. We work the same schedule we ask of you. This is a shared war, not extraction.
Compression of a decade into two years. You'll ship more, own more, and grow faster here than anywhere paying you to coast.
About the company
The company is building the voice AI engineer. Teams use the company to test agents before going live, monitor real production calls, and self-improve continuously. The company doesn't just flag issues and suggest fixes: it reproduces failures in simulation, fixes them, tests the fix thoroughly, and raises PRs. The platform spans pre-production simulation, LLM-powered evaluation, adversarial red-teaming, production monitoring with live drift detection, and cross-provider benchmarking (Vapi, Retell, Pipecat, LiveKit, ElevenLabs, and more).
About the Role
You'll build the core of the company: the simulation engines, evaluation systems, self-improvement loops, and observability pipelines our customers rely on to ship voice agents with confidence.
We deliberately don't split this into "software engineer" vs. "AI engineer." The interesting problems live at the boundary: real-time voice infrastructure meets LLM-as-judge evaluation, distributed systems meet RL-style self-improvement loops, telephony meets audio and speech analysis (ASR quality, barge-in, latency, prosody), and classic NLP meets frontier agentic behavior. You'll work across that whole surface.