You and the process

With a person
Anna, in-house recruiter
9 years hiring engineers
30 minutes with a real recruiter
They read your CV with you, on a call, and say where the offers are being lost.
Didn’t find what you were looking for? Tell us what to build
Be the first to open itNo views yet
Company hidden

AI/ML Engineer

  • Office

Salary

$120,000 - 200,000/ year

AI summary

For members

The whole posting in a few lines. Sign up to read it here and on every role you open.

Sign Up to Read

Description

The company is production monitoring for AI agents. We catch the silent failures your observability tools and evals miss (think bad tool calls, lost context, and infinite loops) before your users find them.

Why this role exists

Agents break silently. They call the wrong tool, forget what the user said three turns ago, and loop until someone pulls the plug. Soon they'll be responsible for the majority of the world’s economic work, and most teams won't even know when they fail.

Making agents reliable is the problem the company exists to solve. That means catching the unknown unknowns, the failures nobody thought to write an eval for, and closing the loop end to end so they get fixed, not just flagged. It's the foundation for building agents people can actually trust and the future of self-improving systems.

The hardest part of our product is deciding what counts as a failure.\ \ There is no ground truth here, and no benchmark to climb. Every customer's agent is different, what "wrong" means changes from one to the next, and we have to get it right across production without anyone telling us what to look for. Being confidently wrong often costs us more trust than being right fifty times earns.

This role owns the intelligence in the loop: what we flag, how sure we are, and whether the fix we propose actually fixes it.

Responsibilities

  • Own detection quality. Find the failures that don't look like failures: compliant but wrong, omissions, and patterns that only show up across thousands of traces
  • Turn implicit signals into evidence. Rephrasing, abandonment, retries, and the other ways users tell you something broke without saying so
  • Build the evals for our own system. If we can't measure precision on a problem with no labels, we can't improve it
  • Make patch generation trustworthy. Reproduce the failure, verify the fix, and know when to stay quiet instead of opening a bad PR
  • Keep it affordable. LLM-as-judge on every event is easy. Doing it at a cost per event that doesn't eat the business is the actual job
  • Read real customer traces every week. The best ideas here come from staring at production, not papers

Requirements

  • High slope over years of experience. New grads and dropouts welcome
  • A track record of shipping things people actually use
  • Hands-on with LLMs in production: evals, LLM-as-judge, embeddings, and knowing when a smaller model or no model is the right call
  • Real research taste. You can tell a real improvement from noise, even when there's no ground truth to check against
  • Bonus: you were the customer once. You ran agents in production and got burned
  • Who you’ll work with
  • You'll join a team of dropout founders and engineers from Amazon, Together, and Zoom. We've been founding operators at unicorns and at startups that went on to be acquired.
  • Onsite in San Francisco. We don’t sponsor visas.
  • Experience: Any (new grads ok)
  • Visa: US citizen/visa only

Where you’d work

From the office

The office

No visa sponsorship

You must already be able to work in the United States

About the company

Company hidden

Office in San Francisco, United States

Your chances

Still hiring, not crowded yet, and a person reads your message.

  • 19 checks run
  • 8 good signs
  • 0 red flags

Still hiring?

12 checks

Actively hiring

In its favour4

  • Still on the company's own careers site, checked 3 h agoModerate evidence
  • Posted 4 days ago: newer than 92% of open rolesModerate evidence
  • States its salarySlight evidence
1 moreFewer
  • A hiring contact is attached to itSlight evidence

How crowded?

7 checks

Low

In its favour4

  • You can message the hiring contact and skip the queueModerate evidence
  • In the office in San Francisco: only people nearby can take itSlight evidence
  • Only for people already authorized to work in United StatesSlight evidence
1 moreFewer
  • Pays below most similar roles, which thins the crowdSlight evidence

Fits Me

How well does this role fit you?

Answer a few questions or drop your CV, and every role gets a fit score with the reasons, this one first.

  • Your field
  • Level
  • Stack
  • Work model
  • Salary floor
  • Must-haves
Details11 facts · Role, Location, Compensation, Employment
Tech stack
  • LLMs
Type
Full-time
Specialty
ML
Region
United States
Pay period
Annual
Show 6 more factsShow less

Role

Category
Data & Analytics
Specialty
ML
Tech stack
  • LLMs

Location

Work model
Office
Region
United States
Office
  • San Francisco, United States
Visa sponsorship
Not sponsored
Must already work in
  • United States

Compensation

Salary
$120,000 - 200,000 / year
Pay period
Annual

Employment

Type
Full-time

Something wrong with this vacancy?

Similar vacancies

  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Data & AI Engineer Intern

    Salary by agreement

    • Hybrid · Paris
    • Junior
  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    AI Engineer

    $80,000 - 210,000 / year

    • Office · CA, Austin
    • Senior
  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Senior AI/ML Engineer

    $200,000 - 260,000 / year

    • Office · San Francisco
    • Senior
  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Member of Technical Staff, ML Infra

    Salary by agreement

    • Office · San Francisco
    • Senior

Share this vacancy

What's wrong with it?

The employer never sees who reported.

Reason