You and the process

With a person
Anna, in-house recruiter
9 years hiring engineers
30 minutes with a real recruiter
They read your CV with you, on a call, and say where the offers are being lost.
Didn’t find what you were looking for? Tell us what to build
Be the first to open itNo views yet
Town

AI Engineer, Evals & Agent Quality

  • Office

Salary

$250,000 - 300,000/ year

AI summary

For members

The whole posting in a few lines. Sign up to read it here and on every role you open.

Sign Up to Read

Description

About Town

Town ( town.com ) is AI that starts from who you are. We build a persistent model of your identity, your voice, your judgment, your relationships, and your priorities, and use it to do real work on your behalf across every tool where you operate: email, calendar, documents, Slack, and more. Town doesn't wait for you to prompt it. It observes, learns, and acts. The more you use it, the more it becomes an extension of you.

Town was founded by Jean-Denis Greze (CEO), former CTO of Plaid, and Tony Vincent (CPO), former Director of Applied AI Product at Google. We're a small, talent-dense team backed by Andreessen Horowitz, Forerunner Ventures, First Round Capital, and Conviction, with more than $73M raised to date.

About the role

Town is building the most personalized, most capable AI assistant for everyone — one that knows you deeply, works across every tool you use, and gets sharper over time. Building the best assistant means proving it's the best: every model, prompt, and system change has to be measurably better, on every surface it touches.

That's what you'll own. You'll build the evals and quality systems that turn assistant performance into numbers the whole team can trust, measuring and improving the full multi-step trajectory the assistant takes to do real work. You'll build the model routing that puts the right model in the right place balancing cost, quality, and speed.

This is a foundational, 0→1 build with ownership to match: the eval framework, the golden datasets and labeling loop, model routing, and online measurement, and you set the bar for what "best" means at Town.

Responsibilities

  • Build a generalized eval system that measures assistant quality across every surface it touches — and, crucially, across multi-step agent trajectories.
  • Stand up golden datasets and the labeling loop that keeps them up to date and constantly checking to validate improvements and avoid regressions.
  • Build model routing and online evaluation tooling to help us learn and route to the best models.
  • Make every prompt and system change measurable, so the team can move fast without breaking what works.
  • Partner with engineers across the product to instrument quality and close the loop from signal to fix.
  • You might thrive here if you...
  • Have built or owned LLM eval systems, or offline/online quality measurement at scale.
  • Think rigorously about measurement. Maybe that came from an MLE or applied-ML background, maybe not, the instinct for how to measure "better" matters more than the exact pedigree.
  • Know the eval landscape hands-on, off-the-shelf tooling and eval frameworks, and have opinions on what to reach for when.
  • Are comfortable reasoning about model routing and the tradeoffs between models.
  • Ship the fixes, not just the dashboards and metrics.
  • Are a senior or staff engineer comfortable in greenfield, where the system doesn't exist yet.
  • Location
  • San Francisco, CA. Five days a week in person at our Financial District office.

Where you’d work

From the office, 4 or more days a week in the office

The office

About the company

Town

  • Industry: AI

Office in San Francisco, United States

Also hiring in Nyc, United States

Your chances

Worth a look before you spend an evening tailoring a CV for it.

  • 22 checks run
  • 3 red flags

Still hiring?

14 checks

2 red flags

How crowded?

8 checks

1 red flag

Fits Me

How well does this role fit you?

Answer a few questions or drop your CV, and every role gets a fit score with the reasons, this one first.

  • Your field
  • Level
  • Stack
  • Work model
  • Salary floor
  • Must-haves
Details11 facts · Role, Location, Compensation, Employment, Company
Tech stack
  • LLMs
Type
Full-time
Industry
AI
Specialty
ML
Region
United States
Pay period
Annual
Show 5 more factsShow less

Role

Category
Data & Analytics
Specialty
ML
Tech stack
  • LLMs

Location

Work model
Office
Region
United States
Office
  • San Francisco, United States
Days in the office
4 or more days

Compensation

Salary
$250,000 - 300,000 / year
Pay period
Annual

Employment

Type
Full-time

Company

Industry
AI

Something wrong with this vacancy?

Similar vacancies

  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Data & AI Engineer Intern

    Salary by agreement

    • Hybrid · Paris
    • Junior
  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    AI Engineer

    $80,000 - 210,000 / year

    • Office · CA, Austin
    • Senior
  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Senior AI/ML Engineer

    $200,000 - 260,000 / year

    • Office · San Francisco
    • Senior
  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Member of Technical Staff, ML Infra

    Salary by agreement

    • Office · San Francisco
    • Senior

Share this vacancy

What's wrong with it?

The employer never sees who reported.

Reason