You and the process

With a person
Anna, in-house recruiter
9 years hiring engineers
30 minutes with a real recruiter
They read your CV with you, on a call, and say where the offers are being lost.
Didn’t find what you were looking for? Tell us what to build
Be the first to open itNo views yet
Exa

Research Engineer, Content Understanding

  • Office
  • 0-3 years

Salary

$180,000 - 350,000/ year

AI summary

For members

The whole posting in a few lines. Sign up to read it here and on every role you open.

Sign Up to Read

Description

Exa is an applied AI lab building a search engine unlike the world has ever seen. We build massive-scale infra to crawl the entire web, train state-of-the-art embedding models to process it, and design super high performant vector databases to retrieve over it. We now power search for Cursor, Cognition, HubSpot, and over 400,000 developers and have raised $350m from Lightspeed, Benchmark, and a16z.

Our ultimate goal is to build perfect search over all the world's information, far beyond Google. If you want to build massive-scale ML systems that will define the way the new AI world consumes information, this is the place for you.

As a backend engineer, you'd play a critical role in our search architecture. We're pretty flexible on what projects people work on based on their skills and interests.

Search quality is bounded by what we understand about a page. Before anything can be retrieved, something has to work out what the page actually says. That means parsing it into the parts that are content and the parts that are furniture, classifying what kind of page it is and what it is about, telling whether the page is usable at all, extracting when it was published, judging how good it is and whether it can be trusted, and working out whether it says anything that a page we already have does not. All of this has to work on every page on the web, in every language, in every shape the web comes in.

Some of this is classic document understanding. Some of it is much more open. Credibility and misinformation, AI-generated and machine-spun content, and pages written to be found rather than read are all unsolved, and search results are only as trustworthy as our answers to them.

We are looking for a research engineer to work on this. There is a lot of room to do it well.

Desired Experience

Graduate-level ML experience (Master’s or PhD with at least 2 years of relevant experience), or an exceptionally strong undergrad

You can build a transformer from scratch in PyTorch, and you have trained models that then had to be cheap enough to run everywhere

You like building large-scale datasets and living in the data. Most of the wins here are in the supervision rather than the architecture

You are comfortable with problems where the ground truth does not exist yet and defining it is part of the job

You care about the problem of finding high quality knowledge and recognize how important this is for the world

Example Projects

Make parsing work on the pages where it currently does not, and prove the improvement rather than assert it

Teach a model to judge page quality, and get everyone to agree on what quality means well enough to supervise it

Work on credibility and misinformation as a modelling problem: what a page claims, whether it is a reliable source of it, and whether it was written for a reader or for a crawler

Decide whether two documents are semantically the same or genuinely different, so we can deduplicate the web without collapsing pages that a user would want to see separately

Build classification and extraction that is accurate at web scale and cheap enough to run on all of it

Design the supervision for something nobody has labels for, and find out whether it is learnable at all

Trace a bad search result back to the page-level prediction that caused it, and fix it at the source

Where you’d work

From the office

The office

About the company

Exa

  • Industry: AI

Office in San Francisco, United States

Your chances

Worth a look before you spend an evening tailoring a CV for it.

  • 21 checks run
  • 2 red flags

Still hiring?

14 checks

No red flags

How crowded?

7 checks

2 red flags

Fits Me

How well does this role fit you?

Answer a few questions or drop your CV, and every role gets a fit score with the reasons, this one first.

  • Your field
  • Level
  • Stack
  • Work model
  • Salary floor
  • Must-haves
Details11 facts · Role, Location, Compensation, Employment, Company
Tech stack
  • HubSpot
  • PyTorch
  • Vector databases
Type
Full-time
Industry
AI
Specialty
ML
Region
United States
Pay period
Annual
Show 5 more factsShow less

Role

Category
Data & Analytics
Specialty
ML
Experience
2+ years
Tech stack
  • HubSpot
  • PyTorch
  • Vector databases

Location

Work model
Office
Region
United States
Office
  • San Francisco, United States

Compensation

Salary
$180,000 - 350,000 / year
Pay period
Annual

Employment

Type
Full-time

Company

Industry
AI

Something wrong with this vacancy?

Similar vacancies

  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Data & AI Engineer Intern

    Salary by agreement

    • Hybrid · Paris
    • Junior
  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    AI Engineer

    $80,000 - 210,000 / year

    • Office · CA, Austin
    • Senior
  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Senior AI/ML Engineer

    $200,000 - 260,000 / year

    • Office · San Francisco
    • Senior
  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Member of Technical Staff, ML Infra

    Salary by agreement

    • Office · San Francisco
    • Senior

Share this vacancy

What's wrong with it?

The employer never sees who reported.

Reason