Still hiring?
14 checksNo red flags
Browse
All Tech JobsThe whole board, newest first.Roles That Fit MeAnswer a few questions, see your matches.Early WindowFound before the big boards.Direct ApplyStraight to the manager, past the ATS.By specialty
Your materials
CV AnalyzerWhat an ATS sees, and what to fix.Tailor CVBrought in line with one posting.Cover LetterWritten from your CV and the role.You and the process
Hey, I’m Wayjo. I find roles before the big boards.
Free to browse. An account unlocks the rest.
Jobs
All Tech JobsThe whole board, newest first.Roles That Fit MeAnswer a few questions, see your matches.Early WindowFound before the big boards.Direct ApplyStraight to the manager, past the ATS.$180,000 - 350,000/ year
The whole posting in a few lines. Sign up to read it here and on every role you open.
Sign Up to ReadExa is an applied AI lab building a search engine unlike the world has ever seen. We build massive-scale infra to crawl the entire web, train state-of-the-art embedding models to process it, and design super high performant vector databases to retrieve over it. We now power search for Cursor, Cognition, HubSpot, and over 400,000 developers and have raised $350m from Lightspeed, Benchmark, and a16z.
Our ultimate goal is to build perfect search over all the world's information, far beyond Google. If you want to build massive-scale ML systems that will define the way the new AI world consumes information, this is the place for you.
We crawl the web continuously and cannot keep all of it, refresh all of it, or spend a GPU on all of it. Something has to decide which pages are worth indexing, which links are worth following, which pages say the same thing as pages we already have, what has gone stale, and what we should never have picked up. These decisions are worth more than most ranking work. The cheapest way to improve search is to stop indexing pages nobody should ever retrieve, and to start indexing the ones we are missing. Today they run on a mix of learned signals and thresholds someone picked.
We are looking for a research engineer to own the signals behind those decisions, and the decisions themselves.
Desired Experience
You are comfortable owning a problem with no ground truth, where the first job is defining what the right answer means
Hands-on ML experience with classifiers, rankers and calibration, plus strong data instincts at web scale
You think in end-to-end impact. A signal only counts if a decision changes and search gets better
You like working across teams. Crawling, indexing and retrieval all consume what you build
You care about the problem of finding high quality knowledge and recognize how important this is for the world
Example Projects
Make parsing work on the pages where it currently does not, and prove the improvement rather than assert it
Teach a model to judge page quality, and get everyone to agree on what quality means well enough to supervise it
Work on credibility and misinformation as a modelling problem: what a page claims, whether it is a reliable source of it, and whether it was written for a reader or for a crawler
Decide whether two documents are semantically the same or genuinely different, so we can deduplicate the web without collapsing pages that a user would want to see separately
From the office
Worth a look before you spend an evening tailoring a CV for it.
Still hiring?
14 checksNo red flags
How crowded?
7 checks1 red flag
How well does this role fit you?
Answer a few questions or drop your CV, and every role gets a fit score with the reasons, this one first.
Something wrong with this vacancy?