Description
You'll build the evaluation systems that tell us whether the company actually works. That sounds simple. It isn't. Our core promise, convert any URL into clean, structured, LLM-ready data reliably, is hard to measure rigorously across millions of different websites, formats, and edge cases. As the systems we're measuring get more complex, the question "did that work?" gets harder, not easier.
This isn't an eval role where you inherit a framework and run benchmarks. You'll design the metrics, build the pipelines, generate the datasets, and own the feedback loop from output quality back to model and product decisions. If you care about what "good" actually means and have the engineering depth to measure it, this is the role.
Salary Range: $250,000–$290,000 USD/year (SF) / $210,000–$224,000 CAD/year (Toronto)
Equity Range: Competitive equity. Details shared during the process.
Location: San Francisco, CA (SF HQ) or Toronto, ON (Toronto Hub). Hybrid, onsite 3+ days a week.
Job Type: Full-Time
Experience: 4+ years in ML, research engineering, or data-heavy backend, with real evaluation work
Work Authorization: Must be authorized to work in the United States or Canada. We're not able to sponsor US visas right now. For Canada, we'll consider sponsorship on a case-by-case basis through our Toronto Hub.
About the company
The company is the easiest way to turn the web into data AI agents can use. One API call converts any URL into clean, LLM-ready markdown or structured data. It's the boring-hard problem everyone building with LLMs eventually hits, solved.
We hit 8 figures in ARR in year one and more than doubled it in year two. Growth like this is rare, and we're just getting started.
We're a small team punching far above our weight. Everyone here owns a real piece of the product and company, end to end, and runs it themselves. No hiding behind process or headcount.
This is a place for people who want to work at the frontier: an AI company building the infrastructure other AI companies run on, not one bolting AI onto an existing product. We move fast, go deep, and are building the tools superintelligence will rely on to gather data from the web. That library is called Alexandria, and it starts now.