You and the process

With a person
Anna, in-house recruiter
9 years hiring engineers
30 minutes with a real recruiter
They read your CV with you, on a call, and say where the offers are being lost.
Didn’t find what you were looking for? Tell us what to build
Be the first to open itNo views yet
Company hidden

Founding Engineer, Data & AI

  • Office
  • 3-6 years

Salary

$150,000 - 250,000/ year

AI summary

For members

The whole posting in a few lines. Sign up to read it here and on every role you open.

Sign Up to Read

Description

Full-time. Founding team. New York City, in-person required. Reports to the Co-founder & CTO and works directly with both founders and domain experts.

About the company

We bring that information together so teams can understand the evidence, apply their own judgment and preserve why they made a decision.

Our ambition is a world model for corporate governance: a system that connects institutional knowledge, policies, decisions and outcomes. Getting there starts with reliable data and software customers trust in their daily work.

Responsibilities

  • Build and own the data systems that make the company useful and trustworthy.
  • This is a backend- and data-intensive applied AI role. You should be as comfortable debugging a production pipeline and evolving a database schema as evaluating a document-extraction model. You will build on an existing codebase, working with messy filings and customer documents and making the results dependable enough for customers to use. The work includes asynchronous jobs, versioned evidence and safe reprocessing, not just prompts and model experiments.
  • Your first mandate is one prioritized pipeline and its downstream use. You will establish what good looks like with domain experts, improve the system, and own it in production. As one of our first engineering hires, you will also help shape how we build, test and operate software.
  • What you'll own
  • Build and improve ingestion and extraction pipelines for public filings, external feeds and customer documents. Choose structured sources, deterministic parsing or models based on the task. Make jobs safe to retry, prevent duplicate or stale-worker writes, recover from partial failures, and make missing or stale coverage visible.
  • Design data models for the entities and relationships your pipeline needs. Preserve source evidence, effective dates and version history so past outputs remain reproducible as data, policies and models change. Distinguish verified facts, model outputs, human decisions and missing information.
  • Build evaluations with domain reviewers using independently human-reviewed reference cases. Measure completeness, citation support, consequential errors and cases requiring human judgment, not just aggregate agreement. Treat model-generated references as proposals, not ground truth, and turn production failures into regression tests.
  • Provide reliable services for research, policy application and recommendations. Make reprocessing safe: retain customer corrections, identify affected outputs and flag results that need renewed review.
  • Protect customer information throughout ingestion, storage, evaluation and reprocessing. Respect tenant boundaries and permitted data uses. Do not assume one customer's documents or decisions can be reused for another.
  • Operate what you build. Own monitoring, diagnosis and tested recovery for your services. Manage provider usage, caching and processing costs without hiding gaps or sacrificing quality. Reduce manual intervention and document enough that another engineer can release and support them.
  • You will work closely with our product engineering counterpart, who owns the customer workflow and how people inspect, correct and use the outputs. You own the underlying pipeline and its data-quality contract. Agree on interfaces and failure behavior together, and follow issues through to resolution rather than stopping at a handoff.
  • The founders set priorities and resolve tradeoffs. Founders and domain reviewers provide policy interpretation and reviewed reference answers; you are not expected to invent governance rules yourself. We will scope your initial work so you can make progress even if the other role has not yet been filled.
  • Problems you might tackle
  • Two disclosures disagree about a director's committee membership. Determine which facts apply at which dates, preserve both sources and make the discrepancy reviewable.
  • A model update improves average extraction accuracy but introduces a consequential error. Catch it with an evaluation and make a defensible release decision.
  • A new filing arrives after an analyst has approved an output. Update the relevant facts without losing their corrections or silently carrying approval over to changed results.
  • Your first 90 days
  • By day 30: Trace one agreed pipeline from source to customer output, ship an improvement, and establish quality, coverage and operating baselines with domain reviewers.
  • By day 60: Own that pipeline in production, with reviewed test cases, monitoring, safe reprocessing and a tested recovery procedure.
  • By day 90: Demonstrate an agreed improvement in reliability, coverage or review effort without sacrificing the quality baseline. Routine releases and diagnosis no longer need a founder to direct each step, and another engineer can operate the service using your documentation.
  • We will choose the initial pipeline and success measures together based on customer needs and the existing system.
  • What you bring
  • Strong backend engineering, SQL and Postgres fundamentals, with experience operating production services or data pipelines.
  • Practical experience with asynchronous or distributed processing: idempotency, retries, leases and fencing, concurrency, and recovery when only part of a job succeeds.
  • Experience turning messy documents or other imperfect source data into structured information people depend on. Here that means public company filings, third-party data feeds and customer spreadsheets.
  • Practical experience shipping model-based systems and improving them through evaluations against human-reviewed references, user feedback and failure analysis, with the observability to explain what a model did on a given input.
  • Sound judgment about schema evolution, entity resolution, historical records and data provenance.
  • The ability to become productive in an existing codebase, debug across services and storage systems, and evolve interfaces or schemas without breaking their consumers. Our services are TypeScript and Go over Postgres and BigQuery.
  • The independence to narrow an ambiguous problem, ask for the context you need and carry the work through production use.
  • Experience with financial or legal documents and human-review tools is useful. Governance expertise and model-training research are not prerequisites. We care more about systems you have made reliable than a particular language, model or framework.
  • Our stack
  • TypeScript services on NestJS and Go services over gRPC, with Postgres as the system of record and BigQuery for analytical and document data. We use Anthropic and OpenAI models with LangSmith tracing, run on Google Cloud with Kubernetes and Terraform, and observe with Datadog and Sentry. CI enforces coverage gates. Python is useful for eval and analysis work.

Conditions

We work in person in New York and stay close to customers. We prototype quickly, use AI development tools where they help, and remain responsible for what we ship.

We narrow scope before compromising correctness, permissions or customer trust. We test representative cases and failure paths, observe what happens after release, and flag risks early. Founding engineers have room to make decisions and are expected to make their reasoning understandable to the team.

Apply

Send a short note and examples of systems you have built. Describe a difficult data or model failure, how you traced it to its source, what you changed and how you verified the fix. Confidential work can be discussed without sharing proprietary code or customer data.

Skills: Go, Google Cloud, Kubernetes, Node.js, PostgreSQL, Python, TypeScript, SQL, Docker, ETL, gRPC, Terraform, LLMs, AI Agents, LangChain, BigQuery, Generative AI, NestJS, OpenAI, LangSmith, Datadog

Experience: 3+ years

Visa: US citizen/visa only

Where you’d work

From the office

No visa sponsorship

You must already be able to work in the United States

About the company

Company hidden

  • Industry: SaaS

Office in New York, United States

Your chances

Still hiring, not crowded yet, and a person reads your message.

  • 19 checks run
  • 8 good signs
  • 0 red flags

Still hiring?

12 checks

Actively hiring

In its favour4

  • Still on the company's own careers site, checked 2 h agoModerate evidence
  • Specific about the basics: pay, place, level, stack and contract all statedSlight evidence
  • States its salarySlight evidence
1 moreFewer
  • A hiring contact is attached to itSlight evidence

How crowded?

7 checks

Low

In its favour4

  • You can message the hiring contact and skip the queueModerate evidence
  • In the office in New York: only people nearby can take itSlight evidence
  • Only for people already authorized to work in United StatesSlight evidence
1 moreFewer
  • Asks for BigQuery, which fewer than 1% of open roles doSlight evidence

Fits Me

How well does this role fit you?

Answer a few questions or drop your CV, and every role gets a fit score with the reasons, this one first.

  • Your field
  • Level
  • Stack
  • Work model
  • Salary floor
  • Must-haves
Details13 facts · Role, Location, Compensation, Employment, Company
Tech stack
  • Python
  • SQL
  • PostgreSQL
  • Google Cloud
  • Kubernetes
  • Docker
  • BigQuery
  • LLMs
  • LangChain
  • TypeScript
  • Node.js
  • Go
  • Terraform
  • JavaScript
  • gRPC
Type
Full-time
Industry
SaaS
Specialty
Data Engineering
Region
United States
Pay period
Annual
Show 7 more factsShow less

Role

Category
Data & Analytics
Specialty
Data Engineering
Experience
3+ years
Tech stack
  • Python
  • SQL
  • PostgreSQL
  • Google Cloud
  • Kubernetes
  • Docker
  • BigQuery
  • LLMs
  • LangChain
  • TypeScript
  • Node.js
  • Go
  • Terraform
  • JavaScript
  • gRPC

Location

Work model
Office
Region
United States
Office
  • New York, United States
Visa sponsorship
Not sponsored
Must already work in
  • United States

Compensation

Salary
$150,000 - 250,000 / year
Pay period
Annual

Employment

Type
Full-time

Company

Industry
SaaS

Something wrong with this vacancy?

Similar vacancies

  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Data Engineer

    $345,000 - 385,000 / year

    • Hybrid · Seattle
    • Mid-Level
  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Lead Data Engineer

    $110,000 - 180,000 / year

    • Office · US
    • Lead & Manager
    Direct apply
  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Senior Data Engineer

    $200,000 - 250,000 / year

    • Office · San Francisco
    • Senior
  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Technical Lead, Data Platform Engineer

    Salary by agreement

    • Hybrid · London
    • Lead & Manager

Share this vacancy

What's wrong with it?

The employer never sees who reported.

Reason