You and the process

With a person
Anna, in-house recruiter
9 years hiring engineers
30 minutes with a real recruiter
They read your CV with you, on a call, and say where the offers are being lost.
Didn’t find what you were looking for? Tell us what to build
Be the first to open itNo views yet
Thinking Machines Lab

Software Engineer, Inference

  • Office

Salary

$300,000 - 400,000/ year

AI summary

For members

The whole posting in a few lines. Sign up to read it here and on every role you open.

Sign Up to Read

Description

About Thinking Machines

The mission of Thinking Machines is to build AI that extends human will and judgment. We are training frontier models with Inkling, developing Tinker to let people make models their own, and crafting interfaces that broaden human-AI communication. We believe the future worth building is human, and we're hiring people who want to build it.

About the Role

We're hiring a Software Engineer, Inference to own the reliability, scale, and efficiency of the systems that serve our models to real users. Our research and inference teams push the limits of model performance and serving efficiency; this role makes sure those gains reach production safely and stay up — powering Tinker's live, multi-tenant serving and the products built on top of our models.

This is a production-facing systems role at the center of the company. You'll be the bridge between cutting-edge inference techniques and the day-to-day reality of serving real traffic: rollouts, capacity, incidents, and everything that keeps a fast-growing platform online.

Responsibilities

  • Operate and scale the production inference systems that serve live traffic, including Tinker's multi-tenant serving platform
  • Own the rollout process for new models, model versions, and inference optimizations, ensuring safe, incremental deployment to production
  • Build and improve observability, alerting, and capacity planning so the team can detect, diagnose, and resolve production issues quickly
  • Partner with inference and research teams to productionize new serving techniques without compromising reliability
  • Lead incident response for production inference issues, driving root cause analysis and durable fixes
  • Design for graceful degradation, failover, and redundancy so that serving stays resilient as usage grows
  • Manage capacity and cost tradeoffs for serving infrastructure as traffic and model sizes scale

Requirements

  • Minimum Qualifications
  • Experience operating large-scale, latency-sensitive production systems
  • Proficiency in Python and Go or another systems language
  • Experience with observability, monitoring, and incident response for production services
  • Strong understanding of distributed systems and how they fail at scale
  • Preferred Qualifications
  • Experience running production inference for large language models or other large-scale ML systems
  • Experience with deployment and rollout systems, such as canarying, blue/green deploys, or feature flags
  • Experience with capacity planning and cost optimization for GPU or TPU infrastructure
  • Familiarity with inference-specific techniques, such as batching, caching, or quantization, and their operational implications
  • Comfortable being on-call and leading incident response for critical production systems
  • Comfortable working with high autonomy in a fast-changing, early-stage environment
  • Logistics

Conditions

Compensation: Depending on background, skills and experience, the expected annual salary range for this position is $300,000 - $400,000 USD.

Visa sponsorship: We sponsor visas. While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together.

Benefits: Thinking Machines offers generous health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.

Where you’d work

From the office

Relocation offered

Visa sponsored

About the company

Thinking Machines Lab

  • Industry: AI

Offices in San Francisco, United States, New York, United States

Your chances

Worth a look before you spend an evening tailoring a CV for it.

  • 23 checks run
  • 3 red flags

Still hiring?

14 checks

No red flags

How crowded?

9 checks

3 red flags

Fits Me

How well does this role fit you?

Answer a few questions or drop your CV, and every role gets a fit score with the reasons, this one first.

  • Your field
  • Level
  • Stack
  • Work model
  • Salary floor
  • Must-haves
Details11 facts · Role, Location, Compensation, Employment, Company
Tech stack
  • Python
  • Go
  • LLMs
Type
Full-time
Industry
AI
Region
United States
Pay period
Annual
Show 6 more factsShow less

Role

Category
DevOps & Infrastructure
Tech stack
  • Python
  • Go
  • LLMs

Location

Work model
Office
Region
United States
Offices
  • San Francisco, United States
  • New York, United States
Relocation
Offered
Visa sponsorship
Sponsored

Compensation

Salary
$300,000 - 400,000 / year
Pay period
Annual

Employment

Type
Full-time

Company

Industry
AI

Something wrong with this vacancy?

Similar vacancies

  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Lead IT Engineer

    Salary by agreement

    • Remote
    • Lead & Manager
  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Security Engineer, New Grad

    Salary by agreement

    • Office · Dublin
    • Junior
  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Site Reliability Engineer

    Salary by agreement

    • Remote
  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Senior Security Operations Analyst

    Salary by agreement

    • Hybrid · Wellington, Auckland
    • Senior

Share this vacancy

What's wrong with it?

The employer never sees who reported.

Reason