You and the process

With a person
Anna, in-house recruiter
9 years hiring engineers
30 minutes with a real recruiter
They read your CV with you, on a call, and say where the offers are being lost.
Didn’t find what you were looking for? Tell us what to build
Early Window: Be the first to open itNo views yetCloses in
Company hidden

Senior Machine Learning Platform Engineer

  • Hybrid
  • 6+ years

Salary

$145,000 - 235,000/ year

AI summary

For members

The whole posting in a few lines. Sign up to read it here and on every role you open.

Sign Up to Read

Description

Role summary

We are hiring a Senior Machine Learning Platform Engineer to build the infrastructure and engineering workflows that take custom models from development into reliable production use. This is a hands-on role for someone who has hosted LLMs on GPU infrastructure, developed the cloud platform around model serving, and built MLOps pipelines that make releases repeatable and observable.

You will work with model developers, application engineers, security, and cloud platform teams to turn a model artifact into a secure, scalable inference service. The work spans serving architecture, infrastructure as code, deployment automation, model lifecycle management, and production operations across AWS and Azure, with potential integration into on-premises environments. You will also help shape agent workflows that use hosted models, so serving choices support the needs of multi-step applications. Success means teams can deploy, update, monitor, and troubleshoot models through clear, reusable platform patterns.

Responsibilities

  • Design and build hosting for custom ML and AI models across AWS and Azure, with particular focus on GPU-backed LLM inference and real-time endpoints; support batch inference where appropriate.
  • Package models and their dependencies into reproducible serving workloads; choose and implement suitable managed services, containers, or Kubernetes-based patterns based on throughput, latency, security, reliability, and cost.
  • Provision and configure Kubernetes clusters or other suitable hosting infrastructure, including compute, storage, API access, identity and access controls, secrets, networking integration, observability, and environment configuration.
  • Help design secure connections and deployment patterns between cloud and on-premises environments as hosting needs evolve.
  • Develop infrastructure as code and deployment automation so model-hosting environments can be provisioned, reviewed, promoted, and maintained consistently.
  • Build MLOps workflows for model registration, versioning, validation, release, rollback, and retirement. Connect training or model preparation to deployment through automated pipelines and appropriate quality gates.
  • Partner with application teams to design and prototype agent workflows, including model and tool orchestration, state handling, failure recovery, and evaluation.
  • Translate agent workload patterns into hosting decisions about model selection, context length, concurrency, latency, cost, tool access, and end-to-end tracing.
  • Establish production monitoring for service health, latency, throughput, errors, GPU and other resource use, and model behavior. Share responsibility for diagnosing incidents and improving capacity, reliability, and cost.
  • Create reusable deployment templates, reference architectures, documentation, and onboarding paths that help other teams ship models safely.
  • Partner with model and application teams on practical tradeoffs such as online versus batch inference, managed versus self-hosted serving, scaling, evaluation, data handling, and operational ownership.
  • Required experience
  • Hands-on experience hosting LLM inference on GPU infrastructure in a production environment. You can explain what you personally built, how models reached production, and how you managed throughput, latency, utilization, reliability, and cost.
  • Experience building the surrounding model-serving platform for custom models, such as inference runtimes, deployment patterns, endpoint access, scaling, and operational tooling.
  • Strong software engineering skills, especially Python, with experience building services, automation, and maintainable production code.
  • Experience developing infrastructure for ML workloads across AWS and Azure, with deep hands-on delivery in at least one and practical ability to work in the other. You have used infrastructure as code such as Terraform or an equivalent tool.
  • Experience provisioning, configuring, and maintaining Kubernetes or another production hosting platform for containerized inference workloads.
  • Experience building or operating ML deployment pipelines with versioned artifacts, automated validation, CI/CD, environment promotion, and rollback.
  • Familiarity with LLM agent patterns, including model invocation, tool calls, and multi-step workflows, and how they affect serving capacity, reliability, and access controls.
  • Working knowledge of production concerns for inference services: scaling, latency, availability, logging and metrics, shared incident response, access control, and cost.
  • Ability to work across model development, application, platform, and security teams; turn ambiguous requirements into a working design; and document the resulting operational approach.
  • Helpful experience
  • Advanced GPU inference optimization, including capacity planning, batching, autoscaling, memory use, model loading, and performance tuning.
  • Serving generative or other compute-intensive custom models beyond LLMs; selecting and tuning inference servers and runtime configurations.
  • AWS SageMaker or EKS; Azure Machine Learning or AKS; or equivalent managed and self-hosted model platforms.
  • Hybrid or on-premises model hosting, including connectivity, security boundaries, hardware constraints, and operational handoff.
  • Model registries, experiment tracking, data or feature pipelines, scheduled retraining, model evaluation, and drift or quality monitoring.
  • Secure enterprise deployment patterns such as private networking, IAM/RBAC, secrets management, auditability, and handling sensitive data.
  • Hands-on development of agent workflows, including tool integration, evaluation, tracing, or guardrails.
  • Who will thrive here
  • You are an engineer who has taken responsibility for what happens after a model is trained: how it is packaged, deployed, secured, scaled, observed, updated, and supported. You are comfortable writing code and infrastructure, investigating production failures, and making clear tradeoffs with partner teams.

Conditions

The pay range for this role is $147,050 to $230,850 USD annually with additional

opportunities for pay in the form of bonus and/or equity (applies to United

States of America candidates only). Pay varies by work location, job-related

knowledge, skills, and experience.

Benefits

  • HP offers a comprehensive benefits package for this position, including:
  • Health insurance
  • Dental insurance
  • Vision insurance
  • Long term/short term disability insurance
  • Employee assistance program
  • Flexible spending account
  • Life insurance
  • Generous time off policies, including;
  • 4-12 weeks fully paid parental leave based on tenure
  • 11 paid holidays
  • Additional flexible paid vacation and sick leave
  • US benefits overview https://hpbenefits.ce.alight.com/
  • The compensation and benefits information is accurate as of the date of this
  • posting. The Company reserves the right to modify this information at any time,
  • with or without notice, subject to applicable law.
  • Job -
  • Software
  • Schedule -
  • Full time
  • Shift -
  • No shift premium (United States of America)
  • Travel -
  • No
  • Relocation -
  • No

Where you’d work

Part of the week in the office

You can work from

  • United States

No relocation

About the company

Company hidden

Office in United States

Your chances

Still hiring, not crowded yet, and you'd be among the first.

  • 18 checks run
  • 7 good signs
  • 0 red flags

Still hiring?

12 checks

Actively hiring

In its favour5

  • Still on the company's own careers site, checked 2 h agoModerate evidence
  • Found in the last 48 hours, before the big job boardsModerate evidence
  • The company posted 59 roles in the last 2 weeksSlight evidence
2 moreFewer
  • Specific about the basics: pay, place, level, stack and contract all statedSlight evidence
  • States its salarySlight evidence

How crowded?

6 checks

Low

In its favour2

  • In its Early Window: not on the big job boards yetStrong evidence
  • Senior level: far fewer people qualifySlight evidence

Fits Me

How well does this role fit you?

Answer a few questions or drop your CV, and every role gets a fit score with the reasons, this one first.

  • Your field
  • Level
  • Stack
  • Work model
  • Salary floor
  • Must-haves
Details14 facts · Role, Location, Compensation, Employment
Tech stack
  • Python
  • AWS
  • Kubernetes
  • Azure
  • LLMs
Seniority
Senior
Type
Full-time
Equity
Equity offered
Specialty
ML
Region
United States
Show 8 more factsShow less

Role

Category
Data & Analytics
Specialty
ML
Seniority
Senior
Experience
6+ years
Tech stack
  • Python
  • AWS
  • Kubernetes
  • Azure
  • LLMs

Location

Work model
Hybrid
Region
United States
Office
  • United States
Remote from
  • United States
Relocation
Not offered

Compensation

Salary
$145,000 - 235,000 / year
Pay period
Annual
Equity
Equity offered

Employment

Type
Full-time

Something wrong with this vacancy?

Similar vacancies

  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Data & AI Engineer Intern

    Salary by agreement

    • Hybrid · Paris
    • Junior
  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    AI Engineer

    $80,000 - 210,000 / year

    • Office · CA, Austin
    • Senior
  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Senior AI/ML Engineer

    $200,000 - 260,000 / year

    • Office · San Francisco
    • Senior
  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Member of Technical Staff, ML Infra

    Salary by agreement

    • Office · San Francisco
    • Senior

Share this vacancy

What's wrong with it?

The employer never sees who reported.

Reason