All open roles

ML Infrastructure Engineer

Engineering · Posted 3 months ago

We are looking for an ML Infrastructure Engineer with 3+ years of experience to own and scale the training and inference stack at a fast-growing AI document processing platform. You'll be a strong generalist who understands the mechanics of how ML models work – from serving and monitoring to building robust data pipeli...

Way of working
On site
Location
San Francisco, CA
Pay range
$200,000 to $300,000
Level
Mid
Experience
3+ years
Type
Full time
Visa sponsorship
Not offered for this role
The company
AI/ML · 50 to 200 people

Skills that matter here

PythonKubernetesPyTorchDockerHelm

The full description

We are looking for an ML Infrastructure Engineer with 3+ years of experience to own and scale the training and inference stack at a fast-growing AI document processing platform. You'll be a strong generalist who understands the mechanics of how ML models work – from serving and monitoring to building robust data pipelines – and can improve inference performance, reliability, and cost efficiency. This is a high-impact IC role where you'll work closely with ML researchers to ensure models are deployed quickly and reliably, and that infrastructure is never a bottleneck for the products being served. The ideal candidate is AI-native from the get-go, comfortable with 1-to-3 node training and single-to-double node serving, and thrives in a fast-paced startup environment.

What you will be doing

- Building and maintaining model serving infrastructure – improving inference speed, monitoring, and reliability to ensure it's never a bottleneck for customers

- Setting up and improving training infrastructure for models ranging from 300M to 30B parameters across 1-to-3 node environments

- Developing observability, logging, and monitoring systems across the ML stack

- Building internal data pipelines and tooling to help ML researchers move faster from experiment to production

- Architecting infrastructure to arbitrate inference between multiple cloud providers while optimizing for accuracy, latency, and cost

How hiring runs

  1. 1Phone Screen
  2. 2ML Infra Debugging & Coding
  3. 3Full-Day Onsite

Interested in this one?

There is no apply button here on purpose. Tell us about yourself, we book a short call, and if this role fits we walk you through the company and ask before anything is sent. Always free for you.

Tell us about you
Tell us about you