All open roles

ML Infrastructure Engineer

Ai ml · Posted 28 days ago

We are looking for an ML Infrastructure Engineer with 7+ years of experience to own training infrastructure end-to-end and turn a multi-cloud GPU fleet into a world-class training engine for massive multimodal models. You'll be the connective tissue between researchers and compute – designing distributed training syste...

Way of working
On site
Location
Redwood City, CA, South Bay Area, CA, San Francisco, CA
Pay range
$220,000 to $350,000
Level
Senior
Experience
5+ years
Type
Full time
Visa sponsorship
Not offered for this role
The company
Technology,Information and Internet, Information and Internet · 50 to 200 people

Skills that matter here

PyTorchDeepSpeedAccelerateFSDPKubernetesSLURMGCPAWSTensorRTTritonNCCLDockerPythonCUDA

The full description

We are looking for an ML Infrastructure Engineer with 7+ years of experience to own training infrastructure end-to-end and turn a multi-cloud GPU fleet into a world-class training engine for massive multimodal models. You'll be the connective tissue between researchers and compute – designing distributed training systems, optimizing GPU utilization, and building the pipelines that ingest terabytes of multimodal robot data. This is a high-ownership role where your work directly accelerates the path from model to deployed robot. We need someone who has led technical projects in HPC or ML infrastructure and is genuinely passionate about the robotics space.

What you will be doing

- Architecting and scaling distributed training infrastructure across large GPU clusters – implementing sharding, activation checkpointing, and memory optimization (ZeRO, FSDP) for multimodal models

- Building researcher-friendly tooling and job scheduling systems (Kubernetes/SLURM) that prioritize fast iteration, automated retries, and seamless failure recovery

- Designing high-throughput data pipelines to ingest and transform terabytes of multimodal robot data (video, proprioception, 3D signals) so dataloaders never starve the GPUs

- Building low-latency inference pipelines for real-time robot control – applying quantization, distillation, and model compilation (TensorRT, Triton) to move models from lab to physical world

- Deep systems profiling – diving into GPU utilization, I/O bottlenecks, and memory fragmentation to squeeze maximum performance out of an expanding compute fleet

How hiring runs

  1. 1Recruiter Screen
  2. 2Coding Interview
  3. 3Technical Interview - Coding Round 1
  4. 4Final Interview

Interested in this one?

There is no apply button here on purpose. Tell us about yourself, we book a short call, and if this role fits we walk you through the company and ask before anything is sent. Always free for you.

Tell us about you
Tell us about you