Infrastructure Engineer — Managed Inference
Engineering · Posted 2 months ago
We need someone with 5+ years of infrastructure engineering experience who has deep hands-on expertise operating production Kubernetes at scale and working with LLM inference serving systems. You should be comfortable debugging NVIDIA GPU systems end-to-end (drivers, CUDA, NCCL, network fabric) and have a track record...
- Way of working
- On site
- Location
- San Francisco, CA
- Pay range
- $200,000 to $260,000
- Level
- Senior
- Experience
- 5 to 10 years
- Type
- Full time
- Visa sponsorship
- Not offered for this role
- The company
- Technology,Information and Internet, Information and Internet · 50 to 200 people
Skills that matter here
What you would be doing
- Design, operate, and scale highly available Kubernetes clusters across AWS, GCP, and specialized GPU cloud providers to serve global inference demand at 1,000+ tokens per second
- Take ownership of the production serving layer — debug NVIDIA systems end-to-end including drivers, CUDA, NCCL, node health, and network fabric
- Work closely with the model optimization team to ensure kernel-level speed improvements survive contact with real traffic, real hardware, and real customers
- Extend deployment tooling so models ship consistently across NVIDIA, Trainium, and TPU hosts through a unified workflow
- Build observability infrastructure that keeps the platform honest: time-to-first-token, inter-token latency, throughput, and availability — measured per model, per chip, and per region
- Own reliability engineering including alerting, automated failover, self-healing infrastructure, and intelligent traffic routing across models, chips, and regions
The full description
What we're looking for:
We need someone with 5+ years of infrastructure engineering experience who has deep hands-on expertise operating production Kubernetes at scale and working with LLM inference serving systems. You should be comfortable debugging NVIDIA GPU systems end-to-end (drivers, CUDA, NCCL, network fabric) and have a track record of building highly available, multi-cloud infrastructure for latency-sensitive AI workloads. Bonus points if you've deployed across heterogeneous accelerator types (NVIDIA GPUs, AWS Trainium, Google TPUs) or contributed to open-source inference frameworks like vLLM or SGLang.
What you'll do:
- Design, operate, and scale highly available Kubernetes clusters across AWS, GCP, and specialized GPU cloud providers to serve global inference demand at 1,000+ tokens per second
- Take ownership of the production serving layer — debug NVIDIA systems end-to-end including drivers, CUDA, NCCL, node health, and network fabric
- Work closely with the model optimization team to ensure kernel-level speed improvements survive contact with real traffic, real hardware, and real customers
- Extend deployment tooling so models ship consistently across NVIDIA, Trainium, and TPU hosts through a unified workflow
- Build observability infrastructure that keeps the platform honest: time-to-first-token, inter-token latency, throughput, and availability — measured per model, per chip, and per region
- Own reliability engineering including alerting, automated failover, self-healing infrastructure, and intelligent traffic routing across models, chips, and regions
Interested in this one?
There is no apply button here on purpose. Tell us about yourself, we book a short call, and if this role fits we walk you through the company and ask before anything is sent. Always free for you.
Tell us about you