All open roles

Member of Technical Staff - Model Optimization and Inference (Experienced)

Ai ml · Posted 7 months ago

We are seeking an experienced ML Infrastructure/Systems Engineer with at least 2 years of full-time experience in building and maintaining production-level ML systems. You should be comfortable designing scalable infrastructure from scratch, making informed design decisions by comparing various technologies, and have a...

Way of working
On site
Location
Seattle, WA
Pay range
$250,000 to $450,000
Level
Staff
Experience
2+ years
Type
Full time
Visa sponsorship
Not offered for this role
The company
Software Development · 20 to 50 people

Skills that matter here

KubernetesTerraformPythonRustGoDagsterRayAirflowWebRTCvLLMTriton Inference ServerTensorRT

What you would be doing

  • Own end-to-end inference optimization across our model stack — LLMs, audio models, and diffusion-based components
  • Implement and tune KV cache strategies for long-context conversations, including eviction policies, compression, and memory-efficient attention
  • Evaluate, deploy, and extend inference serving frameworks (vLLM, SGLang, TensorRT-LLM, etc.) for our specific workloads
  • Profile and benchmark end-to-end latency and throughput; identify and systematically eliminate bottlenecks
  • Build internal tooling that makes optimization work faster and more rigorous — profiling viewers, end-to-end inference test harnesses, and other infrastructure that helps the team move quickly
  • Accelerate diffusion model inference — consistency models, step distillation, caching strategies, and custom kernel optimizations
  • Apply and develop quantization techniques (INT8, INT4, GPTQ, AWQ, and beyond) to reduce memory footprint and increase throughput without meaningfully degrading quality
  • Work closely with research and infrastructure to ensure new models ship with optimized serving from day one

The full description

What we're looking for?

We are seeking an experienced ML Infrastructure/Systems Engineer with at least 2 years of full-time experience in building and maintaining production-level ML systems. You should be comfortable designing scalable infrastructure from scratch, making informed design decisions by comparing various technologies, and have a track record of optimizing systems for latency, throughput, and cost. We're looking for someone with a broad understanding of the ML infra space, including inference infrastructure, real-time video streaming, and data engineering, who can own complex projects and debug distributed systems. Bonus points if you have experience with video or audio models and low-level optimization techniques like CUDA kernels.

What you'll do:

- Own end-to-end inference optimization across our model stack — LLMs, audio models, and diffusion-based components

- Implement and tune KV cache strategies for long-context conversations, including eviction policies, compression, and memory-efficient attention

- Evaluate, deploy, and extend inference serving frameworks (vLLM, SGLang, TensorRT-LLM, etc.) for our specific workloads

- Profile and benchmark end-to-end latency and throughput; identify and systematically eliminate bottlenecks

- Build internal tooling that makes optimization work faster and more rigorous — profiling viewers, end-to-end inference test harnesses, and other infrastructure that helps the team move quickly

- Accelerate diffusion model inference — consistency models, step distillation, caching strategies, and custom kernel optimizations

- Apply and develop quantization techniques (INT8, INT4, GPTQ, AWQ, and beyond) to reduce memory footprint and increase throughput without meaningfully degrading quality

- Work closely with research and infrastructure to ensure new models ship with optimized serving from day one

Interested in this one?

There is no apply button here on purpose. Tell us about yourself, we book a short call, and if this role fits we walk you through the company and ask before anything is sent. Always free for you.

Tell us about you
Tell us about you