Kernel Engineer (ML Accelerators)
Ai ml · Posted 2 months ago
You'll diagnose and resolve performance problems across Decart's ML systems, spanning research, training, and production inference. The largest share of the work is writing and optimizing kernels for TPU and Trainium. You'll also advise researchers on the performance cost of proposed model changes. We're looking for en...
- Way of working
- On site
- Location
- San Francisco, CA
- Pay range
- $200,000 to $240,000
- Level
- Mid
- Experience
- 4+ years
- Type
- Full time
- Visa sponsorship
- Not offered for this role
- The company
- Technology,Information and Internet, Information and Internet · 50 to 200 people
Skills that matter here
What they are looking for
- Production experience squeezing performance out of ML workloads on TPU, Trainium, GPU, or other accelerators
- You've authored kernels for an ML accelerator — not just consumed them
- Depth in computer architecture: you can reason about systolic arrays, memory hierarchies, and interconnect topologies from first principles
- Familiarity with compiler and toolchain internals (e.g., XLA, MLIR, Triton, or vendor stacks)
- Experience scaling training or inference workloads across multi-accelerator clusters
- You've read — or patched — the internals of an ML framework
The full description
About the role
You'll diagnose and resolve performance problems across Decart's ML systems, spanning research, training, and production inference. The largest share of the work is writing and optimizing kernels for TPU and Trainium. You'll also advise researchers on the performance cost of proposed model changes. We're looking for engineers with a demonstrated record in large-scale systems engineering and low-level optimization.
Minimum requirements
- Bachelor's degree in Electrical/Computer Engineering, Computer Science, or a related field, plus 2+ years of relevant experience (or equivalent practical experience)
- 2+ years developing low-level software in C/C++ (Python proficiency a plus)
- Solid grounding in operating systems fundamentals (process/thread scheduling, synchronization, virtual memory), CPU/GPU architecture, and hardware/software co-design
- Working knowledge of PyTorch and machine learning algorithms, with a focus on engineering application
- Demonstrated ability to profile compute and memory behavior, diagnose bottlenecks, and validate improvements with rigorous measurement
What we're looking for
- Production experience squeezing performance out of ML workloads on TPU, Trainium, GPU, or other accelerators
- You've authored kernels for an ML accelerator — not just consumed them
- Depth in computer architecture: you can reason about systolic arrays, memory hierarchies, and interconnect topologies from first principles
- Familiarity with compiler and toolchain internals (e.g., XLA, MLIR, Triton, or vendor stacks)
- Experience scaling training or inference workloads across multi-accelerator clusters
- You've read — or patched — the internals of an ML framework
Projects you might work on
- Cut milliseconds off end-to-end token latency in DOS by restructuring attention and sampling paths for a new accelerator generation
- Design communication schedules that overlap compute with network transfer across multi-chip topologies
- Build analytical performance models to predict where the next 2x is hiding before writing a line of kernel code
- Trace a throughput regression from a framework-level symptom down to instruction scheduling in generated assembly — and fix it
- Port DOS's kernel suite to new hardware and close the gap to theoretical peak
Interested in this one?
There is no apply button here on purpose. Tell us about yourself, we book a short call, and if this role fits we walk you through the company and ask before anything is sent. Always free for you.
Tell us about you