All open roles

Kernel Engineer (ML Accelerators)

Ai ml · Posted 2 months ago

You'll diagnose and resolve performance problems across Decart's ML systems, spanning research, training, and production inference. The largest share of the work is writing and optimizing kernels for TPU and Trainium. You'll also advise researchers on the performance cost of proposed model changes. We're looking for en...

Way of working
On site
Location
San Francisco, CA
Pay range
$200,000 to $240,000
Level
Mid
Experience
4+ years
Type
Full time
Visa sponsorship
Not offered for this role
The company
Technology,Information and Internet, Information and Internet · 50 to 200 people

Skills that matter here

C/C++PythonPyTorchCUDAXLAMLIRTritonTPUTrainiumGPU

What they are looking for

  • Production experience squeezing performance out of ML workloads on TPU, Trainium, GPU, or other accelerators
  • You've authored kernels for an ML accelerator — not just consumed them
  • Depth in computer architecture: you can reason about systolic arrays, memory hierarchies, and interconnect topologies from first principles
  • Familiarity with compiler and toolchain internals (e.g., XLA, MLIR, Triton, or vendor stacks)
  • Experience scaling training or inference workloads across multi-accelerator clusters
  • You've read — or patched — the internals of an ML framework

The full description

About the role

You'll diagnose and resolve performance problems across Decart's ML systems, spanning research, training, and production inference. The largest share of the work is writing and optimizing kernels for TPU and Trainium. You'll also advise researchers on the performance cost of proposed model changes. We're looking for engineers with a demonstrated record in large-scale systems engineering and low-level optimization.

Minimum requirements

- Bachelor's degree in Electrical/Computer Engineering, Computer Science, or a related field, plus 2+ years of relevant experience (or equivalent practical experience)

- 2+ years developing low-level software in C/C++ (Python proficiency a plus)

- Solid grounding in operating systems fundamentals (process/thread scheduling, synchronization, virtual memory), CPU/GPU architecture, and hardware/software co-design

- Working knowledge of PyTorch and machine learning algorithms, with a focus on engineering application

- Demonstrated ability to profile compute and memory behavior, diagnose bottlenecks, and validate improvements with rigorous measurement

What we're looking for

- Production experience squeezing performance out of ML workloads on TPU, Trainium, GPU, or other accelerators

- You've authored kernels for an ML accelerator — not just consumed them

- Depth in computer architecture: you can reason about systolic arrays, memory hierarchies, and interconnect topologies from first principles

- Familiarity with compiler and toolchain internals (e.g., XLA, MLIR, Triton, or vendor stacks)

- Experience scaling training or inference workloads across multi-accelerator clusters

- You've read — or patched — the internals of an ML framework

Projects you might work on

- Cut milliseconds off end-to-end token latency in DOS by restructuring attention and sampling paths for a new accelerator generation

- Design communication schedules that overlap compute with network transfer across multi-chip topologies

- Build analytical performance models to predict where the next 2x is hiding before writing a line of kernel code

- Trace a throughput regression from a framework-level symptom down to instruction scheduling in generated assembly — and fix it

- Port DOS's kernel suite to new hardware and close the gap to theoretical peak

Interested in this one?

There is no apply button here on purpose. Tell us about yourself, we book a short call, and if this role fits we walk you through the company and ask before anything is sent. Always free for you.

Tell us about you
Tell us about you