All open roles

Member of Technical Staff, ML Systems

Engineering · Posted 6 days ago

TensorScale is rebuilding the training and inference stack for world models, from GPU kernels up through distributed serving. You will own speed and efficiency across that stack: low-level kernels, distributed inference engines, and multi-node training and serving systems. You report to the cofounder and CEO and work a...

Way of working
On site
Location
Menlo Park
Pay range
$180,000 to $230,000
Level
Staff
Experience
1+ years
Type
Full time
Visa sponsorship
Not offered for this role

Skills that matter here

CudaTritonPyTorchNsightNCCLRDMA

What you would be doing

  • Optimize GPU and system performance. Training and inference for image, video, and world-model workloads.
  • Profile and remove bottlenecks. Kernel, memory, system, and cluster level, using Nsight and related tooling.
  • Write low-level optimizations. CUDA and Triton, on code paths that run in production.
  • Build distributed engines. Inference and training for diffusion models across multiple GPUs and nodes.
  • Own communication performance. NCCL, RDMA over InfiniBand or RoCE, and disaggregated serving.
  • Make the gains stick. Benchmarking and regression harnesses so performance does not slide back in production.

The full description

TensorScale is rebuilding the training and inference stack for world models, from GPU kernels up through distributed serving. You will own speed and efficiency across that stack: low-level kernels, distributed inference engines, and multi-node training and serving systems. You report to the cofounder and CEO and work alongside four cofounders who between them cover distributed systems, kernel optimization, cloud infrastructure, and research.

You will feel at home here if you would rather make a video model ten times faster than train one.

What you'll do

- Optimize GPU and system performance. Training and inference for image, video, and world-model workloads.

- Profile and remove bottlenecks. Kernel, memory, system, and cluster level, using Nsight and related tooling.

- Write low-level optimizations. CUDA and Triton, on code paths that run in production.

- Build distributed engines. Inference and training for diffusion models across multiple GPUs and nodes.

- Own communication performance. NCCL, RDMA over InfiniBand or RoCE, and disaggregated serving.

- Make the gains stick. Benchmarking and regression harnesses so performance does not slide back in production.

Interested in this one?

There is no apply button here on purpose. Tell us about yourself, we book a short call, and if this role fits we walk you through the company and ask before anything is sent. Always free for you.

Tell us about you
Tell us about you