Member of Technical Staff, ML Systems
Engineering · Posted 6 days ago
TensorScale is rebuilding the training and inference stack for world models, from GPU kernels up through distributed serving. You will own speed and efficiency across that stack: low-level kernels, distributed inference engines, and multi-node training and serving systems. You report to the cofounder and CEO and work a...
- Way of working
- On site
- Location
- Menlo Park
- Pay range
- $180,000 to $230,000
- Level
- Staff
- Experience
- 1+ years
- Type
- Full time
- Visa sponsorship
- Not offered for this role
Skills that matter here
What you would be doing
- Optimize GPU and system performance. Training and inference for image, video, and world-model workloads.
- Profile and remove bottlenecks. Kernel, memory, system, and cluster level, using Nsight and related tooling.
- Write low-level optimizations. CUDA and Triton, on code paths that run in production.
- Build distributed engines. Inference and training for diffusion models across multiple GPUs and nodes.
- Own communication performance. NCCL, RDMA over InfiniBand or RoCE, and disaggregated serving.
- Make the gains stick. Benchmarking and regression harnesses so performance does not slide back in production.
The full description
TensorScale is rebuilding the training and inference stack for world models, from GPU kernels up through distributed serving. You will own speed and efficiency across that stack: low-level kernels, distributed inference engines, and multi-node training and serving systems. You report to the cofounder and CEO and work alongside four cofounders who between them cover distributed systems, kernel optimization, cloud infrastructure, and research.
You will feel at home here if you would rather make a video model ten times faster than train one.
What you'll do
- Optimize GPU and system performance. Training and inference for image, video, and world-model workloads.
- Profile and remove bottlenecks. Kernel, memory, system, and cluster level, using Nsight and related tooling.
- Write low-level optimizations. CUDA and Triton, on code paths that run in production.
- Build distributed engines. Inference and training for diffusion models across multiple GPUs and nodes.
- Own communication performance. NCCL, RDMA over InfiniBand or RoCE, and disaggregated serving.
- Make the gains stick. Benchmarking and regression harnesses so performance does not slide back in production.
Interested in this one?
There is no apply button here on purpose. Tell us about yourself, we book a short call, and if this role fits we walk you through the company and ask before anything is sent. Always free for you.
Tell us about you