Member of Technical Staff
Ai ml · Posted 2 months ago
Our mission at Wafer is to maximize intelligence per watt by using AI to optimize AI infrastructure, achieving orders of magnitude better energy and cost efficiency per token.
- Way of working
- On site
- Location
- San Francisco, CA
- Pay range
- $200,000 to $200,000
- Level
- Staff
- Experience
- 1 to 6 years
- Type
- Full time
- Visa sponsorship
- Not offered for this role
- The company
- Technology, Information and Internet · under 20 people
Skills that matter here
What you would be doing
- Ship day-zero support for new open-source models, tuned for latency and throughput
- Optimize the serving stack: batching, KV cache, speculative decoding, quantization
- Write and tune kernels in CUDA, HIP, and Triton for NVIDIA, AMD, TPU, Trainium, D-Matrix, and more.
- Design, deploy, and operate heterogeneous clusters across vendors
- Run production inference across a mixed fleet: reliability, observability, and cost per token at scale
The full description
Member of Technical Staff
Our mission at Wafer is to maximize intelligence per watt by using AI to optimize AI infrastructure, achieving orders of magnitude better energy and cost efficiency per token.
We believe cheap intelligence is the most essential piece of technology for a future of abundance. We care about building a future where intelligence is "too cheap to meter."
Wafer commercializes these efforts by serving serverless and dedicated inference for open source LLMs at the best performance per dollar. Our core bet is doing this through autonomous optimization of heterogeneous hardware.
What you'll do
- Ship day-zero support for new open-source models, tuned for latency and throughput
- Optimize the serving stack: batching, KV cache, speculative decoding, quantization
- Write and tune kernels in CUDA, HIP, and Triton for NVIDIA, AMD, TPU, Trainium, D-Matrix, and more.
- Design, deploy, and operate heterogeneous clusters across vendors
- Run production inference across a mixed fleet: reliability, observability, and cost per token at scale
How we evaluate
We score every candidate on seven values:
- Infinitely Resourceful
- Exceptionalism
- Unreasonable Standards
- Company Over Self
- High EQ
- Learns Quickly
- First Principles Thinker
How we work
On-site in San Francisco, five days a week. Small team with massive surface area and ownership. You operate with complete autonomy of how to solve problems, and work with the team to set the direction of your work. We don't see engineers as code writers, but as problem solvers. You will do everything from talking to customers to writing custom GPU kernels in esoteric hardware.
Interested in this one?
There is no apply button here on purpose. Tell us about yourself, we book a short call, and if this role fits we walk you through the company and ask before anything is sent. Always free for you.
Tell us about you