Senior GPU Infrastructure Engineer
Engineering · Posted 3 months ago
Hyperbolic is hiring a Senior GPU Infrastructure Engineer to build the multi-tenancy provisioning, virtualization, scheduling, storage, and networking layers behind its open-access GPU cloud marketplace. This foundational role covers the stack from bare-metal hardware through orchestration, serving AI developers and researchers at a Series A company of approximately 30 people.
- Way of working
- Hybrid
- Location
- San Francisco
- Pay range
- $180,000 to $250,000
- Level
- Senior
- Experience
- 5 to 8 years
- Type
- Unknown
- Visa sponsorship
- Available for this role
- The company
- SERIES_A
Skills that matter here
What you would be doing
- No key responsibilities data available for this role
- Please contact the recruiting team for detailed job requirements
What they are looking for
- Kubernetes
- Linux
- Distributed Systems
- DevOps
- CI/CD
- Terraform
- Python
- Docker
Worth knowing
- GPU infrastructure background
- Bare-metal GPU provisioning
- GPU scheduling systems
- Hardware vendor experience
- Startup environment comfort
- Hardware-software breadth
- Founded by elite AI researchers: CEO Dr. Jasper Zhang (Math Ph.D. from UC Berkeley, gold medalist in Alibaba Global Math Competition) and CTO Dr. Yuchen Jin (renowned AI researcher), solving the GPU accessibility bottleneck
- Already deployed with major partners including Hugging Face, Quora, Cornell University, and UC Berkeley, with 40,000+ developers in ecosystem
- Building proprietary Proof of Sampling (PoSP) protocol for verifiable, decentralized AI—claims to be only company delivering scalable, verifiable AI at Web2 performance levels
The full description
**Hyperbolic** Senior GPU Infrastructure Engineer; $180K–$250K + equity; San Francisco (Hybrid)
**About Company**
Hyperbolic Labs is pioneering AI infrastructure with its open-access GPU cloud, aggregating computing resources across the globe to offer an innovative GPU marketplace and AI inference service at up to 75% cost savings compared to traditional cloud providers. The mission is to democratize AI by breaking down barriers to computing power — making it affordable and accessible to developers and researchers everywhere.
Founded by co-founders with PhDs in AI, Math, and Computer Science, Hyperbolic raised a Series A and is preparing for significant growth. The team is ~30 people. The company sits at the intersection of AI and open-source technology, and the culture is engineering-driven with a strong emphasis on ownership, execution, and building at the cutting edge of GPU infrastructure.
**About the Role**
This is a foundational infrastructure role: building the multi-tenancy provisioning and virtualization layer that transforms raw GPUs from diverse global suppliers into a programmable, orchestrated pool serving thousands of AI developers and researchers. You'll work at the cutting edge of cloud infrastructure, building the core orchestration layer that enables Hyperbolic to deliver up to 75% cost savings vs. traditional cloud. This is hands-on, deep infrastructure work — bare-metal provisioning, GPU scheduling, storage systems, and networking — not a wrapper around AWS. This role will report to the VP of Engineering once that hire is in place; in the interim, expect to work closely with the co-founders and existing engineering team.
**Responsibilities**
- Build and scale the GPU Cloud Marketplace by designing multi-tenancy provisioning and virtualization solutions - Manage bare-metal provisioning and lifecycle management: IPMI/Redfish, BMC-based remote management, PXE boot, automated OS deployment - Design and implement GPU scheduling and orchestration: GPU type awareness, memory management, topology considerations, placement strategies for multi-GPU jobs, fragmentation minimization - Build and maintain infrastructure automation using Terraform or Pulumi, CI/CD pipelines, secrets management, configuration management, and observability stacks - Design storage and data infrastructure for AI/ML workloads: object storage, high-IOPS block storage, distributed file systems for training data and checkpoints - Implement API design and cloud-init for automated provisioning and configuration - Work with hardware vendors and vendor engineering teams to troubleshoot issues and optimize integrations
**Required Skills**
A. Deep understanding of bare-metal provisioning and lifecycle management, including IPMI/Redfish, BMC-based remote management, PXE boot, and automated OS deployment workflows B. Deep understanding of GPU scheduling and orchestration: GPU type awareness, memory management, topology considerations, placement strategies for multi-GPU jobs, fragmentation minimization C. Strong infrastructure and DevOps engineering skills: Terraform or Pulumi, CI/CD for infrastructure, secrets management, configuration management, and observability stack implementation D. Experience with storage and data infrastructure for AI/ML workloads: object storage, high-IOPS block storage, and distributed file systems for training data and checkpoints E. Solid understanding of GPU architecture, CUDA, and GPU compute optimization F. Proven experience building and scaling cloud infrastructure or distributed systems in production environments G. Excellent communication skills across technical and non-technical stakeholders; proven ability to work with hardware vendors
**Bonus Skills**
- Familiarity with high-performance networking: InfiniBand and RoCE (RDMA over Converged Ethernet) - Experience with distributed storage systems: Ceph, Weka, or VAST Data - Experience with Kubernetes GPU operators, Slurm, or Ray for distributed training - Background at GPU cloud or AI infrastructure companies
**Logistical Info**
- Location: San Francisco, Hybrid — ideally 2–3 days per week in office. US-based required. - Compensation: $180K–$250K base + equity. Flexibility for very senior candidates. - Number of openings: 1 - **Other:** ~30-person team, Series A stage. Visa: US citizen / green card preferred; open to H-1B transfers.
**Ideal Background**
- Strongest signal — GPU cloud infrastructure companies: CoreWeave, Lambda, SF Compute, Modal, Together AI, Shadeform, RunPod, Crusoe, or a hyperscaler's GPU/AI infra team - Understands the full stack from bare-metal hardware up through orchestration and scheduling - Sweet spot is 5–8+ years with real GPU cluster operations in production (not just training models on them) - Has dealt with hardware vendor relationships and built provisioning systems from scratch - Cloud-native experience alone (AWS/GCP without bare-metal) is insufficient — this role requires understanding the physical layer
**Green Flags**
- Background at GPU cloud / AI infrastructure companies (CoreWeave, Lambda, SF Compute, Modal, Together AI, Shadeform, RunPod, Crusoe) - Hands-on experience with bare-metal GPU provisioning — not just cloud abstraction layers - Has built or meaningfully contributed to a GPU scheduling/orchestration system - Experience working directly with hardware vendors (NVIDIA, Supermicro, etc.) - Startup environment comfort — high agency, wears multiple hats, ships fast - Understands both the hardware and software sides of GPU infrastructure
**Red Flags**
- Wrong company background — enterprise IT, pure software companies without infrastructure depth, SaaS companies (this is the #1 historical rejection reason for Hyperbolic roles: 29/33 prior submissions rejected for this) - Pure cloud-native without bare-metal experience — must understand hardware-level provisioning - ML researcher or data scientist who "also does infra" — this is a deep infrastructure role, not an ML engineering role - No experience with GPU-specific challenges (scheduling, topology, CUDA optimization) - Overconfident / high ego — small team culture requires collaborative, low-ego engineers
**Interview Process**
1. Recruiter screen (Austin Dupuy) 2. Technical interview (systems design / infrastructure deep-dive) 3. Onsite / final loop
Interested in this one?
There is no apply button here on purpose. Tell us about yourself, we book a short call, and if this role fits we walk you through the company and ask before anything is sent. Always free for you.
Tell us about you