All open roles

Principal Infrastructure Engineer / Tech Lead

Engineering · Posted 14 days ago

We need an experienced infrastructure engineer with 7+ years of experience who has led infrastructure projects and is comfortable building and scaling cloud systems (AWS preferred) in a fast-paced startup environment. You should have strong CS fundamentals, Python proficiency, and a track record of owning critical infr...

Way of working
On site
Location
San Francisco, CA
Pay range
$200,000 to $250,000
Level
Principal
Experience
7+ years
Type
Full time
Visa sponsorship
Not offered for this role
The company
Software Development · under 20 people

Skills that matter here

AWSPythonCI/CDInfrastructure-as-CodeGPU Workloads

What you would be doing

  • Own and scale the cloud infrastructure (AWS) that powers custom AI model training and inference for enterprise customers
  • Design and maintain reliable deployment processes, CI/CD pipelines, and infrastructure-as-code for getting models and services into production
  • Manage GPU allocation, instance sizing, and cost optimization across cloud resources
  • Build monitoring, logging, and automated testing infrastructure that gives the team visibility into what's running, failing, and why
  • Own data and ML ops including ETL pipelines from varied customer data sources, model versioning, and lifecycle management
  • Drive security engineering including data isolation between customers, infrastructure hardening, and SOC2 compliance
  • Mentor and partner with the existing infrastructure engineer to raise the bar for the team

The full description

What we're looking for:

We need an experienced infrastructure engineer with 7+ years of experience who has led infrastructure projects and is comfortable building and scaling cloud systems (AWS preferred) in a fast-paced startup environment. You should have strong CS fundamentals, Python proficiency, and a track record of owning critical infrastructure end-to-end. Bonus points if you have familiarity with ML workflows or MLOps.

What you'll do:

- Own and scale the cloud infrastructure (AWS) that powers custom AI model training and inference for enterprise customers

- Design and maintain reliable deployment processes, CI/CD pipelines, and infrastructure-as-code for getting models and services into production

- Manage GPU allocation, instance sizing, and cost optimization across cloud resources

- Build monitoring, logging, and automated testing infrastructure that gives the team visibility into what's running, failing, and why

- Own data and ML ops including ETL pipelines from varied customer data sources, model versioning, and lifecycle management

- Drive security engineering including data isolation between customers, infrastructure hardening, and SOC2 compliance

- Mentor and partner with the existing infrastructure engineer to raise the bar for the team

How hiring runs

  1. 1Intro Call
  2. 2Coding Screen
  3. 3Take-Home Project
  4. 4Take-Home Presentation
  5. 5Final Onsite

Interested in this one?

There is no apply button here on purpose. Tell us about yourself, we book a short call, and if this role fits we walk you through the company and ask before anything is sent. Always free for you.

Tell us about you
Tell us about you