All open roles

Staff / Principal Platform Engineer

Engineering · Posted last month

We need a Staff or Principal-level Platform Engineer with 8+ years of experience in software engineering and deep expertise in Kubernetes and infrastructure-as-code. You should be comfortable taking end-to-end ownership of cloud infrastructure at scale and have a track record of building reliable, high-performance syst...

Way of working
Hybrid
Location
Mountain View, CA, San Francisco Bay Area, CA
Pay range
$280,000 to $350,000
Level
Principal
Experience
8+ years
Type
Full time
Visa sponsorship
Not offered for this role
The company
Software Development · 50 to 200 people

Skills that matter here

KubernetesTerraformTerragruntArgoCDHelmKustomizeGitHub ActionsAnsibleGoogle Cloud PlatformMicrosoft AzureOracle CloudGoPythonBashCI/CD

What you would be doing

  • Design, deploy, and maintain reliable, high-performance, and secure cloud infrastructure for Inworld's TTS and LLM Router products
  • Take ownership of CI/CD pipelines and infrastructure deployments using Terraform, ArgoCD, GitHub Actions, and related tooling
  • Work closely with engineers across the organization to deploy and evolve services across major cloud providers (GCP, Azure, Oracle Cloud)
  • Drive engineering velocity by identifying and building AI-powered tooling and workflows that improve how teams develop and deploy software
  • Facilitate a "you build it, you run it" culture by providing the necessary tools and processes for monitoring reliability, availability, and performance of services
  • Conduct root cause analysis to identify critical issues and develop automated solutions to prevent recurrence
  • Manage and scale Kubernetes clusters, including creating Kustomize manifests and Helm charts for application deployments

The full description

What we're looking for:

We need a Staff or Principal-level Platform Engineer with 8+ years of experience in software engineering and deep expertise in Kubernetes and infrastructure-as-code. You should be comfortable taking end-to-end ownership of cloud infrastructure at scale and have a track record of building reliable, high-performance systems that support consumer-facing AI products. Bonus points if you have experience with AI/ML infrastructure or realtime low-latency workloads.

What you'll do:

- Design, deploy, and maintain reliable, high-performance, and secure cloud infrastructure for Inworld's TTS and LLM Router products

- Take ownership of CI/CD pipelines and infrastructure deployments using Terraform, ArgoCD, GitHub Actions, and related tooling

- Work closely with engineers across the organization to deploy and evolve services across major cloud providers (GCP, Azure, Oracle Cloud)

- Drive engineering velocity by identifying and building AI-powered tooling and workflows that improve how teams develop and deploy software

- Facilitate a "you build it, you run it" culture by providing the necessary tools and processes for monitoring reliability, availability, and performance of services

- Conduct root cause analysis to identify critical issues and develop automated solutions to prevent recurrence

- Manage and scale Kubernetes clusters, including creating Kustomize manifests and Helm charts for application deployments

Interested in this one?

There is no apply button here on purpose. Tell us about yourself, we book a short call, and if this role fits we walk you through the company and ask before anything is sent. Always free for you.

Tell us about you
Tell us about you