Staff / Principal Platform Engineer
Engineering · Posted last month
We need a Staff or Principal-level Platform Engineer with 8+ years of experience in software engineering and deep expertise in Kubernetes and infrastructure-as-code. You should be comfortable taking end-to-end ownership of cloud infrastructure at scale and have a track record of building reliable, high-performance syst...
- Way of working
- Hybrid
- Location
- Mountain View, CA, San Francisco Bay Area, CA
- Pay range
- $280,000 to $350,000
- Level
- Principal
- Experience
- 8+ years
- Type
- Full time
- Visa sponsorship
- Not offered for this role
- The company
- Software Development · 50 to 200 people
Skills that matter here
What you would be doing
- Design, deploy, and maintain reliable, high-performance, and secure cloud infrastructure for Inworld's TTS and LLM Router products
- Take ownership of CI/CD pipelines and infrastructure deployments using Terraform, ArgoCD, GitHub Actions, and related tooling
- Work closely with engineers across the organization to deploy and evolve services across major cloud providers (GCP, Azure, Oracle Cloud)
- Drive engineering velocity by identifying and building AI-powered tooling and workflows that improve how teams develop and deploy software
- Facilitate a "you build it, you run it" culture by providing the necessary tools and processes for monitoring reliability, availability, and performance of services
- Conduct root cause analysis to identify critical issues and develop automated solutions to prevent recurrence
- Manage and scale Kubernetes clusters, including creating Kustomize manifests and Helm charts for application deployments
The full description
What we're looking for:
We need a Staff or Principal-level Platform Engineer with 8+ years of experience in software engineering and deep expertise in Kubernetes and infrastructure-as-code. You should be comfortable taking end-to-end ownership of cloud infrastructure at scale and have a track record of building reliable, high-performance systems that support consumer-facing AI products. Bonus points if you have experience with AI/ML infrastructure or realtime low-latency workloads.
What you'll do:
- Design, deploy, and maintain reliable, high-performance, and secure cloud infrastructure for Inworld's TTS and LLM Router products
- Take ownership of CI/CD pipelines and infrastructure deployments using Terraform, ArgoCD, GitHub Actions, and related tooling
- Work closely with engineers across the organization to deploy and evolve services across major cloud providers (GCP, Azure, Oracle Cloud)
- Drive engineering velocity by identifying and building AI-powered tooling and workflows that improve how teams develop and deploy software
- Facilitate a "you build it, you run it" culture by providing the necessary tools and processes for monitoring reliability, availability, and performance of services
- Conduct root cause analysis to identify critical issues and develop automated solutions to prevent recurrence
- Manage and scale Kubernetes clusters, including creating Kustomize manifests and Helm charts for application deployments
Interested in this one?
There is no apply button here on purpose. Tell us about yourself, we book a short call, and if this role fits we walk you through the company and ask before anything is sent. Always free for you.
Tell us about you