All open roles

Director of Site Reliability Engineering

Engineering · Posted 3 months ago

Director of Site Reliability Engineering at Stellar, leading a distributed team of four SREs and defining the vision, operating model, infrastructure foundations, and enablement practices that help engineering teams own and operate production services reliably. This hands-on leadership role reports to the CTO and spans cloud infrastructure, Kubernetes, CI/CD, observability, incident response, developer productivity, and operational maturity.

Way of working
Hybrid
Location
NYC/SF
Pay range
$210,000 to $310,000
Level
Principal
Experience
3+ years
Type
Unknown
Visa sponsorship
Not offered for this role

Skills that matter here

Site Reliability Engineering (SRE)DevOpsKubernetes

What you would be doing

  • No key responsibilities data available for this role
  • Please contact the recruiting team for detailed job requirements

What they are looking for

  • Site Reliability Engineering (SRE)
  • DevOps
  • Kubernetes

Worth knowing

  • Lean SRE leadership
  • Distributed operations experience
  • SRE function building
  • AI workflow adoption
  • Pragmatic tooling judgment
  • Strong executive communication
  • Founded by Jed McCaleb (creator of Mt. Gox, co-founder of Ripple) and Joyce Kim; Stellar launched Soroban smart contract platform on mainnet in February 2024 backed by $100M adoption fund, expanding from payments-only to DeFi/RWA ecosystem.
  • The Stellar network facilitated Franklin Templeton's first tokenized US mutual fund in 2021 and serves financial institutions globally including IBM banking partnerships in the South Pacific.
  • In March 2026, XLM (Stellar's native token) was classified as a US digital commodity by the SEC/CFTC, placing it alongside Bitcoin and Ethereum in regulatory clarity.

The full description

**Stellar** Director of Site Reliability Engineering; $210,000 – $310,000 + lumen-denominated grants; New York / San Francisco (Hybrid)

**About Company** Stellar is a decentralized, public blockchain that gives developers the tools to create experiences that are more like cash than crypto. The network is faster, cheaper, and far more energy-efficient than most blockchain-based systems. It's designed so Stellar's ecosystem can make a real-world, lasting impact.

Since 2014, the mission-driven team at the Stellar Development Foundation (SDF) has helped fuel the tremendous growth of the Stellar blockchain network, an open-source platform that operates at high-scale today. Developers and companies around the world build on it, and the SDF team is expanding to support the rapidly growing and changing Stellar ecosystem.

**About the Role** SDF is hiring a Director of Site Reliability Engineering to lead a team of 4 SREs and shape how engineering teams own, operate, and improve production services. This is a backfill — the function has been led at this level before — reporting directly to the CTO.

The Director will set the vision, operating model, and culture for SRE while owning the core infrastructure services that help SDF engineering teams build, deploy, observe, and operate software with confidence. Engineering teams at SDF own the services they build; SRE provides the frameworks, standards, shared infrastructure, tooling, observability practices, and enablement model that make strong service ownership possible across engineering.

This is a hands-on leadership role. SDF is a small, mission-driven foundation with a broad technical surface area, so this role requires leverage, ownership, and a bias toward solving the right problems over creating processes for its own sake. The ideal candidate brings strong technical judgment, pragmatic leadership, and the ability to influence through trust, clarity, and execution.

**Responsibilities** - Lead, coach, and develop a distributed SRE team of 4, setting a clear vision, charter, operating model, priorities, and success measures. - Define and roll out a Service Ownership & Maturity Framework across engineering, with expectations that vary appropriately by service criticality. - Own and improve core engineering infrastructure services, including cloud foundations, Kubernetes and compute patterns, CI/CD, observability, secrets management, GitHub workflows, and infrastructure automation. - Help engineering teams become stronger owners and operators of their services through better standards, dashboards, runbooks, alerting, escalation paths, operational readiness, and deployment practices. - Make reliability, operational maturity, infrastructure health, and developer productivity more measurable through trusted metrics and practical operational intelligence. - Improve deployment automation, resilience, self-healing patterns, disaster recovery readiness, and service reliability based on actual impact and risk. - Mature incident response, escalation, postmortems, and on-call health across a geographically distributed team. - Build paved paths and self-service infrastructure that reduce toil, lower cognitive load, and help engineering teams move faster while strengthening ownership and reliability. - Partner closely with Security, Compliance, Legal, Finance, Procurement, and Corporate IT where infrastructure, access management, cloud operations, vendor review, or controls intersect with engineering. - Pragmatically evaluate AI-assisted and agentic workflows where they can improve infrastructure operations, service ownership, developer workflows, or toil reduction.

**Required Skills** A. 10+ years of experience in SRE, infrastructure engineering, platform engineering, cloud infrastructure, production operations, or closely related engineering roles. B. 5+ years of experience leading, managing, or formally developing infrastructure, SRE, platform, or reliability engineers. C. 3+ years of experience with modern cloud infrastructure in AWS, GCP, or similar environments. D. 3+ years of experience with Kubernetes, container orchestration, infrastructure-as-code, declarative systems, CI/CD, and deployment safety.

**Bonus Skills** - Experience leading SRE, infrastructure, or platform work in a lean, high-agency organization. - Experience supporting globally distributed teams or 24/7 operational coverage. - Experience improving developer productivity through paved paths, self-service infrastructure, automation, and reduced toil. - Experience with infrastructure security fundamentals, secrets management, access controls, cloud security practices, or compliance-related infrastructure controls. - Experience in financial services, regulated environments, blockchain, crypto, Web3, or other high-reliability technical ecosystems. - Experience evaluating vendors and infrastructure platforms with skepticism, technical rigor, and cost discipline. - Practical experience applying AI-assisted or agentic workflows to infrastructure, reliability, operations, observability, or developer productivity.

**Logistical Info** - **Location: **Based in NY/SF or willing to relocate (in-person collaboration is critical). In-person component is important (if in NYC/SF, 3 days a week; if within 30 miles outside the city - 1 day a week) - **Compensation: **$210,000 – $310,000 depending on job-related knowledge, skills, experience, and location. Lumen-denominated grants are offered in addition to base salary. - Number of openings: 1 - **Other:** Can do visa transfer or provide TN visa, but not fresh visa or OPT

**Green Flags** - Hands-on technical leaders who can set vision AND roll up their sleeves — not just manage from a distance. - Experience at mid-sized companies where SRE/platform leaders wore multiple hats and had high ownership. - Track record of building or maturing SRE functions — defining charters, operating models, and frameworks. - Has actively adopted AI-assisted workflows into infrastructure or operations work. - Pragmatic approach to tooling — knows when to build, buy, simplify, or retire. - Strong executive communication — can partner directly with a CTO.

**Red Flags** - Candidates without any hands-on technical ability — this is not a purely managerial role. Stellar expects technical leaders to be somewhat hands-on. - Purely process-oriented leaders who default to creating bureaucracy over solving problems. - Candidates whose experience is only at very large organizations who may struggle with the breadth and ownership expected at a smaller, mission-driven foundation.

**Interview Process** - Recruiter Screen - Hiring Manager Video Interview - Panel 1 Interview — (a) Virtual: SRE-focused interview (non-coding), (b) Security-focused interview (non-coding), (c) Engineering-focused interview (non-coding) - Panel 2 Interview — In-Person: CTO, Lunch with the team

Interested in this one?

There is no apply button here on purpose. Tell us about yourself, we book a short call, and if this role fits we walk you through the company and ask before anything is sent. Always free for you.

Tell us about you
Tell us about you