All open roles

Senior AI/ML Engineer

Ai ml · Posted last month

We're looking for a Senior AI/ML Engineer to build the AI systems layer behind Drawbridge's process intelligence platform.

Way of working
On site
Location
San Francisco, CA
Pay range
$180,000 to $275,000
Level
Senior
Experience
4 to 10 years
Type
Full time
Visa sponsorship
Not offered for this role
The company
under 20 people

Skills that matter here

PythonGoPostgreSQLRedisReactTypeScriptGCPTerraformDockerDatadogOpenAI APIAnthropic APIGeminiCohereRAGVector SearchLLM Evals

The full description

About this role

We're looking for a Senior AI/ML Engineer to build the AI systems layer behind Drawbridge's process intelligence platform.

This is a full-time role for a senior applied AI engineer who cares about making AI behavior measurable, reliable, safe, and useful in production. We build production AI systems on top of hosted frontier models from providers like OpenAI, Anthropic, etc.

You will own the systems that turn messy enterprise data into structured process intelligence: retrieval, context construction, model/provider selection, structured extraction, evals, orchestration, quality loops, and monitoring.

- Impact: You'll directly improve the quality, reliability, and usefulness of Drawbridge's AI harness

- Applied AI ownership: You'll build and own the production systems around LLM APIs.

- Evaluation discipline: You'll create the datasets, tests, and quality gates that help us know when the system is improving or regressing.

- Product partnership: You'll work closely with backend, product, and forward deployed engineering to solve real customer workflow problems.

- Technical leverage: Your work will shape how Drawbridge chooses models, constructs context, evaluates outputs, manages cost/latency, and safely ships AI behavior.

What we're looking for

Must-haves

- 4-6+ years building production software, ML systems, or applied AI systems

- Production AI engineering: You have strong skills and experience building reliable, observable AI services and pipelines, including handling latency, cost, retries, failures, and debugging in production

- Practical LLM experience: You have hands-on experience turning hosted LLM APIs into product behavior through structured outputs, retrieval and context construction, tool use, or multi-step workflows.

- AI quality ownership: You can define what good means, build representative datasets and evaluations, diagnose failures in data or model behavior, and ship regression tests, guardrails, and fallback paths.

- AI evaluation and guardrails: Experience with golden datasets, LLM-as-judge, human review, regression testing, launch gates, hallucination reduction, constraint enforcement, verification, and fallback paths.

- Systems thinking: You can reason about pipelines end-to-end, from raw customer data to generated outputs and user-facing product behavior

- Pragmatic mindset: You know when to refactor vs. ship, when to build vs. buy, and how to manage technical debt intentionally

- Ownership mentality: You take end-to-end responsibility for features, from design through deployment and monitoring

- Comfort with ambiguity: You can turn a fuzzy product goal into experiments, implementation, measurement, and shipped improvements

Strongly preferred

- Customer-facing communication: Clear written and verbal communicator with users/customers

- Retrieval and knowledge systems: Hands-on experience with embeddings, hybrid search, chunking, ranking, reranking, and assembling useful context from messy knowledge sources.

- Agentic workflows: Experience making multi-step, tool-using AI systems reliable when plans, tools, or intermediate outputs can fail.

- Document understanding: Experience extracting structured, trustworthy information from long or messy enterprise documents.

- AI-assisted development: You already use tools like Claude Code, Cursor, Codex, Copilot, or similar to accelerate your work

Nice to have

- Process mining, task mining, or enterprise workflow experience

- Multimodal inputs: Experience with documents, screenshots, video, transcripts, or mixed enterprise data sources

- Go or backend service experience

Our tech stack

Go, Python, PostgreSQL, Redis/Valkey, React 19 + TypeScript, GCP, Terraform, Docker, Datadog, GitHub

AI systems: hosted LLM APIs (OpenAI, Anthropic, Gemini), embeddings and retrieval, structured extraction, agent/tool-calling workflows, eval pipelines, and AI observability

Interested in this one?

There is no apply button here on purpose. Tell us about yourself, we book a short call, and if this role fits we walk you through the company and ask before anything is sent. Always free for you.

Tell us about you
Tell us about you