Senior AI/ML Engineer
Ai ml · Posted last month
We're looking for a Senior AI/ML Engineer to build the AI systems layer behind Drawbridge's process intelligence platform.
- Way of working
- On site
- Location
- San Francisco, CA
- Pay range
- $180,000 to $275,000
- Level
- Senior
- Experience
- 4 to 10 years
- Type
- Full time
- Visa sponsorship
- Not offered for this role
- The company
- under 20 people
Skills that matter here
The full description
About this role
We're looking for a Senior AI/ML Engineer to build the AI systems layer behind Drawbridge's process intelligence platform.
This is a full-time role for a senior applied AI engineer who cares about making AI behavior measurable, reliable, safe, and useful in production. We build production AI systems on top of hosted frontier models from providers like OpenAI, Anthropic, etc.
You will own the systems that turn messy enterprise data into structured process intelligence: retrieval, context construction, model/provider selection, structured extraction, evals, orchestration, quality loops, and monitoring.
- Impact: You'll directly improve the quality, reliability, and usefulness of Drawbridge's AI harness
- Applied AI ownership: You'll build and own the production systems around LLM APIs.
- Evaluation discipline: You'll create the datasets, tests, and quality gates that help us know when the system is improving or regressing.
- Product partnership: You'll work closely with backend, product, and forward deployed engineering to solve real customer workflow problems.
- Technical leverage: Your work will shape how Drawbridge chooses models, constructs context, evaluates outputs, manages cost/latency, and safely ships AI behavior.
What we're looking for
Must-haves
- 4-6+ years building production software, ML systems, or applied AI systems
- Production AI engineering: You have strong skills and experience building reliable, observable AI services and pipelines, including handling latency, cost, retries, failures, and debugging in production
- Practical LLM experience: You have hands-on experience turning hosted LLM APIs into product behavior through structured outputs, retrieval and context construction, tool use, or multi-step workflows.
- AI quality ownership: You can define what good means, build representative datasets and evaluations, diagnose failures in data or model behavior, and ship regression tests, guardrails, and fallback paths.
- AI evaluation and guardrails: Experience with golden datasets, LLM-as-judge, human review, regression testing, launch gates, hallucination reduction, constraint enforcement, verification, and fallback paths.
- Systems thinking: You can reason about pipelines end-to-end, from raw customer data to generated outputs and user-facing product behavior
- Pragmatic mindset: You know when to refactor vs. ship, when to build vs. buy, and how to manage technical debt intentionally
- Ownership mentality: You take end-to-end responsibility for features, from design through deployment and monitoring
- Comfort with ambiguity: You can turn a fuzzy product goal into experiments, implementation, measurement, and shipped improvements
Strongly preferred
- Customer-facing communication: Clear written and verbal communicator with users/customers
- Retrieval and knowledge systems: Hands-on experience with embeddings, hybrid search, chunking, ranking, reranking, and assembling useful context from messy knowledge sources.
- Agentic workflows: Experience making multi-step, tool-using AI systems reliable when plans, tools, or intermediate outputs can fail.
- Document understanding: Experience extracting structured, trustworthy information from long or messy enterprise documents.
- AI-assisted development: You already use tools like Claude Code, Cursor, Codex, Copilot, or similar to accelerate your work
Nice to have
- Process mining, task mining, or enterprise workflow experience
- Multimodal inputs: Experience with documents, screenshots, video, transcripts, or mixed enterprise data sources
- Go or backend service experience
Our tech stack
Go, Python, PostgreSQL, Redis/Valkey, React 19 + TypeScript, GCP, Terraform, Docker, Datadog, GitHub
AI systems: hosted LLM APIs (OpenAI, Anthropic, Gemini), embeddings and retrieval, structured extraction, agent/tool-calling workflows, eval pipelines, and AI observability
Interested in this one?
There is no apply button here on purpose. Tell us about yourself, we book a short call, and if this role fits we walk you through the company and ask before anything is sent. Always free for you.
Tell us about you