All open roles

Research Engineer

Engineering · Posted 2 months ago

You'll build the evaluation systems that tell us whether Firecrawl actually works. That sounds simple. It isn't. Our core promise, convert any URL into clean, structured, LLM-ready data reliably, is hard to measure rigorously across millions of different websites, formats, and edge cases. As the systems we're measuring...

Way of working
Hybrid
Location
San Francisco, CA, Bay Area, CA
Pay range
$250,000 to $290,000
Level
Mid
Experience
3+ years
Type
Full time
Visa sponsorship
Not offered for this role
The company
Technology,Information and Internet, Information and Internet · 20 to 50 people

Skills that matter here

PythonMachine LearningLLMsData PipelinesEvaluation FrameworksWeb ScrapingNLPStatistical Analysis

What you would be doing

  • Design the metrics that define what "good output" actually means across millions of sites, formats, and edge cases
  • Build the pipelines and harnesses that measure quality rigorously and at scale
  • Generate and curate the datasets that make evaluation trustworthy
  • Own the feedback loop from output quality back to model and product decisions
  • Turn "did that work?" into an answer the whole team can act on

What they are looking for

  • You have the engineering depth to build real evaluation systems, not just run existing ones
  • You care deeply about what "good" means and how to measure it rigorously
  • You're comfortable owning ambiguous problems where the metric itself has to be invented
  • You move fast and close the loop - you'd rather ship, measure, and iterate than perfect on paper

The full description

Research Engineer Role

You'll build the evaluation systems that tell us whether Firecrawl actually works. That sounds simple. It isn't. Our core promise, convert any URL into clean, structured, LLM-ready data reliably, is hard to measure rigorously across millions of different websites, formats, and edge cases. As the systems we're measuring get more complex, the question "did that work?" gets harder, not easier.

This isn't an eval role where you inherit a framework and run benchmarks. You'll design the metrics, build the pipelines, generate the datasets, and own the feedback loop from output quality back to model and product decisions. If you care about what "good" actually means and have the engineering depth to measure it, this is the role.

About Firecrawl

Firecrawl is the easiest way to turn the web into data AI agents can use. One API call converts any URL into clean, LLM-ready markdown or structured data - the boring-hard problem everyone building with LLMs eventually hits, solved.

We hit 8 figures in ARR in year one and more than doubled it in year two. We have 150k+ GitHub stars, and developers, agents, and category-defining AI companies build on us every day. Growth like this is rare, and we're just getting started.

We're a small team punching far above our weight. Everyone here owns a real piece of the product and company, end to end, and runs it themselves - no hiding behind process or headcount.

This is a place for people who want to work at the frontier: an AI company building the infrastructure other AI companies run on, not one bolting AI onto an existing product. We move fast, go deep, and are building the tools superintelligence will rely on to gather data from the web.

What You'll Do

- Design the metrics that define what "good output" actually means across millions of sites, formats, and edge cases

- Build the pipelines and harnesses that measure quality rigorously and at scale

- Generate and curate the datasets that make evaluation trustworthy

- Own the feedback loop from output quality back to model and product decisions

- Turn "did that work?" into an answer the whole team can act on

What We're Looking For

- You have the engineering depth to build real evaluation systems, not just run existing ones

- You care deeply about what "good" means and how to measure it rigorously

- You're comfortable owning ambiguous problems where the metric itself has to be invented

- You move fast and close the loop - you'd rather ship, measure, and iterate than perfect on paper

What We're NOT Looking For

- Someone who only wants to run benchmarks someone else designed

- A pure researcher who won't build the systems, or a pure engineer who won't think about methodology

- Someone who needs a fully-specced ticket to start

A Note On Pace

We operate at an absurd level of urgency because the window for what we're building won't stay open forever. If that excites you, keep reading. If it doesn't, no hard feelings — but this role probably isn't for you.

How hiring runs

  1. 1Intro Chat
  2. 2Technical Chat
  3. 3Founders Chat
  4. 4Paid Work Trial

Interested in this one?

There is no apply button here on purpose. Tell us about yourself, we book a short call, and if this role fits we walk you through the company and ask before anything is sent. Always free for you.

Tell us about you
Tell us about you