All open roles

Research Crawling Engineer

Data · Posted 3 months ago

We need someone with 4+ years of experience (2+ for candidates with PhDs) building web crawlers and large-scale data acquisition systems who has hands-on experience designing high-throughput, fault-tolerant pipelines. You should be comfortable operating at the boundary of scale and reliability in adversarial web enviro...

Way of working
Remote
Location
Remote
Pay range
$160,000 to $250,000
Level
Mid
Experience
2+ years
Type
Full time
Visa sponsorship
Not offered for this role
The company
IT Services and IT Consulting,Software Development · 20 to 50 people

Skills that matter here

GoRustPythonJavaC++Distributed SystemsHeadless BrowsersPlaywrightPuppeteerChrome DevTools ProtocolProxy SystemsIP RotationHTTP/NetworkingNLP PipelinesData PipelinesCloud InfrastructureBare-Metal Infrastructure

What you would be doing

  • Build and maintain large-scale web crawlers across diverse domains (social media, travel, multi-language sites) that power dataset creation for frontier AI labs
  • Design high-throughput, fault-tolerant systems for data collection handling millions to billions of URLs/day
  • Handle anti-bot systems, rate limits, and dynamic/JS-heavy sites — thinking creatively when standard protocols fail
  • Develop pipelines for cleaning, deduplication, filtering, and normalization of web data at TB–PB scale
  • Construct and maintain datasets for research and model training, collaborating directly with research teams to align data collection with modeling needs
  • Monitor crawl performance, coverage, and data quality; iterate quickly as web environments constantly change
  • Optimize infrastructure for cost, latency, and reliability across cloud or bare-metal environments

The full description

What we're looking for:

We need someone with 4+ years of experience (2+ for candidates with PhDs) building web crawlers and large-scale data acquisition systems who has hands-on experience designing high-throughput, fault-tolerant pipelines. You should be comfortable operating at the boundary of scale and reliability in adversarial web environments and have a track record of processing millions to billions of URLs/day. Bonus points if you have experience with NLP pipelines, LLM pretraining data, or dataset curation for ML.

What you'll do:

- Build and maintain large-scale web crawlers across diverse domains (social media, travel, multi-language sites) that power dataset creation for frontier AI labs

- Design high-throughput, fault-tolerant systems for data collection handling millions to billions of URLs/day

- Handle anti-bot systems, rate limits, and dynamic/JS-heavy sites — thinking creatively when standard protocols fail

- Develop pipelines for cleaning, deduplication, filtering, and normalization of web data at TB–PB scale

- Construct and maintain datasets for research and model training, collaborating directly with research teams to align data collection with modeling needs

- Monitor crawl performance, coverage, and data quality; iterate quickly as web environments constantly change

- Optimize infrastructure for cost, latency, and reliability across cloud or bare-metal environments

About the team:

We build infrastructure that delivers massive amounts of web data to the companies training the world's most powerful AI models. Frontier AI labs are our customers — you'll be working as an extension of their data teams on cutting-edge pre-training and inference model development. We own one of the largest repositories of public web data and have more resources available for working with data at scale than basically any other company. We're a lean, flat organization (~42 people) with no people managers — just builders pushing to expand what's possible for open web data and AI. We're cash flow positive and growing quickly.

How hiring runs

  1. 1Initial Call with CEO or Technical Interview
  2. 2Technical Interview or CEO Call
  3. 3CTO or Additional Technical Interview
  4. 4COO Cultural Interview (Optional)
  5. 5Reference Check

Interested in this one?

There is no apply button here on purpose. Tell us about yourself, we book a short call, and if this role fits we walk you through the company and ask before anything is sent. Always free for you.

Tell us about you
Tell us about you