waypointjobs

Wintermeyer Ventures

Research Crawling Engineer

los angeles, CA

Check who can apply and the requirements below before continuing.

This job is closed

Applications are no longer available for this announcement. Explore current related opportunities below.

Job description

About the Role

This is a hands-on engineering role focused on building and operating large-scale web crawlers and data acquisition systems that power dataset creation for frontier AI model training. You will work as an extension of AI research teams, helping to collect, clean, and curate web data at a scale that few organizations can match. The work directly shapes the pretraining and inference pipelines used by some of the most advanced AI labs in the world.

What You'll Do

Build and maintain large-scale web crawlers across diverse domains including social media, travel, and multi-language sites.

Design high-throughput, fault-tolerant data collection systems capable of handling millions to billions of URLs per day.

Navigate anti-bot systems, rate limits, and JavaScript-heavy sites, finding creative solutions when standard protocols fall short.

Develop pipelines for cleaning, deduplication, filtering, and normalization of web data at TB to PB scale.

Construct and maintain datasets for research and model training in close collaboration with research teams.

Monitor crawl performance, coverage, and data quality, iterating quickly as web environments change.

Optimize infrastructure for cost, latency, and reliability across cloud and bare-metal environments.

What We're Looking For

3 or more years building and operating web crawlers at scale (2 or more years considered for candidates with a PhD).

Proficiency in one or more of: Go, Rust, Python, Java, or C++.

Experience running data pipelines at TB or greater scale.

Deep knowledge of HTTP, networking, and browser behavior.

Hands‑on experience with distributed systems or parallel processing.

Experience with headless browsers such as Playwright, Puppeteer, or Chrome DevTools Protocol.

Familiarity with proxy systems, IP rotation, or request orchestration.

Experience with data quality evaluation, scoring, or benchmarking at scale.

Experience running crawling or data workloads on cloud platforms (AWS, GCP) or bare-metal infrastructure.

Background in NLP pipelines, ML dataset curation, or AI lab work is a strong plus.

Compensation & Benefits

Salary range: $160,000 to $250,000 USD annually. Visa sponsorship is not available for this role.

Location

This role is fully remote. The primary location is Los Angeles, CA, United States, though candidates based in other major US cities are welcome.

#J-18808-Ljbffr

Who can apply

Review the original listing for work authorization, qualifications and employer requirements.

Explore related searches

Current related jobs

Amazon Web Services, Inc.

WhatJobs

Senior DevOps Engineer – Cloud Consulting

herndon, VA

Salary not specified

Amazon Web Services, Inc. seeks a Senior Delivery Consultant DevOps to lead complex cloud and DevOps engagements within AWS ProServe for public s…

Last received from source 2026-10-10View job

CNH Industrial

WhatJobs

Design Engineer II

oak brook, IL

Salary not specified

The Intermediate Electrical Design Engineer develops wire harnesses, cables, mounting brackets, and tractor electrical routing solutions from con…

Last received from source 2026-10-10View job

BNY

WhatJobs

Senior Specialist, Full-Stack Engineer

pittsburgh, PA

Salary not specified

hackajob is collaborating with BNY to connect them with exceptional professionals for this role. Full-Stack Engineer (Senior Specialist) - Risk E…

Last received from source 2026-10-10View job