Hands-on experience building and scaling web crawlers or scraping systems, ideally in support of machine learning training data.
Strong engineering skills in distributed systems at scale (e.g., Kubernetes, queue-based architectures, or custom pipelines processing billions of documents).
The capacity to autonomously evaluate the quality, coverage, and compliance of crawled data, and to build the tooling to measure it.
What you'll be doing
Building and operating large-scale, distributed web crawlers that discover, fetch, and extract data across billions of pages reliably and efficiently.
Solving hard crawling problems such as content extraction from messy HTML, deduplication at web scale, freshness and recrawl strategies, and politeness and rate-limit handling.
Designing targeted crawling pipelines that find high-value data sources, including audio, video, and multilingual content, and turn them into clean training-ready datasets.
Creating tooling and infrastructure that lets researchers request, monitor, and explore newly crawled web data quickly and reliably.
Perks and Benefits
Innovative culture: You’ll be part of a generational opportunity to define the trajectory of AI, surrounded by a team pushing the boundaries of what’s possible.
Growth paths: Joining ElevenLabs means joining a dynamic team with countless opportunities to drive impact - beyond your immediate role and responsibilities.
Learning & development: ElevenLabs proactively supports professional development through an annual discretionary stipend.
Social travel: We also provide an annual discretionary stipend to meet up with colleagues each year, however you choose.
Annual company offsite: Each year, we bring the entire team together in a new location - past offsites have included Croatia and Italy.
Co-working: If you’re not located near one of our main hubs, we offer a monthly co-working stipend.