Wynd Labs logo
Posted 1mo ago•Remote

Data Engineer

MiddleRemoteSalary undisclosed
Required Skills
PythonKubernetesDockerKafkaCI/CDLinux
Job Description

Who We Are:

We build infrastructure that delivers massive amounts of web data to the companies training the world’s most powerful AI models.

We're the team that helps to power and support Grass, a bandwidth-sharing network that lets us operate a massive distributed crawler, giving us unique access to high-quality public web data at global scale. On top of that, we’ve built pipelines for ingesting, segmenting, and annotating billions of videos, transcripts, and audio files, powering dataset creation for frontier labs.

We’re lean, technical, and move fast. No red tape, no slow decision-making; just a team of builders pushing to expand what’s possible for open web data and AI.

The Role:

We are seeking a Data Engineer to support and improve large-scale data pipelines and infrastructure. You’ll work across data collection, processing, transformation, validation, and delivery, with a focus on scalability, reliability, and performance.
This is a hands-on role where you’ll work with distributed systems, large datasets, web scraping infrastructure, and production data workloads.

Please note: This role requires a work schedule that overlaps sufficiently with EST business hours to collaborate effectively with the team.

Similar Openings in Data & Analytics

View all in category➔