--- title: 'Data Engineer at Wynd Labs' canonical: 'https://feeny.ai/job/data-engineer-wynd-labs-remote-za2c14y2h3e5' type: 'job' last_seen: '2026-09-06' --- # Data Engineer at Wynd Labs - **Company:** Wynd Labs - **Location:** Remote - **Employment:** full-time - **Work type:** remote - **Posted:** 2026-09-04 - **Last confirmed live:** 2026-09-06 - **Apply:** https://jobs.ashbyhq.com/wynd-labs/216d74ef-d925-4bf1-ae4f-9eedc3e145f8 ## Job description Who We Are: We build infrastructure that delivers massive amounts of web data to the companies training the world’s most powerful AI models. We're the team that helps to power and support Grass, a bandwidth-sharing network that lets us operate a massive distributed crawler, giving us unique access to high-quality public web data at global scale. On top of that, we’ve built pipelines for ingesting, segmenting, and annotating billions of videos, transcripts, and audio files, powering dataset creation for frontier labs. We’re lean, technical, and move fast. No red tape, no slow decision-making; just a team of builders pushing to expand what’s possible for open web data and AI. The Role: We are seeking a Data Engineer to support and improve large-scale data pipelines and infrastructure. You’ll work across data collection, processing, transformation, validation, and delivery, with a focus on scalability, reliability, and performance. This is a hands-on role where you’ll work with distributed systems, large datasets, web scraping infrastructure, and production data workloads. Please note: This role requires a work schedule that overlaps sufficiently with EST business hours to collaborate effectively with the team. Who You Are: - Bachelor’s degree or equivalent work experience - Python (advanced) — strong grasp of async programming, multiprocessing, and writing production-grade code for long-running data jobs - Web scraping at scale — hands-on experience with high-volume scraping (proxies, rate limiting, anti-bot evasion). Experience with platform APIs and large media/metadata datasets (video platforms, social media) - Distributed data pipelines — experience designing and operating pipelines across many workers/servers using task queues (Celery, Kafka, RabbitMQ, or similar) - Data warehousing — practical experience with columnar/analytical warehouses; Databend, ClickHouse, or BigQuery strongly preferred; comfortable with complex analytical queries, partitioning strategies, cost-aware querying on cloud warehouses - Docker & Kubernetes — containerizing workloads, writing Helm charts/manifests, managing deployments, autoscaling scraping/processing workloads - Linux & bare-metal ops — comfortable managing services on Linux servers, debugging performance issues (disk I/O, network, memory) without managed-cloud abstractions - CI/CD for data workflows (GitHub Actions, ArgoCD) - Writing Scalable API What You'll Be Doing: - Maintain, optimize, and troubleshoot database queries and related data systems to support efficient data access, processing, and reliability. - Assist in creating, maintaining, and improving data pipelines used to collect, process, transform, validate, and deliver large-scale datasets. - Support web scraping and data collection initiatives, including developing, testing, and maintaining scripts or tools used to gather publicly available data in accordance with Company requirements. - Monitor and troubleshoot data pipeline issues, identify data quality concerns, and help implement timely fixes to maintain data accuracy and operational continuity. - Document engineering work, including database queries, pipeline processes, scraping workflows, technical decisions, issues encountered, and resolutions implemented. - Participate in research and development projects to improve the Company’s data products and workflows. Why Work With Us: - Opportunity. We are at the forefront of developing a web-scale crawler and knowledge graph that improves access to public web data and extends the value of AI to the people. - Culture. We're a lean team with a high bar. We come to work not to be comfortable, but to find out what we're capable of and to do work that matters. We're not calling for people who keep things moving. We're calling for people who make everyone around them better. We prioritize low ego and high output. This is a fully remote team. - Compensation. You’ll receive a competitive salary, benefits and equity package. ## About Wynd Labs ## Company Overview - **One-liner**: Wynd Labs builds a decentralized proxy network to provide internet-scale web data for AI training, research, analytics, and business intelligence. - **Entity Type**: Private (Seed stage) - **Headquarters**: Spokane, Washington, United States (with offices in Toronto, Canada, and a distributed team in the Philippines, Singapore, and Denmark) - **Founded**: 2022 - **Founders**: Andrej R. (Co-Founder & CEO) ## Core Business - **Primary industry/industries**: Software Development, AI Infrastructure, Decentralized Data Networks - **Target customers**: B2B — AI development companies, research institutions, analytics firms, and business intelligence teams that require large-scale access to public web data - **Mission or purpose statement**: To make data acquisition faster, more transparent, and accessible, fostering a fair and equitable internet through a decentralized approach to data access. ## Products & Services - **Grass (Core Product)**: A decentralized residential proxy network that enables large-scale, high-speed access to public web data. Designed to improve how public web data is accessed for AI applications, including AI training data, price scraping, and financial data analytics. Type: Decentralized Network / Infrastructure. ## Market Standing - **Valuation/Market Cap**: Not publicly available - **Key Metric**: Total Funding — $4.5M - **Notable Investors/Partners**: Big Brain Holdings, Bitscale Capital, Tribe Capital, Advisors Anonymous, Polychain Capital, and 11+ additional investors [cbinsights.com](https://www.cbinsights.com/company/wynd-network) - **Growth Signals**: - **Headcount Growth**: +750% YoY (from ~1 to 10 employees), with +21.4% monthly growth [linkedin.com](https://www.linkedin.com/company/wyndlabs) - **Active Hiring**: 13 open positions across engineering, operations, and growth [jobs.ashbyhq.com](https://jobs.ashbyhq.com/wynd-labs) - **International Presence**: Team distributed across the US, Canada, Philippines, Singapore, and Denmark. ## Competitive Advantages - **Decentralized Proxy Network**: Unlike centralized scraping services, Wynd Labs uses a distributed residential proxy network (Grass) that is harder to block, more scalable, and cost-effective. - **Focus on AI Training Data**: Directly addresses the critical bottleneck of high-quality, large-scale web data for training advanced AI models. - **Lean, High-Velocity Team**: Described as a highly motivated team driven by ambitious goals and a strong sense of urgency, allowing for rapid iteration and impact. ## Strategic Focus - **Scale the Grass Network**: Expanding the decentralized proxy network to handle even larger data volumes and more use cases. - **Deepen AI Integration**: Continuing to serve as critical infrastructure for AI companies needing reliable, large-scale web data. - **Expand the Team**: Aggressively hiring to support rapid growth, with a focus on backend engineering, machine learning, and network infrastructure roles. ## Why Work Here - **High-Impact Role**: As a small team (10 people), every hire has significant ownership and direct influence on the product and company direction. - **Remote-First Culture**: The team is distributed across North America, Southeast Asia, and Europe, supporting a remote work environment. - **Competitive Compensation**: For example, the Machine Learning Engineer role lists a salary range of $150K – $220K [getaicareers.com](https://getaicareers.com/companies/wynd-labs). - **Innovative Tech Stack**: Work on cutting-edge problems in decentralized systems, web scraping at scale, and AI infrastructure. - **Strong Growth Trajectory**: With 750% YoY headcount growth and active hiring, the company is in a rapid expansion phase, offering career growth opportunities. ## Sources 1. [wyndlabs.ai](https://wyndlabs.ai/) 2. [jobs.ashbyhq.com](https://jobs.ashbyhq.com/wynd-labs) 3. [linkedin.com](https://www.linkedin.com/company/wyndlabs) 4. [cbinsights.com](https://www.cbinsights.com/company/wynd-network) 5. [getaicareers.com](https://getaicareers.com/companies/wynd-labs) ## Other roles at Wynd Labs - [Junior Software Engineer](https://feeny.ai/job/junior-software-engineer-wynd-labs-remote-62ae7aedx9fa) - [Growth & Communications Lead](https://feeny.ai/job/growth-communications-lead-wynd-labs-remote-4hhb1fx8kyeg) - [Research Crawling Engineer](https://feeny.ai/job/research-crawling-engineer-wynd-labs-remote-hp85wk6x6wen) - [Senior Software Engineer (Backend)](https://feeny.ai/job/senior-software-engineer-backend-wynd-labs-remote-axhwarnw3301) - [Machine Learning Engineer](https://feeny.ai/job/machine-learning-engineer-wynd-labs-remote-whe42v314npy) - [Backend Engineer (TypeScript)](https://feeny.ai/job/backend-engineer-typescript-wynd-labs-remote-97sac88vf4np) - [Web Scraping Specialist](https://feeny.ai/job/web-scraping-specialist-wynd-labs-remote-z2c7v86byynz) - [Network Infrastructure Engineer (DevOps)](https://feeny.ai/job/network-infrastructure-engineer-devops-wynd-labs-ashburn-2ap4z98nryet) — Ashburn, VA - [Data Engineer](https://feeny.ai/job/data-engineer-centerfield-los-angeles-dea9e54mc8jk) — Los Angeles, CA - [Data Engineer](https://feeny.ai/job/data-engineer-fanduel-new-york-sgsnta6t8jzb) — New York, NY