--- title: 'Member of Technical Staff, Machine Learning at Sieve' canonical: 'https://feeny.ai/job/member-of-technical-staff-machine-learning-sieve-san-francisco-1hgtcn0595d1' type: 'job' last_seen: '2026-09-10' --- # Member of Technical Staff, Machine Learning at Sieve - **Company:** Sieve - **Location:** San Francisco, CA - **Employment:** full-time - **Work type:** onsite - **Posted:** 2026-07-13 - **Last confirmed live:** 2026-09-10 - **Apply:** https://jobs.ashbyhq.com/sieve/dbff6c5e-341c-4daf-9563-16dbeb4b1fad ## Job description ## About Us Sieve is a multi-modal lab curating the world's highest-quality training datasets — spanning video, audio, images, text, and 3D. We combine exabyte-scale data infrastructure and novel multimodal understanding techniques that push the frontier of foundation models. Video alone makes up 80% of internet traffic, and across modalities, data has become the enabling medium powering creativity, communication, gaming, AR/VR, and robotics. Sieve exists to solve the biggest bottleneck in the growth of these applications: high-quality training data. We partner with top AI labs and did $XXM last quarter alone, as a team of ~30 people. We also raised our Series A from Tier 1 firms such as [Matrix Partners](https://matrix.vc/), [Swift Ventures](https://www.swift.vc/), [Y Combinator](https://www.ycombinator.com/), and [AI Grant](https://aigrant.com/). ## Why Now Sieve is one of the most capital-efficient teams in AI — roughly 30 people serving the world's leading AI labs across every major data modality. You'll join early, own problems end-to-end, and watch your work ship directly into the models defining the frontier. ## About the Role As a Machine Learning Engineer at Sieve, you'll own the entire ML lifecycle — from understanding customer problems, to designing datasets, improving models, building evaluation systems, and shipping production pipelines that deliver measurable improvements in dataset quality. You'll work directly with frontier AI labs to understand difficult data problems, then build end-to-end systems that solve them. One week you might fine-tune a multimodal model to improve recall on a difficult edge case. The next you might engineer a VLM-based QA pipeline, design a new evaluation framework, or run a large-scale filtering pipeline on millions of hours of multimodal data. We're looking for engineers who enjoy owning problems end-to-end, from understanding customer requirements through shipping production ML systems that measurably improve dataset quality. ## What You'll Do - Own model quality for customer-facing video understanding problems - Fine-tune vision-language and multimodal foundation models for specialized tasks - Build automated evaluation and QA pipelines using frontier models like Gemini, GPT, Claude, and open-source VLMs - Design high-precision filtering, ranking, retrieval, and labeling systems over internet-scale video datasets - Create datasets, benchmarks, and evaluation frameworks that continuously improve model quality - Develop production ML pipelines spanning preprocessing, inference, post-processing, and quality validation - Work directly with frontier AI labs to translate ambiguous requirements into scalable ML systems - Ship improvements quickly, measure results, and iterate based on real-world performance ## Requirements - Strong Python engineer with experience building production ML systems - Experience training, fine-tuning, or deploying modern deep learning models - Comfortable working with PyTorch and modern foundation models - Excellent intuition for evaluation, dataset quality, precision/recall tradeoffs, and edge cases - Enjoys rapidly prototyping with new AI models and APIs - Comfortable owning projects from customer problem to internal pipelines to deployed solution - Strong communicator who enjoys working directly with customers and cross-functional teams - Excited by video, multimodal AI, and frontier foundation models - In-person at our SF HQ *all roles at Sieve require you to be onsite in San Francisco 5 days per week ## About Sieve ## Company Overview - **One-liner**: Sieve is an AI research lab and platform focused on building high-quality multimodal data (video, audio, image, and interaction data) and infrastructure for training frontier AI models. - **Entity Type**: Private (Seed-stage) - **Headquarters**: San Francisco, California, USA - **Founded**: 2022 - **Founders**: Mokshith Voodarla (CEO) and Abhinav "Abhi" Ayalur (CTO) ## Core Business - Primary industry: Artificial Intelligence / Data Infrastructure / Computer Vision - Target customers: Frontier AI labs, Fortune 100 companies, and fast-growing AI startups working on generative media, robotics, computer use, world models, and agentic systems (primarily B2B Enterprise). - Mission statement: To build the data and environments that frontier AI labs use to train the next generation of multimodal systems. ## Products & Services - **Sieve Platform (sievedata.com)**: A development platform to discover, design, and run AI features at scale. It offers pre-built, production-ready apps and models (e.g., `sieve/speech_transcriber` for transcription, `sieve/seamless_text2text` for translation) that can be used via a single API call. The platform provides no-code interfaces, reliable infrastructure, and primitives like queues and async jobs. It is focused on making it easy to combine multiple AI models into complex applications (e.g., video dubbing, translation). [docs.sievedata.com] - **Sieve Data Lab (sieve.ai)**: A multimodal data lab offering high-quality video, audio, image, and interaction datasets for frontier AI. Services include custom data collection, indexing billions of data points, filtering for semantics and rights, adding dense annotations (transcripts, object labels, UI events), and secure delivery with end-to-end encryption and SOC 2 Type 2 compliance. [sieve.ai] - **Video AI API (Beta)**: A recently launched API that allows developers to use Sieve's pre-built AI apps and models for video processing in a single call. [sievedata.com] ## Market Standing - **Valuation/Market Cap**: Not publicly disclosed. - **Key Metric**: Raised approximately **$4M** in a seed round in 2022. [ycombinator.com, sievedata.com] - **Notable Investors/Partners**: Y Combinator (Winter 2022 batch, with Diana Hu as Primary Partner), AI Grant (Nat Friedman & Daniel Gross). [ycombinator.com, sieve.ai/about] - **Growth Signals**: The company has evolved from a general AI platform to a specialized research lab focusing on the critical bottleneck of training data for multimodal AI. It has attracted the trust of "frontier AI labs, Fortune 100 companies, and fast-growing AI startups." The team size is listed as 18 on Y Combinator (2022) and they are actively hiring for multiple engineering roles through the end of 2026, indicating scaling. [ycombinator.com, jobs.ashbyhq.com] ## Competitive Advantages - **Focus on Data as a Moat**: Sieve is explicitly structured as a "research lab exclusively focused on video data," identifying that high-quality training data is the primary bottleneck for the next generation of AI (video, audio, robotics, agentic systems). [ycombinator.com] - **End-to-End Data Service**: They provide a complete pipeline from custom collection and indexing to dense annotation and secure delivery, unlike many competitors who only offer one piece. - **Deep Research Partnerships**: The team works directly with leading AI researchers to understand model bottlenecks and design high-signal datasets, creating a tight feedback loop between research needs and data production. - **Strong Founder-Market Fit**: Founders Mokshith Voodarla (Scale AI, NVIDIA) and Abhinav Ayalur (NVIDIA, Niantic) have direct experience in building computer-vision systems and experiencing the data difficulties first-hand. ## Strategic Focus - **Deliver Training Data for Frontier AI**: Sieve’s current priority is to scale its exabyte-scale infrastructure and novel multimodal understanding techniques to create datasets for video, audio, images, software, robotics, and interactive worlds. [sieve.ai/about] - **Enable Complex AI Use Cases**: The platform side focuses on making it easy for developers to combine multiple models to build production-ready AI features (like translation, dubbing) quickly. - **Data Security and Compliance**: Emphasizing secure, compliant data delivery with SOC 2 controls is a key priority for servicing Fortune 100 companies. ## Why Work Here - **Culture**: The company describes itself as "tight-knit, fast-moving, and deeply customer-oriented." It is a "research lab to the core," combining frontier research with infrastructure and customer partnership. [sieve.ai/about] - **Team**: The team brings together experience from NVIDIA, Scale AI, Zoox, and Niantic, providing exposure to top-tier expertise in AI and infrastructure. [sieve.ai/about] - **Impact**: Employees will be working on the critical layer of data for the next generation of AI models (generative media, robotics, agents), offering a rare position at the forefront of AI research. [ycombinator.com] - **Remote/Hybrid/Office**: The company is based in San Francisco, CA. - **Open Roles**: As of mid-2026, they are actively hiring for **Applied Research Engineer**, **Distributed Systems Engineer**, **Product Engineer**, and **Software Engineer**, with salary ranges listed as $150K - $300K depending on role and experience. [jobs.ashbyhq.com, ycombinator.com] ## Sources 1. [docs.sievedata.com](https://docs.sievedata.com/guide/intro) 2. [jobs.ashbyhq.com](https://jobs.ashbyhq.com/sieve) 3. [sieve.ai](https://www.sieve.ai/) 4. [sieve.ai/about](https://www.sieve.ai/about) 5. [ycombinator.com](https://www.ycombinator.com/companies/sieve) ## Other roles at Sieve - [Business Development](https://feeny.ai/job/business-development-sieve-san-francisco-crb7611fxcpw) — San Francisco, CA - [Growth Associate](https://feeny.ai/job/growth-associate-sieve-india-8qjbpndm06dg) — India - [Member of Technical Staff, Forward Deployed](https://feeny.ai/job/member-of-technical-staff-forward-deployed-sieve-san-francisco-kcexg4yd4whn) — San Francisco, CA - [Product Operations Lead](https://feeny.ai/job/product-operations-lead-sieve-san-francisco-3sj7sbk4ncqa) — San Francisco, CA - [Member of Technical Staff, Reliability](https://feeny.ai/job/member-of-technical-staff-reliability-sieve-san-francisco-xefhm0xxfcm0) — San Francisco, CA - [Member of Technical Staff, Product](https://feeny.ai/job/member-of-technical-staff-product-sieve-san-francisco-71h5z7nn9gvd) — San Francisco, CA - [Research & Product Lead](https://feeny.ai/job/research-product-lead-sieve-san-francisco-vcvayswg9p90) — San Francisco, CA - [Member of Technical Staff](https://feeny.ai/job/member-of-technical-staff-sieve-san-francisco-wfh094cr8jd1) — San Francisco, CA - [Member of Technical Staff, Infrastructure](https://feeny.ai/job/member-of-technical-staff-infrastructure-sieve-san-francisco-ajs7dj77750a) — San Francisco, CA - [Member of Technical Staff, Machine Learning](https://feeny.ai/job/member-of-technical-staff-machine-learning-bjak-china-pjzg7hagkbjz) — China