Sieve

Member of Technical Staff, Forward Deployed at Sieve (San Francisco, CA)

Sieve· San Francisco, CA·

Role details

Work type
Onsite
Employment
Full-Time

Job description

About Us

Sieve is a multi-modal lab curating the world's highest-quality training datasets — spanning video, audio, images, text, and 3D. We combine exabyte-scale data infrastructure and novel multimodal understanding techniques that push the frontier of foundation models. Video alone makes up 80% of internet traffic, and across modalities, data has become the enabling medium powering creativity, communication, gaming, AR/VR, and robotics. Sieve exists to solve the biggest bottleneck in the growth of these applications: high-quality training data.

We partner with top AI labs and did $XXM last quarter alone, as a team of ~30 people. We also raised our Series A from Tier 1 firms such as Matrix Partners, Swift Ventures, Y Combinator, and AI Grant.

Why Now

Sieve is one of the most capital-efficient teams in AI — roughly 30 people serving the world's leading AI labs across every major data modality. You'll join early, own problems end-to-end, and watch your work ship directly into the models defining the frontier.

About the Role

As a Forward Deployed Engineer at Sieve, you’ll work on highly specific dataset problems for frontier AI labs. We're looking for someone with a strong bias to action who likes working closely with customers, untangling messy requirements, and shipping fast. You’ll work closely with customers and internal teams to understand exactly what data is needed, then turn ambiguous requirements into production systems that can find, generate, filter, transform, evaluate, and package high-quality video datasets at scale.

Requirements

  • Comfortable working directly with customers or external teams to translate ambiguous needs into concrete technical systems
  • Strong Python developer with hands-on experience in PyTorch or similar ML frameworks
  • Experience building custom algorithms, model workflows, or large-scale data pipelines
  • Strong intuition for dataset quality, filtering, labeling, evaluation, and edge cases
  • Able to break customer-level goals down into the models, heuristics, infrastructure, and QA steps needed to deliver
  • Writes clean, maintainable code and can move quickly without creating brittle systems
  • Deep passion for video, media technologies, and frontier AI applications
  • Motivated by delivering end-to-end outcomes, not just training models or writing research code
  • Bonus: Experience with large-scale video, audio, or multimodal data processing
  • Bonus: Active contributor to open source projects
  • Bonus: Experience as an early hire at a startup
  • In-person at our SF HQ

Benefits

  • 401k + Full Health Insurance
  • Breakfast, Lunch, and Dinner covered and your choice of snacks
  • Ubers covered home

*all roles at Sieve require you to be onsite in San Francisco 5 days per week

Why work at Sieve

  • Culture: The company describes itself as "tight-knit, fast-moving, and deeply customer-oriented." It is a "research lab to the core," combining frontier research with infrastructure and customer partnership. [sieve.ai/about]
  • Team: The team brings together experience from NVIDIA, Scale AI, Zoox, and Niantic, providing exposure to top-tier expertise in AI and infrastructure. [sieve.ai/about]
  • Impact: Employees will be working on the critical layer of data for the next generation of AI models (generative media, robotics, agents), offering a rare position at the forefront of AI research. [ycombinator.com]
  • Remote/Hybrid/Office: The company is based in San Francisco, CA.
  • Open Roles: As of mid-2026, they are actively hiring for Applied Research Engineer, Distributed Systems Engineer, Product Engineer, and Software Engineer, with salary ranges listed as $150K - $300K depending on role and experience. [jobs.ashbyhq.com, ycombinator.com]

Application questions