--- title: 'Research Engineer, Multimodal Data at Eventual' canonical: 'https://feeny.ai/job/research-engineer-multimodal-data-eventual-san-francisco-q7myvdm4xpe4' type: 'job' last_seen: '2026-09-09' --- # Research Engineer, Multimodal Data at Eventual - **Company:** Eventual - **Location:** San Francisco, CA - **Compensation:** $150k–$250k - **Employment:** full-time - **Posted:** 2026-04-29 - **Last confirmed live:** 2026-09-09 - **Apply:** https://jobs.ashbyhq.com/eventualcomputing/20e2b01d-969b-454b-98f6-bdce0af5b32c ## Job description ## ABOUT EVENTUAL From humanoid robots to autonomous vehicles, every Physical AI model is trained on petabytes of video, lidar, radar, and sensor data. Today's data platforms (Databricks, Snowflake) were built for spreadsheet-like analytics, not video corpora. And understanding that video still means paying a person to watch it, ten dollars an hour of footage at the low end. So teams check a sample and hope it represents the rest. The footage grows every year; the budget to look at it doesn't. Eventual was founded in 2022 to close that gap. Our open-source engine, Daft http://daft, is purpose-built for multimodal AI: 2 PB/day at Amazon, 60-100 PB at another FAANG company, and in production at companies like Mobileye, TogetherAI. On top of it we're building the infrastructure that finds any situation you can describe across a fleet's entire video history, and turns it into a training set or an alert someone can still act on. We fine-tune and run the vision models ourselves, which makes indexing every hour cheaper than annotating a sample. We're building this with the top Physical AI labs and GPU cloud providers. We've raised $30M from investors like Felicis, CRV, Y Combinator, and angels from the co-founders of Databricks and Perplexity. Our team comes from AWS, Lyft, and Tesla. We powered the last generation of Physical AI in self-driving; now we're doing it for the next. Join our small (but powerful!) team, 4 days/week in our SF Mission District office. ## YOUR ROLE As a Research Engineer on the Visual Understanding team, you'll own the layer that makes petabytes of video queryable by content. Physical AI teams have video, lidar, radar, and sim outputs scattered across object stores with no way to find what they need without weeks of human annotation. Eventual runs vision/language models and pipelines over every clip in a corpus along axes the customer cares about (gripper type, failure mode, object class, scene, motion density), so a researcher can ask "left-arm grasp failures on deformable objects" and get a curated dataset in minutes. You'll define the roadmap for our visual understanding capabilities, train and select the models that make corpus-scale annotation tractable at single-digit cents per hour of video, and build the rich datasets that go on to train customer models. This is a applied research role — meaning you'll read papers and run experiments, but you ship to production and your work has a real impact on our customers’ models and robots. ## KEY RESPONSIBILITIES - Own the visual understanding roadmap end-to-end: from picking the model family for a customer's taxonomy to landing it in production inference at corpus scale. - Train, fine-tune, and evaluate VLMs, VQA models, embedding models, and CNNs against customer datasets and benchmarks. - Drive down per-clip annotation cost — model selection, distillation, batching, decode pipelining — so "annotate every clip in a 10K-hour corpus" stays economical. - Build the rich, queryable datasets that customers train on: design taxonomies with researchers, instrument quality, version the outputs. - Partner with the dataloading and storage teams so visual understanding outputs flow into the index and on to the GPU without re-engineering. - Work directly with researchers at our partner labs — your shortest feedback loop is their next training iteration. ## WHAT WE LOOK FOR - Strong familiarity with modern vision and multimodal models — convolution nets, VLMs, VQA, embeddings — and a sense for the SOTA that's actually deployable today vs. on a leaderboard. - Experience running these models at scale on real video and sensor data, ideally for perception tasks (detection, tracking, segmentation, retrieval, captioning). - Background from a perception team at a self-driving, robotics, or visual-data company — or equivalent depth from a research lab. - Comfortable with cloud infrastructure and large-scale data processing — you don't need to be a distributed-systems engineer, but you've shipped jobs that run on thousands of GPU-hours of video. ## NICE TO HAVE - Experience training and evaluating vision or multimodal models (not just calling APIs). - ML/AI research background — papers, citations, or a research org on your resume. - Worked on embeddings, retrieval, or content-aware search at scale. - Experience designing labeling taxonomies or running annotation programs. ## PERKS & BENEFITS - In-person, tight-knit team — 4 days/week in our SF Mission office. - Competitive comp and meaningful startup equity. - Catered lunches and dinners for SF employees. - Commuter benefit. - Team-building events and poker nights. - Health, vision, and dental coverage. - Flexible PTO. - Latest Apple equipment. - 401(k) plan with match. If you're excited about being on the team that turns petabytes of raw video into the training data for the next generation of Physical AI, we'd love to talk. ## About Eventual ## Company Overview - **One-liner**: Eventual builds an open-source, high-performance data engine (Daft) purpose-built for AI and multimodal workloads, making querying images, video, audio, and text as intuitive as working with tables. - **Entity Type**: Private (Startup, Y Combinator W22 batch) - **Headquarters**: San Francisco, California, USA (Mission District) - **Founded**: 2022 - **Founders**: Sammy Sidhu (CEO) and Jay Chia ## Core Business - **Primary Industry**: AI Infrastructure / Data Engineering / Multimodal Data Processing - **Target Customers**: B2B, Enterprise AI teams, foundation model developers, autonomous vehicle companies, and any organization processing large-scale multimodal data (images, video, audio, text). - **Mission / Purpose**: To make querying any kind of data—images, video, audio, text—as intuitive as working with tables, and powerful enough to scale to production AI workloads. ## Products & Services - **Daft (Open-Source Engine)**: The company’s core product. A high-performance, distributed data engine designed for AI and multimodal data. It handles petabytes of data daily, coordinates with external APIs, manages GPU clusters, and handles failures that traditional engines (like Spark/Databricks) cannot. Available as an open-source Python library on GitHub (5,570+ stars). - **Daft CLI**: A command-line tool for spinning up and managing Ray clusters for the Daft Query Engine. - **Daft Examples & Benchmarking**: Public repositories providing usage examples and distributed query benchmarking tools. ## Market Standing - **Valuation / Market Cap**: Not publicly disclosed. - **Key Metric**: Total Funding – Backed by Y Combinator, Caffeinated Capital, Array.vc, and angel investors including the co-founders of Databricks and Perplexity. Specific funding amounts are not publicly disclosed. - **Notable Investors / Partners**: Y Combinator (Winter 2022 batch), Caffeinated Capital, Array.vc. Key customers/partners include **Amazon**, **Mobileye**, **Together AI**, and **CloudKitchens**. - **Growth Signals**: Quadrupled team size within the past year (from ~4 to ~18 employees). Daft already processes petabytes of data daily at major companies. Actively hiring to double the team again. ## Competitive Advantages - **Purpose-Built for AI**: Unlike Databricks, Snowflake, or Spark, which were designed for tabular/analytics workloads, Daft is built from the ground up for multimodal AI data (images, video, audio, text). - **Open-Source + Performance**: Daft is open-source (Apache 2.0), giving teams full control, while delivering performance competitive with proprietary engines. It handles GPU cluster orchestration and API coordination natively. - **Founding Team & Advisors**: Founders have deep backgrounds in HPC, deep learning, self-driving cars (DeepScale/Tesla, Lyft L5), and ML infrastructure (Freenome). The team includes engineers from Databricks, AWS, Nvidia, Pinecone, and GitHub Copilot. - **Traction with Top-Tier Customers**: Already deployed at Amazon, Mobileye, Together AI, and CloudKitchens, validating product-market fit with demanding enterprise workloads. ## Strategic Focus - **Scale the Team**: Aggressively hiring across engineering (Product, Systems, HPC, Research) to double the current headcount of 18. - **Expand Daft’s Capabilities**: Continue building out Daft’s multimodal query engine to handle even more complex AI workloads and modalities. - **Deepen Enterprise Adoption**: Grow usage within existing customers (Amazon, etc.) and acquire new large-scale AI teams. - **Community Growth**: Grow the open-source community around Daft (currently 5,570+ stars on GitHub). ## Why Work Here - **High Impact**: Work directly on infrastructure that powers breakthrough AI applications (foundation models, autonomous vehicles). “Your work directly determines whether AI teams can build breakthrough applications or get stuck rebuilding infrastructure.” - **World-Class Team**: Join a small, elite team of 18 engineers from Databricks, AWS, Nvidia, Pinecone, GitHub Copilot, and Tesla. “Quadrupling our size within a year” signals rapid scaling and opportunity. - **In-Office Culture**: 4 days/week in the San Francisco Mission District office. Tight-knit, collaborative environment with catered lunch/dinner, poker nights, and team-building events. - **Compensation & Benefits**: Competitive salary ($150K–$250K for engineering roles) + startup equity. Benefits include health/vision/dental, flexible PTO, 401k with match, commuter benefits, and latest Apple equipment. - **Values-Driven**: “Take pride in our work,” “Create clarity,” “Be curious,” “Own the outcome.” Emphasis on writing high-quality code, tackling hard distributed systems problems, and direct ownership. - **Growth Trajectory**: Early-stage (18 people) with strong traction and backing. Opportunity to shape the product and culture as the company scales. ## Sources 1. [Y Combinator – Eventual Company Profile](https://www.ycombinator.com/companies/eventual) 2. [Eventual Official Careers Page](https://www.eventual.ai/careers) 3. [Eventual GitHub Organization](https://github.com/eventual-inc) 4. [Y Combinator – Eventual Jobs Page](https://www.ycombinator.com/companies/eventual/jobs) 5. [Eventual Job Board (Ashby)](https://jobs.ashbyhq.com/eventualcomputing) ## Other roles at Eventual - [Software Engineer, Multimodal Backend Systems](https://feeny.ai/job/software-engineer-multimodal-backend-systems-eventual-san-francisco-szayste712fy) — San Francisco, CA - [Software Engineer, Data Systems](https://feeny.ai/job/software-engineer-data-systems-eventual-san-francisco-je7sws8z9p6d) — San Francisco, CA