--- title: 'Research Member of Technical Staff- Training Systems at Rhoda AI' canonical: 'https://feeny.ai/job/research-member-of-technical-staff-training-systems-rhoda-ai-mountain-view-n7d7bvemaev0' type: 'job' last_seen: '2026-09-09' --- # Research Member of Technical Staff- Training Systems at Rhoda AI - **Company:** Rhoda AI - **Location:** Mountain View, CA - **Employment:** full-time - **Posted:** 2026-05-17 - **Last confirmed live:** 2026-09-09 - **Apply:** https://jobs.ashbyhq.com/rhoda-ai/643d697d-45c4-4505-8967-433df7873134/application **Skills:** PyTorch, Distributed Training, Multimodal Training, Tensor Parallelism, Pipeline Parallelism, Sharded Training, FSDP, ZeRO, Performance Measurement, Debugging, JAX, CUDA, Triton, Graph Capture, Operator Fusion, Video Training, Large-scale Training Frameworks, Cluster Topology, Networking, Scheduling > The Staff / Principal ML Systems Engineer owns training systems performance end-to-end, focusing on large-scale multimodal training efficiency, scalability, and correctness. This role involves diagnosing performance bottlenecks, designing parallelism strategies, and partnering with researchers to translate model inn... ## Job description At Rhoda AI, we’re building the next generation of generalist intelligent robots. We own the full robotics stack from high-performance hardware and robot systems to the infrastructure and state-of-the-art foundation world models that control our robots. Our robots are designed to be generalists capable of operating in complex, real-world environments and handling long-tail edge cases, made possible by our cutting edge research and end-to-end system design. We've raised over $450M and are investing aggressively in model research, infrastructure, hardware development, and manufacturing scale-up to make generalist robotics a reality. We're looking for a Staff / Principal ML Training Systems Engineer to own training systems performance end-to-end. You will define how our models train at scale — driving efficiency, scalability, and correctness across large-scale multimodal training. This is a core systems role, not infrastructure support. Your work directly determines how efficiently we use compute, how well models scale across thousands of GPUs, and how quickly research can iterate. ## What You'll Do Own training performance end-to-end - Diagnose and improve performance of large-scale multimodal training (vision, video, proprioception, actions, language) - Build systematic performance attribution: step-time decomposition (compute vs communication vs input pipeline), scaling curves across cluster sizes, and bottleneck identification and prioritization - Drive measurable gains in: - Distributed efficiency (comm/compute overlap, bucketization, topology-aware mapping, parallelism strategies) - Compute efficiency (kernel hotspots, operator fusion, attention optimization, framework/runtime overhead) - Memory efficiency (activation checkpointing, sequence packing/bucketing, fragmentation reduction) Design training systems (not just tune them) - Define and evolve parallelism strategies: data / tensor / pipeline / sharding / hybrid approaches - Improve execution efficiency through communication scheduling and overlap, graph capture and execution optimization, and runtime-level improvements - Contribute to and extend training frameworks where needed Make performance observable and measurable - Establish source-of-truth performance metrics: step-time breakdowns, MFU / throughput / scaling efficiency - Build tools to identify bottlenecks quickly, track performance across model families, and compare scaling behavior across configurations - Develop regression detection: microbenchmarks, performance baselines, and automated detection of efficiency regressions Partner deeply with researchers - Work side-by-side with research scientists and research engineers — no silos - Translate model innovations into scalable, efficient implementations - Advise on training tradeoffs for robotics world models: long-horizon sequences, rollout/evaluation cadence, multimodal and variable-length data Collaborate on cluster-level efficiency - Work with infrastructure/SRE teams to improve utilization across large distributed jobs, impact of network and collective performance on training, and topology-aware job placement and scaling behavior ## What We're Looking For - Proven track record improving large-scale distributed training performance - Deep hands-on experience with modern ML stacks (PyTorch required; JAX a plus) - Strong understanding of data / tensor / pipeline parallelism, sharded training (FSDP / ZeRO-style), communication patterns and overlap strategies, and scaling behavior across large GPU clusters - Strong systems intuition — ability to reason across compute, communication, and memory bottlenecks - Exceptional debugging and measurement ability: turn "training is slow" into clear bottlenecks, experiments, and validated improvements - High ownership mindset and comfort in a fast-moving environment ## Nice to Have (But Not Required) - GPU kernel or compiler-level experience (CUDA, Triton, graph capture, operator fusion) - Experience with multimodal or video training (variable-length sequences, packing/bucketing) - Experience working on large-scale training frameworks or distributed runtimes - Familiarity with cluster topology, networking, and large-scale scheduling effects ## Why This Role - Direct leverage on research velocity — every efficiency gain you make accelerates model iteration across the entire research team - Own the scalability and performance of large-scale multimodal training for real-world embodied intelligence, not static benchmarks - Improvements you make compound across every training run the company executes — high ownership, high impact, small elite team ## About Rhoda AI ## Company Overview - **One-liner**: Rhoda AI builds robot foundation models that learn from internet-scale video to enable manipulation-capable robots to generalize in real-world industrial environments. - **Entity Type**: Private (Seed + Series A; $680M total funding) - **Headquarters**: Palo Alto, CA, United States - **Founded**: 2024 - **Founders**: Jagdeep Singh (CEO), Eric Chan (Chief Scientist), Changan Chen (Chief Research Officer), Andrew Wooten, Gordon Wetzstein (Scientific Advisor) [rhoda.ai/team](https://www.rhoda.ai/team) ## Core Business - **Primary industry**: Robotics, Artificial Intelligence, Hardware - **Target customers**: B2B – industrial enterprises in automotive, manufacturing, logistics, and e-commerce - **Mission/purpose**: “Building at the frontier of embodied intelligence” and “Designing intelligence for the real world” – deploying general-purpose robot foundation models that adapt to the variability of commercial and industrial environments. [rhoda.ai](https://www.rhoda.ai/) | [linkedin.com](https://www.linkedin.com/company/rhoda-ai) ## Products & Services - **Rhoda Robot Platform (Hardware + Software)**: A turnkey robotic system with custom actuators (25kg rated, 40kg peak payload), safety-rated vision, wheel-base, and brakes in every actuator. The platform runs the **Direct Video Action** model, which pre-trains on over a million real-world videos and post-trains on task-specific robot data. Long-context windows enable single-shot learning from human demonstrations without retraining. [rhoda.ai](https://www.rhoda.ai/) - **FutureVision Intelligence Layer**: A proprietary architecture that handles real-world industrial tasks autonomously by generalizing beyond conventional VLA pipelines. [rhoda.ai](https://www.rhoda.ai/) - **Deployment Services**: End-to-end solutions for tasks such as returns processing, bearing decanting, container breakdown, and material transport in logistics and e-commerce facilities. [rhoda.ai](https://www.rhoda.ai/) ## Market Standing - **Valuation**: Not publicly disclosed - **Key Metric**: Total funding **$680M** – Seed round $67.4M (April 2024), Series A $162.6M (April 2025), and Series A $450M (March 2026) with 4 investors. [linkedin.com](https://www.linkedin.com/company/rhoda-ai) - **Notable Investors/Partners**: Not individually named in available sources; partners include customers in automotive, manufacturing, logistics, and e-commerce. [rhoda.ai](https://www.rhoda.ai/) - **Growth Signals**: Headcount growth of **+20% monthly** (LinkedIn reports 55 employees; Built In reports 73 employees – conflicting data). Talent sourced from companies like QuantumScape, Flexiv Robotics, Scale AI, Google, Meta, NVIDIA, and Stanford University. [linkedin.com](https://www.linkedin.com/company/rhoda-ai) | [builtin.com](https://builtin.com/company/rhoda-ai) ## Competitive Advantages - **Direct Video Action architecture** with web-scale video pre-training gives a strong prior on motion, physics, and dynamics, enabling generalization beyond lab settings. - **Custom high-performance actuators** (25kg rated, 40kg peak) with safety-rated brakes and vision – designed for 3+ years of continuous operation at rated payload. - **Long-context in-context learning** allows the robot to perform tasks single-shot from human demonstrations without retraining. - **Turnkey industrial deployments** that handle ambiguity and variability (e.g., debris, transparent bags, heavy boxes) – tasks customers initially believed could not be automated. [rhoda.ai](https://www.rhoda.ai/) ## Strategic Focus - Scaling robot intelligence through a pre-training/post-training paradigm, moving from controlled labs to reliable production environments. - Expanding commercial deployments across automotive, manufacturing, logistics, and e-commerce verticals. - Pursuing a path to **physical AGI** by combining foundation models with custom hardware for real-world adaptability. [rhoda.ai](https://www.rhoda.ai/) ## Why Work Here - **Culture**: “Building at the frontier of embodied intelligence” – emphasis on hard problems, first-principles thinking, and category-defining systems. [linkedin.com](https://www.linkedin.com/company/rhoda-ai) - **Work policy**: Hybrid workspace (employees combine remote and on-site work), with offices in Mountain View and Palo Alto, CA. [builtin.com](https://builtin.com/company/rhoda-ai) - **Engineering culture**: Strong research focus (Chief Research Officer, Chief Scientist, many research engineers/scientists). Recent job postings include Research Scientist, Robot Perception Engineer, Mechanical Engineer, Fullstack Engineer, and Product Manager. [builtin.com](https://builtin.com/company/rhoda-ai) - **Notable perks**: Not explicitly listed, but the company is well-funded ($680M) and at an early stage (founded 2024), offering significant growth opportunity and impact. ## Sources 1. [rhoda.ai](https://www.rhoda.ai/) – Company website (product, technology, commercial applications) 2. [rhoda.ai/careers](https://www.rhoda.ai/careers) – Careers page 3. [rhoda.ai/team](https://www.rhoda.ai/team) – Leadership team and co-founders 4. [builtin.com](https://builtin.com/company/rhoda-ai) – Built In profile (employee count, offices, job listings, hybrid policy) 5. [linkedin.com](https://www.linkedin.com/company/rhoda-ai) – LinkedIn company page (funding, headcount, executive list, talent sources) ## Other roles at Rhoda AI - [Robotics Test Engineer](https://feeny.ai/job/robotics-test-engineer-rhoda-ai-mountain-view-fawz7jcaj9pv) — Mountain View, CA - [Senior Data Project Manager](https://feeny.ai/job/senior-data-project-manager-rhoda-ai-mountain-view-3r8qd12vvp3g) — Mountain View, CA - [Senior Hardware Reliability Engineer](https://feeny.ai/job/senior-hardware-reliability-engineer-rhoda-ai-mountain-view-tb5a3r2wpf2a) — Mountain View, CA - [Supply Chain Manager](https://feeny.ai/job/supply-chain-manager-rhoda-ai-mountain-view-at62gnb3hppz) — Mountain View, CA - [Robot Pilot](https://feeny.ai/job/robot-pilot-rhoda-ai-munich-ae615mvnydam) — Munich, Germany - [Senior Fleet Software Engineer](https://feeny.ai/job/senior-fleet-software-engineer-rhoda-ai-mountain-view-bbmg1fzakq3b) — Mountain View, CA - [Fleet Response Engineer](https://feeny.ai/job/fleet-response-engineer-rhoda-ai-mountain-view-qk40e501tskz) — Mountain View, CA - [Senior Electrical Engineer (Hands)](https://feeny.ai/job/senior-electrical-engineer-hands-rhoda-ai-mountain-view-63xfzs2xajpn) — Mountain View, CA - [Research Member of Technical Staff- Robot Learning Systems & Reliability](https://feeny.ai/job/research-member-of-technical-staff-robot-learning-systems-reliability-rhoda-ai-azp0647ma7en) — Mountain View, CA - [Strategic Finance Manager](https://feeny.ai/job/strategic-finance-manager-rhoda-ai-mountain-view-efncjc6jbpjr) — Mountain View, CA