--- title: 'Research Scientist / Engineer – Reinforcement Learning Infrastructure at Luma' canonical: 'https://feeny.ai/job/research-scientist-engineer-reinforcement-learning-infrastructure-luma-sf-bay-v3kqd6ee41ag' type: 'job' last_seen: '2026-09-11' --- # Research Scientist / Engineer – Reinforcement Learning Infrastructure at Luma - **Company:** Luma - **Location:** Sf Bay Area, CA - **Compensation:** $188k–$395k - **Employment:** full-time - **Work type:** hybrid - **Posted:** 2026-07-21 - **Last confirmed live:** 2026-09-11 - **Apply:** https://jobs.gem.com/lumalabs-ai/am9icG9zdDogDs0MGNruilpcY5YMWl-O ## Job description ## About Luma AI Luma's mission is to build multimodal AI to expand human imagination and capabilities. We believe that multimodality is critical for intelligence. To go beyond language models and build more aware, capable and useful systems, the next step function change will come from vision. So, we are working on training and scaling up multimodal foundation models for systems that can see and understand, show and explain, and eventually interact with our world to effect change. ## About the Role Reinforcement learning is how our foundation models go from capable to useful — learning to reason, use tools, and act over long horizons. The RL Infrastructure team builds the systems that make this possible at scale: high-throughput distributed training that couples policy optimization with large fleets of inference workers, environments that expose models to realistic multi-step tasks, and the reward, verification, and evaluation systems that turn model behavior into learning signal. Unlike pretraining, RL at scale is a full-loop systems problem — training, rollout generation, environment execution, and reward computation all run concurrently across thousands of GPUs and must stay fast, stable, and correct together. We are looking for engineers and scientists who have lived this problem: people who have post-trained LLMs with RL, built environments and verifiers from scratch, and debugged what happens when an asynchronous rollout pipeline meets a frontier-scale training run. You will work alongside our research team to design and operate the RL stack for our largest multimodal models. ## Responsibilities - Design, build, and scale distributed RL post-training systems for large multimodal models — orchestrating trainer, rollout, environment, and reward workloads across thousands of GPUs - Build and optimize high-throughput rollout generation, including efficient integration of inference engines (e.g. vLLM, SGLang) into the training loop, weight synchronization, and asynchronous / off-policy training schemes - Design and implement RL environments for agentic and multi-step tasks — sandboxed code execution, tool use, computer use, and multimodal interaction — that are reproducible, hermetic, and scalable to millions of episodes - Build reward infrastructure: verifiable / programmatic rewards, reward model serving, LLM-as-judge pipelines, and defenses against reward hacking - Develop the evaluation, monitoring, and debugging tooling needed to keep large RL runs stable, diagnose convergence and throughput regressions, and understand model behavior mid-run - Advance RL training efficiency and stability: sequence packing for long multi-turn trajectories, KV cache reuse across rollouts, curriculum and task sampling, and resource scheduling across heterogeneous training/inference workloads - Collaborate closely with researchers to turn new post-training ideas (RLVR, agentic RL, long-horizon credit assignment, self-improvement loops) into production-quality training runs ## Experience - Hands-on experience post-training LLMs with reinforcement learning (e.g. PPO / GRPO-family methods, RLHF, RLVR / RL from verifiable rewards) at meaningful scale - Extensive experience with distributed PyTorch training and parallelization strategies (FSDP, Tensor / Pipeline / Expert Parallel) for foundation models - Experience building RL environments, reward functions, verifiers, or evaluation harnesses for LLM agents — including sandboxed execution and multi-turn tool use - Deep familiarity with RL post-training frameworks and their systems tradeoffs (e.g. veRL, OpenRLHF, TRL, Ray-based orchestration) and inference engines used for rollouts (vLLM, SGLang) - Strong understanding of GPU clusters, networking, and communication libraries (NCCL, MPI), and how they behave under mixed training + inference workloads - (Preferred) Experience running RL training across >100 GPUs, including asynchronous or disaggregated trainer/rollout architectures - (Preferred) Experience with containerization and orchestration (Kubernetes, Ray) for large environment fleets and sandboxed workloads - (Preferred) Research contributions in RL for LLMs — reasoning, agents, reward modeling, or long-horizon tasks — or open-source contributions to RL training frameworks ## About Luma ## Company Overview - **One-liner**: Luma builds multimodal AI models that can generate, understand, and operate in the physical world, shipping them in products like Dream Machine for creators and teams. - **Entity Type**: Private (Series C) - **Headquarters**: Redwood City, California, USA - **Founded**: 2021 - **Founders**: Not publicly listed on official sources ## Core Business - **Primary industry**: Artificial Intelligence / Generative AI / Creative AI - **Target customers**: B2B (enterprises, agencies, developers) and B2C (creators, filmmakers) - **Mission or purpose statement**: To build unified general intelligence that can generate, understand, and operate in the physical world. ## Products & Services - **Luma Dream Machine**: Consumer-facing product for turning ideas into compelling visuals (video, image, audio, text) using generative AI. - **Ray 3.2 API**: Production-grade video generation API with full creative control—multi-keyframe (up to 16), V2V up to 20 seconds, reframe, 1080p output, native HDR, and 16-bit EXR export. - **Uni-1.1 API**: Multimodal reasoning model API for image generation (text-to-image, reference-guided) and natural-language image editing, with up to nine reference inputs per request. - **Luma Agents**: Agentic creative workflows that plan, generate, iterate, and refine across every stage of creative work—available in the Luma App and trusted by teams at Publicis Groupe, Serviceplan, Mazda, Dentsu, and Humain. - **Open Physical AI Lab**: An open science effort to solve generalization in physical AI, built in the open for the benefit of all humanity. ## Market Standing - **Valuation/Market Cap**: Not disclosed - **Key Metrics**: Total Funding: $1.3B+ (official site) / $1.057B (LinkedIn); Annual Revenue: $5M (LinkedIn estimate); Users worldwide: 30M+; Team Members: 250+ - **Notable Investors/Partners**: HUMAIN (led Series C), Andreessen Horowitz (led Series B), Amplify Partners (led Series A), Matrix (led Seed Round). Partners include Envato, Comfy, Runware, Flora, Krea, Magnific, Fal, LovArt. - **Growth Signals**: 250+ team members, 30M+ users, $1.3B+ total funding, 45% yearly LinkedIn follower growth, recent product launches (Ray3.2, Uni-1.1, Luma Skills, Wonder Project partnership), $900M Series C raised November 2025. ## Competitive Advantages - **Full-stack AI**: Luma builds its own foundation models end-to-end, rather than wrapping third-party models—giving it unique control over quality, latency, and cost. - **Multimodal reasoning**: Uni-1 understands intention and responds to direction, not just prompts—offering brand intelligence at the model level. - **Cinematic-grade output**: Ray 3.2 delivers 1080p, native HDR, 16-bit EXR export—production-ready for professional film and advertising pipelines. - **Open science commitment**: The Open Physical AI Lab publishes research openly, attracting top-tier research talent and fostering community trust. - **Speed & cost**: Claims less than half the price and latency of comparable models. ## Strategic Focus - **Scaling enterprise adoption**: Custom deployment for creative teams at scale across advertising, film, and gaming production pipelines. - **Agentic workflows**: Building Luma Agents as a force multiplier for creative teams—automating repetitive tasks while maintaining brand consistency. - **Physical AI research**: Advancing the Open Physical AI Lab to solve generalization in physical AI, bridging digital generation with real-world operation. - **API ecosystem growth**: Expanding the developer platform with two tiers (Build and Scale) and SDKs for Python and JavaScript/TypeScript. ## Why Work Here - **Culture & mission**: "We believe real-world physics is the path to general intelligence. We unite research, product, and go-to-market into one engine." Described as a "lean, high-achieving team" building the future of creative intelligence. - **Compensation & benefits**: Competitive compensation, 100% covered medical/dental/vision for employee and family, $1,500/year learning stipend, catered meals, team events, home office setup. - **Work mode**: On-site roles in Redwood City, CA with additional offices in New York, NY; Los Angeles, CA; and international locations (Munich, Paris, London, Berlin). - **Engineering culture**: Emphasis on shipping frontier research directly into products—"frontier research ships straight into the hands of working creatives." Hiring across research, systems engineering, infrastructure, design, and product. - **Hiring process**: Transparent and fast—aims to complete within 2-3 weeks. Application review within a week, then recruiter screen, hiring manager interview, and team/leadership interviews. - **Employee ratings**: 4.1/5.0 on LinkedIn (9 reviews)—Work-Life: 3.6, Compensation: 4.4, Culture: 4.1, Career: 4.2. - **Open roles**: 46+ open positions including Forward Deployed Engineer, Research Scientist, Software Engineer, Robotics Engineer, Simulation Researcher, Site Reliability Engineer, Product Marketing Manager, Enterprise Account Executive, and more. ## Sources 1. [lumalabs.ai - Careers](https://lumalabs.ai/careers) 2. [lumalabs.ai - Official Site](https://lumalabs.ai/) 3. [lumalabs.ai - LLM Info](https://lumalabs.ai/llm-info) 4. [LinkedIn - Luma](https://linkedin.com/company/lumalabsai) ## Other roles at Luma - [Tech Lead Manager, Inference](https://feeny.ai/job/tech-lead-manager-inference-luma-sf-bay-area-wv3qjs47nghh) — Sf Bay Area, CA - [Creative Technologist](https://feeny.ai/job/creative-technologist-luma-london-1edm2ph4be1t) — London, United Kingdom - [Sales Manager - West Coast](https://feeny.ai/job/sales-manager-west-coast-luma-los-angeles-qdhxhg06y85x) — Los Angeles, CA - [Forward Deployed Engineer (EU)](https://feeny.ai/job/forward-deployed-engineer-eu-luma-london-h5nr5x8w2bhd) — London, United Kingdom - [Product Manager, Applied Research](https://feeny.ai/job/product-manager-applied-research-luma-sf-bay-area-tgmbcragfeqt) — Sf Bay Area, CA - [Product Manager, Core Product](https://feeny.ai/job/product-manager-core-product-luma-sf-bay-area-dx7f4b6fnr4d) — Sf Bay Area, CA - [Social Media Editor](https://feeny.ai/job/social-media-editor-luma-los-angeles-ekyz2de2wh6f) — Los Angeles, CA - [Senior Video Editor](https://feeny.ai/job/senior-video-editor-luma-los-angeles-1ynx82dj6hbw) — Los Angeles, CA - [Product Marketing Manager](https://feeny.ai/job/product-marketing-manager-luma-sf-bay-area-s0n3bqsd0qy1) — Sf Bay Area, CA - [Applied Research Scientist / Engineer](https://feeny.ai/job/applied-research-scientist-engineer-luma-new-york-2k805g9e3656) — New York, NY