--- title: 'Research Scientist - Post-training / RL at Epsilon Labs, Inc.' canonical: 'https://feeny.ai/job/research-scientist-post-training-rl-epsilon-labs-inc-san-francisco-88p5ss3dcaga' type: 'job' last_seen: '2026-09-16' --- # Research Scientist - Post-training / RL at Epsilon Labs, Inc. - **Company:** Epsilon Labs, Inc. - **Location:** San Francisco, CA - **Employment:** full-time - **Work type:** onsite - **Posted:** 2026-08-31 - **Last confirmed live:** 2026-09-16 - **Apply:** https://jobs.ashbyhq.com/epsilon-health/b4131b45-ea61-4826-8e33-bc6688e9fca7 ## Job description ## About Us We're tackling one of healthcare's most critical challenges in medical imaging and diagnostics. Our company operates at the intersection of cutting-edge AI and clinical practice, building technology that directly impacts patient outcomes. We've assembled one of the industry's most comprehensive and diverse medical imaging datasets and have a proven product-market fit with a substantial customer pipeline already in place. ## Role Overview We're seeking a Research Scientist with deep expertise in post-training and reinforcement learning to join our ML Research team. You'll be at the forefront of developing and deploying state-of-the-art multimodal models for clinical use in radiology settings. This role owns every stage after pretraining: supervised fine-tuning, reward modeling, reinforcement learning against verifiable and learned reward signals, reasoning and tool-use training, and inference-time strategy. You'll work with one of the largest and most diverse medical imaging datasets in the industry, advancing the state-of-the-art in grounded report generation, reward design, and inference-time reasoning while maintaining the clinical rigor required for healthcare deployment. ## Key Responsibilities - Design reinforcement learning with verifiable rewards for report generation, including clinical label and entity-relation matching, grounding IoU, measurement accuracy, and reporting schema compliance. - Extend reinforcement learning to unverifiable and noisy objectives such as report quality and clinical usefulness, using learned reward models and radiologist feedback pipelines (RLHF) built on expert preferences and report edits. - Run GRPO-family algorithms with complex multi-reward objectives, tuning reward composition and diagnosing reward hacking, entropy collapse, and diversity loss. - Train explicit reward models, including multimodal reward models conditioned on the image, with both outcome and process supervision. - Train chain-of-thought reasoning over image regions, including evidence localization and verification loops that keep reasoning grounded in the image rather than in language priors. - Train multimodal tool use — windowing, zoom and crop, detector and segmentation calls, prior study retrieval — with credit assignment across multi-turn trajectories. - Develop inference-time methods including best-of-N sampling against reward models and grounding-aware decoding, and distill the resulting gains back into the policy. - Tune output stylization to institutional reporting conventions, keeping style rewards separated from clinical content rewards. - Stay current with cutting-edge research in reinforcement learning, reward modeling, and multimodal post-training. - Drive research and technical excellence through conference publications and technical blog posts, establishing best practices for post-training medical VLMs at scale. ## Qualifications - 6+ years of academia/industry experience in reinforcement learning, post-training, or multimodal machine learning - Deep expertise in post-training large language or vision-language models (e.g., Qwen-VL, InternVL, LLaVA, or similar architectures) - Strong foundation in modern post-training and reinforcement learning techniques including: - Group-relative policy optimization and its successors (GRPO, DAPO, GSPO, CISPO) with multi-reward objectives - Reinforcement learning with verifiable rewards, and with noisy, sparse, or learned reward signals - Reward model training: pairwise and generative reward models, outcome and process supervision - Preference optimization methods (DPO, IPO, ORPO, KTO) and RLHF - Inference-time compute scaling, including best-of-N sampling and verifier-guided decoding - Practical experience diagnosing and mitigating reward hacking and reward over-optimization - Track record of implementing complex models from research papers and adapting them to new domains - Proficiency in PyTorch or JAX, with experience training large models on multi-GPU/distributed systems - Experience with reinforcement learning infrastructure at scale, including rollout generation (vLLM, SGLang) and frameworks such as verl, TRL, or OpenRLHF - Experience with autoregressive language modeling and instruction tuning - Strong software engineering skills and ability to write production-quality code ## Preferred Qualifications - Publications at top-tier conferences (NeurIPS, ICML, ICLR, CVPR, ACL, EMNLP, MICCAI) - Hands-on experience with medical imaging applications, particularly radiology report generation - Experience with agentic or multi-turn reinforcement learning, including credit assignment over tool-use trajectories - Experience with grounded generation tasks (visual grounding, referring expression comprehension) - Knowledge of evaluation methodologies for long-form generation, including factuality assessment and hallucination detection - Experience mitigating catastrophic forgetting of supervised capabilities during reinforcement learning - Familiarity with clinical NLP and medical knowledge representation - Experience with model interpretability, explainability, and uncertainty quantification in safety-critical applications ## About Epsilon Labs, Inc. ## Company Overview - **One-liner**: Epsilon Health is a tech-enabled teleradiology practice that combines board-certified radiologists with proprietary AI to deliver accurate, rapid diagnostic reports. - **Entity Type**: Private (venture-backed startup, pre-Series A) - **Headquarters**: San Francisco, CA, USA - **Founded**: 2024 - **Founders**: Not publicly disclosed (key executives include Dr. Roi Bittane – Chief Medical Officer, and Rustin Rassoli – Founder/Executive; team also includes leaders from Google DeepMind, Meta, and Twitch) ## Core Business - **Primary industry**: Healthcare – Teleradiology / Medical Imaging / Health Tech - **Target customers**: B2B – Imaging centers, hospitals, health systems, and payers in the United States - **Mission**: “To make sure no diagnosis is missed, delayed, or wrong.” ## Products & Services - **Epsilon Platform (Proprietary AI + Workflow)**: A SaaS-enabled interpretation service that integrates AI models for triage, flagging critical findings, and automating busywork. Radiologists use the platform to increase speed and accuracy while reducing burnout. - **Teleradiology Services**: 24/7 remote radiology interpretation with a turnaround time of 24 hours or less; critical findings flagged immediately. Combines human expertise with AI assistance. ## Market Standing - **Valuation/Market Cap**: Not publicly available (early-stage startup) - **Key Metric**: Total funding not disclosed; backed by venture studio Atomic (based on team background). Headcount ~6 employees (as of mid-2025). - **Notable Investors/Partners**: Implied backing from Atomic; no formal funding announcement found. - **Growth Signals**: Founded in 2024, actively hiring for Research Scientist, Research Engineer, and Senior Backend Engineer roles. Addresses a massive market gap: 700 million scans/year in the US, with a radiologist shortage projected to reach 15,000 by 2030. ## Competitive Advantages - **Integrated AI + Human Workflow**: Unlike point solutions that fail in production, Epsilon builds AI directly into its practice, allowing continuous iteration across hospitals and imaging centers. - **Speed & Accuracy**: 75% of critical findings flagged immediately; reports delivered in 24 hours or less. - **Radiologist-Centric Design**: Removes administrative busywork to let radiologists focus on interpretation, reducing burnout. - **Team Depth**: Combines clinical leadership (former CMO of Envision Radiology) with top-tier ML engineering (Google DeepMind, Meta, Twitch alumni). ## Strategic Focus - **Scale the practice** to meet growing imaging demand by onboarding more radiologists and imaging center partners. - **Deepen AI capabilities** for automated triage, detection, and workflow optimization. - **Expand partnerships** with health systems and payers to improve patient outcomes and reduce costs. ## Why Work Here - **Culture**: “Radiology rebuilt from the ground up” – a mission-driven environment focused on solving a critical healthcare crisis. Emphasis on collaboration between radiologists, engineers, and technologists. - **Work Policy**: Hybrid – in-office presence in San Francisco (SOMA area) with remote flexibility for certain roles (Built In lists both “In-Office” and “Remote Workspace” options). - **Engineering Culture**: Small, high-impact team with autonomy; roles span ML infrastructure, computer vision, and backend systems. Opportunity to shape the product from an early stage. - **Perks**: Not detailed, but typical for early-stage health tech (likely equity, health benefits, and the chance to work on life-saving technology). ## Sources 1. [epsilon.health](https://www.epsilon.health/) – Company homepage and product description 2. [epsilon.health/about](https://www.epsilon.health/about) – Mission, background, and statistics 3. [epsilon.health/team](https://www.epsilon.health/team) – Leadership and engineering team 4. [builtin.com/company/epsilon-health](https://builtin.com/company/epsilon-health) – Office location, headcount, and work policy 5. [jobs.ashbyhq.com/epsilon-health](https://jobs.ashbyhq.com/epsilon-health) – Active job openings ## Other roles at Epsilon Labs, Inc. - [Software Engineer - Product](https://feeny.ai/job/software-engineer-product-epsilon-labs-inc-san-francisco-y7ttr13zgn08) — San Francisco, CA - [Software Engineer - ML Infrastructure](https://feeny.ai/job/software-engineer-ml-infrastructure-epsilon-labs-inc-san-francisco-fbhbgrb0sk47) — San Francisco, CA - [Software Engineer - Core Systems](https://feeny.ai/job/software-engineer-core-systems-epsilon-labs-inc-san-francisco-55xn9qbzrags) — San Francisco, CA - [Research Scientist - VLM Pretraining](https://feeny.ai/job/research-scientist-vlm-pretraining-epsilon-labs-inc-san-francisco-e9sv1335xj0a) — San Francisco, CA - [Research Scientist - Vision Foundation Models](https://feeny.ai/job/research-scientist-vision-foundation-models-epsilon-labs-inc-san-francisco-mzemvmd5nd0b) — San Francisco, CA - [Research Engineer - Data Quality & Evals](https://feeny.ai/job/research-engineer-data-quality-evals-epsilon-labs-inc-san-francisco-bvtb488px6jv) — San Francisco, CA - [Reading Radiologists](https://feeny.ai/job/reading-radiologists-epsilon-labs-inc-remote-jr85tcf6zpp3)