--- title: 'Software Engineer - ML Infrastructure at Epsilon Labs, Inc.' canonical: 'https://feeny.ai/job/software-engineer-ml-infrastructure-epsilon-labs-inc-san-francisco-fbhbgrb0sk47' type: 'job' last_seen: '2026-09-09' --- # Software Engineer - ML Infrastructure at Epsilon Labs, Inc. - **Company:** Epsilon Labs, Inc. - **Location:** San Francisco, CA - **Employment:** full-time - **Work type:** onsite - **Posted:** 2026-08-31 - **Last confirmed live:** 2026-09-09 - **Apply:** https://jobs.ashbyhq.com/epsilon-health/b5ef89d4-1a4c-4302-8939-b6dbf3bb8c70 ## Job description ## ABOUT US We're tackling one of healthcare's most critical challenges in medical imaging and diagnostics. Our company operates at the intersection of cutting-edge AI and clinical practice, building technology that directly impacts patient outcomes. We've assembled one of the industry's most comprehensive and diverse medical imaging datasets and have a proven product-market fit with a substantial customer pipeline already in place. ## ROLE OVERVIEW We’re looking for an ML infrastructure engineer to design and build the core systems that enable scalable, efficient training of large models for deployment and research. Your goal is to make experimentation and training at Epsilon Health fast and reliable to ensure our research teams can focus on science rather than system bottlenecks. Sitting in the Engineering team and working closely with research, you'll own the distributed training and reinforcement learning infrastructure our foundation-model and post-training work runs on, and the inference and evaluation systems that carry models from experimentation into production. ## KEY RESPONSIBILITIES - Partner directly with researchers to deeply understand their workflows, then anticipate and design for how those needs will change - Build a distributed training infrastructure for foundation models on large-scale medical imaging, including the long-context parallelism and checkpointing that volumetric CT/MR training demands. - Build high-throughput data loading and preprocessing that keeps GPUs saturated on large volumetric and multimodal datasets. - Partner with researchers to prototype new ideas and translate them into production-ready code, owning end-to-end delivery from experimentation through deployment and monitoring. - Contribute to production serving and deployment pipelines (model rollout, canary deployments, and monitoring) alongside the backend team. - Build the reinforcement learning training stack (high-throughput rollout generation, reward-model serving, and experience collection), enabling the research team to run online, multi-reward RL at scale. ## QUALIFICATIONS - 6+ years of experience designing, building, and operating large-scale distributed systems or infrastructure in production - Have 2+ years of experience building ML infrastructure or systems in production - Strong Python skills and expertise in PyTorch or JAX - Experience and familiarity with the compute, tooling, and workflow needs of large-scale machine learning research - Experience building infrastructure or platforms specifically for research or machine learning workflows - Deep experience building and operating Kubernetes and cloud infrastructure at scale - Experience with distributed training at scale (FSDP, DeepSpeed, or Megatron-style parallelism) and the systems concerns of keeping large GPU jobs efficient - Prior experience as a technical lead or mentor for other engineers ## PREFERRED QUALIFICATIONS - Experience operating in a startup or startup-like environment, i.e. a small, fast-moving team with high autonomy - Experience building reinforcement learning training infrastructure: rollout generation, reward-model serving, or online/off-policy learning systems - Experience with high-performance inference and serving (vLLM, SGLang, TensorRT, or Triton) for both training-time rollouts and production - Experience optimizing inference and serving for large models: batching, KV/prompt caching, quantization, and low-latency, high-throughput sampling. - Experience optimizing training performance: parallelism, distributed communication, mixed/low precision, and utilization. - Experience building internal training or experimentation platforms used by research teams, supporting A/B testing and experimentation workflows - Familiarity with vision-language models (VLMs) or multimodal architectures The anticipated annual base salary for this position is up to $250,000. This range does not include any other compensation components or other benefits for which an individual may be eligible. The actual base salary offered depends on a variety of factors, which may include as applicable, the qualifications of the individual applicant for the position, years of relevant experience, specific and unique skills, level of education attained, certifications or other professional licenses held, and the location in which the applicant lives and/or from which they will be performing the job. ## About Epsilon Labs, Inc. ## Company Overview - **One-liner**: Epsilon Health is a tech-enabled teleradiology practice that combines board-certified radiologists with proprietary AI to deliver accurate, rapid diagnostic reports. - **Entity Type**: Private (venture-backed startup, pre-Series A) - **Headquarters**: San Francisco, CA, USA - **Founded**: 2024 - **Founders**: Not publicly disclosed (key executives include Dr. Roi Bittane – Chief Medical Officer, and Rustin Rassoli – Founder/Executive; team also includes leaders from Google DeepMind, Meta, and Twitch) ## Core Business - **Primary industry**: Healthcare – Teleradiology / Medical Imaging / Health Tech - **Target customers**: B2B – Imaging centers, hospitals, health systems, and payers in the United States - **Mission**: “To make sure no diagnosis is missed, delayed, or wrong.” ## Products & Services - **Epsilon Platform (Proprietary AI + Workflow)**: A SaaS-enabled interpretation service that integrates AI models for triage, flagging critical findings, and automating busywork. Radiologists use the platform to increase speed and accuracy while reducing burnout. - **Teleradiology Services**: 24/7 remote radiology interpretation with a turnaround time of 24 hours or less; critical findings flagged immediately. Combines human expertise with AI assistance. ## Market Standing - **Valuation/Market Cap**: Not publicly available (early-stage startup) - **Key Metric**: Total funding not disclosed; backed by venture studio Atomic (based on team background). Headcount ~6 employees (as of mid-2025). - **Notable Investors/Partners**: Implied backing from Atomic; no formal funding announcement found. - **Growth Signals**: Founded in 2024, actively hiring for Research Scientist, Research Engineer, and Senior Backend Engineer roles. Addresses a massive market gap: 700 million scans/year in the US, with a radiologist shortage projected to reach 15,000 by 2030. ## Competitive Advantages - **Integrated AI + Human Workflow**: Unlike point solutions that fail in production, Epsilon builds AI directly into its practice, allowing continuous iteration across hospitals and imaging centers. - **Speed & Accuracy**: 75% of critical findings flagged immediately; reports delivered in 24 hours or less. - **Radiologist-Centric Design**: Removes administrative busywork to let radiologists focus on interpretation, reducing burnout. - **Team Depth**: Combines clinical leadership (former CMO of Envision Radiology) with top-tier ML engineering (Google DeepMind, Meta, Twitch alumni). ## Strategic Focus - **Scale the practice** to meet growing imaging demand by onboarding more radiologists and imaging center partners. - **Deepen AI capabilities** for automated triage, detection, and workflow optimization. - **Expand partnerships** with health systems and payers to improve patient outcomes and reduce costs. ## Why Work Here - **Culture**: “Radiology rebuilt from the ground up” – a mission-driven environment focused on solving a critical healthcare crisis. Emphasis on collaboration between radiologists, engineers, and technologists. - **Work Policy**: Hybrid – in-office presence in San Francisco (SOMA area) with remote flexibility for certain roles (Built In lists both “In-Office” and “Remote Workspace” options). - **Engineering Culture**: Small, high-impact team with autonomy; roles span ML infrastructure, computer vision, and backend systems. Opportunity to shape the product from an early stage. - **Perks**: Not detailed, but typical for early-stage health tech (likely equity, health benefits, and the chance to work on life-saving technology). ## Sources 1. [epsilon.health](https://www.epsilon.health/) – Company homepage and product description 2. [epsilon.health/about](https://www.epsilon.health/about) – Mission, background, and statistics 3. [epsilon.health/team](https://www.epsilon.health/team) – Leadership and engineering team 4. [builtin.com/company/epsilon-health](https://builtin.com/company/epsilon-health) – Office location, headcount, and work policy 5. [jobs.ashbyhq.com/epsilon-health](https://jobs.ashbyhq.com/epsilon-health) – Active job openings ## Other roles at Epsilon Labs, Inc. - [Research Scientist - VLM Pretraining](https://feeny.ai/job/research-scientist-vlm-pretraining-epsilon-labs-inc-san-francisco-e9sv1335xj0a) — San Francisco, CA - [Research Scientist - Vision Foundation Models](https://feeny.ai/job/research-scientist-vision-foundation-models-epsilon-labs-inc-san-francisco-mzemvmd5nd0b) — San Francisco, CA - [Research Scientist - Post-training / RL](https://feeny.ai/job/research-scientist-post-training-rl-epsilon-labs-inc-san-francisco-88p5ss3dcaga) — San Francisco, CA - [Research Engineer - Data Quality & Evals](https://feeny.ai/job/research-engineer-data-quality-evals-epsilon-labs-inc-san-francisco-bvtb488px6jv) — San Francisco, CA - [Software Engineer - Product](https://feeny.ai/job/software-engineer-product-epsilon-labs-inc-san-francisco-y7ttr13zgn08) — San Francisco, CA - [Software Engineer - Core Systems](https://feeny.ai/job/software-engineer-core-systems-epsilon-labs-inc-san-francisco-55xn9qbzrags) — San Francisco, CA - [Reading Radiologists](https://feeny.ai/job/reading-radiologists-epsilon-labs-inc-remote-jr85tcf6zpp3) - [Software Engineer - ML Infrastructure](https://feeny.ai/job/software-engineer-ml-infrastructure-anduril-industries-costa-mesa-california-wkj0bexsg204) — Costa Mesa California, United States - [Software Engineer - ML Infrastructure](https://feeny.ai/job/software-engineer-ml-infrastructure-arena-intelligence-inc-bay-area-8m987gkt2sxd) — Bay Area - [Software Engineer - ML Infrastructure](https://feeny.ai/job/software-engineer-ml-infrastructure-specter-san-francisco-tt1dzsmvjbcd) — San Francisco, CA