--- title: 'Machine Learning Scientist at Rime Labs' canonical: 'https://feeny.ai/job/machine-learning-scientist-rime-labs-united-states-r8cqrr7exaes' type: 'job' last_seen: '2026-09-06' --- # Machine Learning Scientist at Rime Labs - **Company:** Rime Labs - **Location:** United States - **Employment:** full-time - **Work type:** remote - **Posted:** 2026-05-15 - **Last confirmed live:** 2026-09-06 - **Apply:** https://jobs.ashbyhq.com/rime/b76243da-3ce0-463e-9f29-2b6d1b78666e ## Job description ## MACHINE LEARNING SCIENTIST Rime builds voice AI for enterprises running customer experiences at scale. Our text-to-speech models are purpose-built for high-volume conversational deployments, engineered for the pronunciation accuracy, latency, and deployment flexibility that production environments actually demand. We started from a different premise than the rest of the field: voice AI isn't bottlenecked by model architecture. It's bottlenecked by data. So before we trained a single model, we built our own corpus: full-duplex, studio-quality conversational speech, recorded and annotated by PhD linguists. That's our moat. It's also why enterprises pick Rime when pilots need to convert into production. We're backed by top-tier investors including Unusual Ventures, and we've built a team at the intersection of product, research, and craft. Building voice models is an art. We intend to master it. ## ROLE OVERVIEW We're hiring a Machine Learning Scientist to push the frontier of speech synthesis and speech understanding at Rime. ## What You'll Own - Design, train, and evaluate speech synthesis models, autoregressive and non-autoregressive. - Drive research on full-duplex and half-duplex multi-modal architectures, including unified S2S systems. - Choose and iterate on speech representations: neural codecs, semantic tokens, mel features, continuous latents. - Build rigorous evaluation, objective and perceptual. Hold the bar on quality and prosodic control. - Collaborate with our linguists on TTS frontend behavior so modeling and frontend choices reinforce each other. ## WHAT WE'RE LOOKING FOR - Deep familiarity with the speech synthesis literature, contemporary and historical — Tacotron, FastSpeech, VITS, VALL-E, the codec-LM lineage. Opinions on what worked and why. - Hands-on training with neural codecs (EnCodec, DAC, Mimi, etc.) and multiple representation choices. - Experience with full- or half-duplex multi-modal modeling (Moshi, LLaMA-Omni, streaming S2S). - Strong attention to detail on data quality. You notice when an annotation pipeline is silently degrading or when an eval set has leakage. - Willing to roll up your sleeves on unglamorous data and training work — paired with the agency to build pipelines so the team isn't stuck doing it by hand. - Working knowledge of TTS frontend (G2P, normalization, prosody) and experience working with linguists. - Strong PyTorch fundamentals. Comfortable with training loops, distributed training, model internals. - PhD or equivalent research experience in speech, audio, ML, or computational linguistics or a track record that makes the credential irrelevant. ## Nice to have - Multilingual TTS experience. - Background in prosody or paralinguistics. - Published work in speech, audio, or core ML venues. - Experience taking research models to production: quantization, distillation, streaming inference. ## WHY JOIN RIME - Category-defining voice AI infrastructure, not incremental research deltas. - Direct collaboration with founders, including a CEO with a Stanford computational linguistics PhD. - Real impact on company trajectory. - Meaningful equity upside. - High ownership, high standards, low bureaucracy. ## What We Offer - Competitive base + meaningful early-stage equity - Remote-friendly - Visa sponsorship available - Access to a proprietary, full-duplex, studio-quality conversational speech corpus - Compute and tooling to do the work - Direct influence on the future of voice AI At Rime, we... - Are outliers - Cut through the hype to focus on the craft - Move fast with agency and freedom - Maintain a growth mindset, finding joy in the struggle - Do the right things, knowing that it'll lead to making money If that sounds like you too, you'll be a great fit for Rime! ## About Rime Labs ## Company Overview - **One-liner**: Rime Labs develops enterprise-grade, ultra-realistic text-to-speech AI voice models designed for production contact center, healthcare, and financial applications. - **Entity Type**: Private (venture-backed, Seed stage) - **Headquarters**: San Francisco, California, United States - **Founded**: 2022 - **Founders**: Lily Clifford (CEO), Brooke Larson (PhD Linguist, ex-Amazon Alexa), Ares Geovanos (COO) ## Core Business - **Primary Industry**: Artificial Intelligence / Voice AI (Text-to-Speech), Enterprise Software - **Target Customers**: B2B, Enterprise (contact centers, healthcare providers, financial institutions, telecom) - **Mission Statement**: "Make voice AI sound and feel human." ## Products & Services - **[Coda](https://rime.ai/)**: Flagship TTS model blending human-quality voice with enterprise-grade speed and concurrency. Optimized for production deployments. - **[Arcana](https://rime.ai/)**: Most expressive and human-like TTS model with 300+ voices, bilingual/multilingual support, and instant code-switching (English, Spanish, Spanglish). Designed for warm, emotionally resonant conversations. - **[Mist](https://rime.ai/)**: High-speed enterprise TTS with sub-200 ms cloud latency and <100 ms on-prem. Supports high-volume streaming and extensive customization. - **[SpeechQA](https://rime.ai/)**: Tool to flag low-confidence words before deployment, enabling proactive pronunciation fixes. ## Market Standing - **Valuation**: Conflicting reports – PitchBook indicates a post-money valuation of ~$9.6M as of May 2025; LinkedIn reports total funding of $8.6M. Not publicly disclosed. - **Key Metric (Private)**: Total Funding – $8.6M across two seed rounds (June 2023: $3.1M; June 2025: $5.5M). PitchBook lists a $6.5M seed round in May 2025 with 22 investors. - **Notable Investors**: Unusual Ventures (lead), Cadenza Capital, Founders You Should Know, and angel investors. - **Growth Signals**: 33 employees (+245.5% YoY); operates in 9 countries; powers tens of millions of conversations monthly across industries including healthcare and food service. ## Competitive Advantages - **Proprietary Dataset**: One of the largest collections of expressive, multilingual conversational speech, captured in-studio and across diverse U.S. locations. - **Deep Linguistic Expertise**: Founders include a Stanford NLP PhD dropout and a PhD linguist from Amazon Alexa; models trained on real-world speech patterns (interruptions, laughter, disfluencies). - **Enterprise-Ready Compliance**: SOC 2 Type II and HIPAA compliant; deployable on-prem, in VPC, or via cloud API. - **Pronunciation Control**: Fine-grained tools to fix pronunciations in minutes without retraining; built-in handling of brand names, drug names, and numbers. - **Low Latency at Scale**: Sub-200 ms cloud latency; <100 ms on-prem; designed for production loads, not just demos. ## Strategic Focus - Deepening enterprise adoption in regulated industries (healthcare, finance, telecom). - Expanding multilingual and code-switching capabilities (English/Spanish/Spanglish). - Scaling the engineering and go-to-market teams to meet growing demand (VP of Engineering, Founding Account Executive roles open). ## Why Work Here - **Culture**: Small, passionate, lean team based in San Francisco. Emphasis on pushing boundaries in voice AI and creating technology that connects with real people. - **Work Environment**: Likely hybrid/office in San Francisco (HQ); on-site recording studio. Notable perks: direct impact on product direction, access to cutting-edge ML and linguistics research. - **Open Roles** (7 as of mid-2026): Solutions Engineer, Solutions Engineering Manager, Developer Advocate, Founding Account Executive, VP of Engineering, Revenue Operations, Product Marketing Manager. - **Engineering Culture**: Strong focus on ML, low-latency systems, and production reliability. Team includes engineers from Meta, Stanford, and speech startups. ## Sources 1. [rime.ai - Homepage](https://rime.ai/) 2. [rime.ai - Company Page](https://rime.ai/company) 3. [LinkedIn - Rime AI](https://www.linkedin.com/company/rime-ai) 4. [PitchBook - Rime Labs Profile](https://pitchbook.com/profiles/company/528897-34) 5. [Ashby - Rime Labs Careers](https://jobs.ashbyhq.com/rime) ## Other roles at Rime Labs - [Forward Deployed Engineer](https://feeny.ai/job/forward-deployed-engineer-rime-labs-san-francisco-hsbwehyshys0) — San Francisco, CA - [Founding Account Executive](https://feeny.ai/job/founding-account-executive-rime-labs-san-francisco-ecc248tbhg9z) — San Francisco, CA - [Machine Learning Scientist](https://feeny.ai/job/machine-learning-scientist-flagship-pioneering-inc-somerville-ma-98f7vrcztgjm) — Somerville MA, United States - [Machine Learning Scientist](https://feeny.ai/job/machine-learning-scientist-eli-health-montreal-b0xn6yyc72ff) — Montréal, Canada - [Machine Learning Scientist](https://feeny.ai/job/machine-learning-scientist-tacit-san-francisco-pt1w314e43ge) — San Francisco, CA - [Machine Learning Scientist](https://feeny.ai/job/machine-learning-scientist-latent-labs-london-9m8twmtyrvpj) — London, United Kingdom - [Machine Learning Scientist](https://feeny.ai/job/machine-learning-scientist-spotter-culver-city-california-wgbw208cbr43) — Culver City California, United States - [Machine Learning Scientist](https://feeny.ai/job/machine-learning-scientist-adyen-amsterdam-qttbfda3bx3z) — Amsterdam, Netherlands - [Machine Learning Scientist](https://feeny.ai/job/machine-learning-scientist-relace-san-francisco-x6cd3a3vkg65) — San Francisco, CA - [Machine Learning Scientist](https://feeny.ai/job/machine-learning-scientist-suno-boston-fj9mt2har4hv) — Boston, MA