--- title: 'Machine Learning Engineer, Inference & Serving (Speech LLM) - San Francisco at Plaud' canonical: 'https://feeny.ai/job/machine-learning-engineer-inference-serving-speech-llm-san-francisco-plaud-san-vgv6qsq1pe0x' type: 'job' last_seen: '2026-09-09' --- # Machine Learning Engineer, Inference & Serving (Speech LLM) - San Francisco at Plaud - **Company:** [Plaud](https://feeny.ai/companies/plaud) - **Location:** San Francisco, CA - **Compensation:** $180k–$270k - **Employment:** full-time - **Work type:** hybrid - **Posted:** 2026-05-08 - **Last confirmed live:** 2026-09-09 - **Apply:** https://jobs.ashbyhq.com/Plaud/01f7453c-0f8f-456a-9182-3fe539ffec41/application **Skills:** Python, PyTorch, TensorFlow, Git, NVIDIA GPU, Kubernetes, WebSockets, WebRTC, LLM Serving Frameworks, KV Cache Management, PagedAttention, vLLM, TensorRT-LLM, SGLang, NVIDIA Triton Inference Server, FP8, INT8, GPTQ, Autoscaling Infrastructure, Speculative Decoding > Build and deploy high-throughput, ultra-low-latency inference engines for large language models and speech LLMs. Collaborate with core ML and backend teams to optimize real-time streaming environments, manage GPU architectures, and ensure low latency for conversational AI systems. ## Job description ABOUT PLAUD INC. Plaud is building the real-world AI interface for professionals to amplify intelligence, elevate productivity and performance, loved by over 2,500,000 users worldwide since 2023. With a mission to amplify human intelligence, Plaud captures, structures, and compounds the intelligence generated in conversations — so humans can think better, decide faster, and execute with clarity. Plaud Inc. is a Delaware-incorporated, San Francisco-based company pushing the boundary of human–AI intelligence through a hardware–software combination. With full ISO 27001, ISO 27701, SOC 2, GDPR, EN18031, and HIPAA compliances, Plaud is committed to the highest standards of data security and privacy protection. To learn more about Plaud, please visit http://www.plaud.ai/https://www.Plaud.ai and follow along on Instagram https://www.instagram.com/plaud_official/, https://twitter.com/PLAUDAI%22HYPERLINK%20%22https:/twitter.com/PLAUDAI%22%20%5C%22HYPERLINK%20%22https:/twitter.com/PLAUDAI%22%20%5CX https://twitter.com/PLAUDAI, Facebook https://www.facebook.com/plaudai, LinkedIn https://www.linkedin.com/company/plaudai/?viewAsMember=true, and YouTube https://www.youtube.com/@PLAUDAI ## Why You Should Join Us Plaud is building the next generation intelligence infrastructure and interfaces to capture, extract, and utilize intelligence from what people say, hear, see, and think. - Plaud is a bootstrapped, skyrocketing, profitable company with a $300M revenue run rate achieved in just three years. - Define the next-gen paradigm for human-AI interaction. - Gain exposure to cutting-edge AI for Pro tools and play a direct role in our global expansion. - Work with passionate teammates who value innovation, collaboration, and customer success. - Grow your career in a culture that champions continuous learning and fast career development. - Market-competitive compensation, global exposure, and a vibrant, creativity-fueled work atmosphere. YOU MAY BE A GOOD FIT IF YOU: - Have hands-on experience building and deploying high-throughput, ultra-low-latency inference engines for large language models or foundational speech models. - Understand the intricate tradeoffs between latency, throughput, and Time-To-First-Token (or Time-To-First-Audio) in real-time streaming environments. - Have practical experience with continuous batching, KV cache management (e.g., PagedAttention), and stateful connections necessary for real-time conversational AI. - Possess a deep understanding of GPU architectures (NVIDIA Ampere/Hopper) and the memory hierarchy, allowing you to identify and eliminate hardware bottlenecks. - Communicate clearly and collaborate effectively, as you will sit at the critical intersection between the core ML training team and the backend infrastructure team. - Thrive in fast-moving environments and genuinely enjoy the systems-engineering challenge of squeezing every last drop of performance out of a cluster of GPUs. - Are obsessed with building AI systems that natively understand and generate speech, ultimately creating a hardware-software AI companion that amplifies human productivity. STRONG CANDIDATES MAY ALSO HAVE EXPERIENCE WITH: - Frontier Serving Frameworks: Deep, under-the-hood familiarity with modern LLM serving frameworks like vLLM, TensorRT-LLM, SGLang, or NVIDIA Triton Inference Server (bonus points for active open-source contributions to these repositories). - Real-Time Audio Streaming: Experience handling continuous audio streams over WebSockets or WebRTC, deploying neural audio codecs, and managing chunked audio generation to minimize conversational latency. - Advanced Inference Techniques: Implementing cutting-edge generation algorithms such as speculative decoding, lookahead decoding, or chunked prefill. - Model Compression & Quantization: Hands-on experience with post-training quantization (PTQ), deploying models in FP8, INT8, AWQ, or GPTQ, without degrading audio naturalness or ASR accuracy. - Large-Scale Distributed Systems: Deploying multi-GPU (Tensor Parallelism) and multi-node inference pipelines, and managing autoscaling infrastructure using Kubernetes. ## WHAT WE OFFER - Founding Team Initiative: Opportunity to be an early, foundational member of our core SpeechLLM lab, with meaningful ownership and impact on a fast-growing startup. - Competitive Compensation: $195K - $365K base salary + performance bonus + Equity. - Comprehensive Benefits: Top-tier healthcare for employees and dependents, including dental and vision, and a generous employer subsidy. - Retirement Planning: 401(k) plan for full-time employees with company matching. - Paid Time Off: Unlimited PTO, plus 13 paid holidays. - New Parent Leave: 12 weeks of paid time off to spend time with your new family, regardless of gender. - Hybrid Office: Minimum of 3x in-office per week to foster highly collaborative, fast-paced research. - Gear & Perks: Choice of top-of-the-line laptops/workstations, annual offsites, and a fully stocked office. Plaud is and will continue to be an equal opportunity employer. We do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristics. ## About Plaud ## Company Overview - **One-liner**: Plaud builds an AI-powered hardware and software note-taking platform that automatically records, transcribes, and summarizes conversations to help professionals stay present and productive. - **Entity Type**: Private (funding stage not publicly disclosed) - **Headquarters**: San Francisco, California, USA - **Founded**: 2023 - **Founders**: Not publicly listed on the website or in standard search results ## Core Business - Primary industry/industries: AI-powered productivity software and consumer hardware - Target customers: B2C (professionals, knowledge workers, medical practitioners) and B2B (teams and enterprises) - Mission or purpose statement: "Amplify human intelligence" – building the next-generation intelligence infrastructure to capture, extract, and utilize what you say, hear, see, and think. ## Products & Services - **Plaud Note (Hardware)**: An AI note-taking device that records conversations and uses AI to generate transcripts, summaries, action items, and key points. Award-winning hardware design. - **Plaud App (Software)**: Companion app for transcription (112 languages), AI-powered summaries ("Plaud Intelligence"), search ("Ask Plaud"), and export/sharing/integration. Available on mobile and web. - **Plaud Cloud**: Encrypted private cloud sync for cross-device access, with user-controlled data storage (on-device or cloud). - **Subscription Plans**: Free tier (1,200 mins/mo transcription), Pro Plan ($99.99/year), Unlimited Plan ($239.99/year – unlimited transcription). ## Market Standing - **Valuation/Market Cap**: Not publicly disclosed - **Key Metric**: Over 2,000,000 users worldwide since 2023 (as of mid-2026); used in over 170 countries - **Notable Investors/Partners**: Not publicly listed. Featured in major media outlets (unspecified on site). - **Growth Signals**: Rapid user growth to 2M+ in ~3 years; expansion from a single product to a suite with multiple subscription tiers; global office presence (San Francisco, Singapore, Tokyo, Shenzhen, Seattle, Beijing); compliance with ISO27001, ISO27701, GDPR, SOC 2, HIPAA, and EN 18031 – signaling enterprise readiness and expansion. ## Competitive Advantages - **Hardware-Software Integration**: A dedicated hardware device for note-taking (unlike pure software competitors like Otter.ai or Fireflies.ai), enabling frictionless, ambient capture. - **Privacy-First Architecture**: Data is never sold or used to train AI models; encryption in transit and at rest; user controls what syncs; multiple global security certifications (SOC 2, HIPAA, GDPR, ISO). - **Multilingual & Multimodal**: Transcribes in 112 languages and supports multimodal input (voice, potentially more). - **Strong Mission & Brand**: Positioned as "World's No.1 AI Note-taking Brand" with a clear philosophy around "Conversation as Intelligence." ## Strategic Focus - **Scaling the user base**: Targeting 50 million professionals by 2030. - **Enterprise & Healthcare Expansion**: HIPAA compliance and healthcare-specific messaging ("complete charts instantly") indicate a push into medical and enterprise verticals. - **Global Office Build-out**: Establishing engineering, product, and go-to-market teams across US, Asia, and Europe. - **Product Depth**: Moving beyond transcription to become an "intelligence infrastructure" that captures, extracts, and utilizes all forms of human communication. ## Why Work Here - **Mission-Driven**: The company's philosophy ("Amplify Human Intelligence") is central to its culture, and employees are likely working on a product that directly impacts how millions of professionals work. - **Growth Stage**: Rapidly scaling from 2M to a target of 50M users, offering significant career growth opportunities. - **Global & Remote-Friendly**: Offices in San Francisco, Singapore, Tokyo, Shenzhen, Seattle, and Beijing – suggests a distributed, cross-cultural work environment. - **Cutting-Edge Tech**: Working on the frontier of AI (LLMs, speech recognition, hardware-software integration) with a "Pursue SOTA" (State of the Art) value. - **Values-Driven Culture**: Listed values include "Human First," "Act from first principles," "Dare to change," "Use AI to its fullest," and "Let others thrive." - **Perks**: One-year warranty on hardware, lifetime customer support, and a focus on data security (ISO/SOC/HIPAA compliance) indicates a mature engineering and operations culture. ## Sources 1. [Plaud.ai - Homepage](https://www.plaud.ai/) 2. [Plaud.ai - About Us](https://www.plaud.ai/pages/about-us) 3. [Plaud.ai - Our Story](https://www.plaud.ai/pages/our-story) 4. [Plaud - Careers Page](https://jobs.ashbyhq.com/Plaud) ## Other roles at Plaud - [Influencer Marketing Manager - Korea](https://feeny.ai/job/influencer-marketing-manager-korea-plaud-singapore-vwkcsgenbz6z) — Singapore - [Full Stack Engineer (Contractor) - Palo Alto](https://feeny.ai/job/full-stack-engineer-contractor-palo-alto-plaud-san-francisco-0srbbjq9zfsc) — San Francisco, CA - [Full-Stack Engineer - Palo Alto](https://feeny.ai/job/full-stack-engineer-palo-alto-plaud-san-francisco-pg93b9rtbnpd) — San Francisco, CA - [Legal Counsel, Data Security & Compliance](https://feeny.ai/job/legal-counsel-data-security-compliance-plaud-san-francisco-b1me6n84zpk6) — San Francisco, CA - [Senior Intellectual Property Specialist - San Francisco](https://feeny.ai/job/senior-intellectual-property-specialist-san-francisco-plaud-san-francisco-na99bkdagmpm) — San Francisco, CA - [Senior Backend Engineer - Palo Alto](https://feeny.ai/job/senior-backend-engineer-palo-alto-plaud-san-francisco-fyrbr3brtc7j) — San Francisco, CA - [Senior Backend Engineer - SG](https://feeny.ai/job/senior-backend-engineer-sg-plaud-singapore-764n3bt604m1) — Singapore - [Staff Product Manager - San Francisco](https://feeny.ai/job/staff-product-manager-san-francisco-plaud-san-francisco-f7a0pytx5wws) — San Francisco, CA - [Senior Full Stack Engineer - Singapore](https://feeny.ai/job/senior-full-stack-engineer-singapore-plaud-singapore-caph7s9eqhq8) — Singapore - [Graphic Designer - London](https://feeny.ai/job/graphic-designer-london-plaud-london-e89tzsegrqj4) — London, United Kingdom