--- title: 'Site Reliability Engineer at Boson AI' canonical: 'https://feeny.ai/job/site-reliability-engineer-boson-ai-toronto-69xyfzggnjaj' type: 'job' last_seen: '2026-09-11' --- # Site Reliability Engineer at Boson AI - **Company:** Boson AI - **Location:** Toronto, Canada - **Compensation:** $125k–$250k - **Employment:** full-time - **Work type:** onsite - **Posted:** 2026-07-14 - **Last confirmed live:** 2026-09-11 - **Apply:** https://jobs.lever.co/bosonai/15b2b0cc-7c08-4626-a7a1-17d1334d638e ## Job description ## About The Role Boson AI builds production-grade AI systems that make communication with AI more natural, capable, and useful. We are looking for a Site Reliability Engineer to help build and operate the infrastructure behind that work. Based in Toronto or remote, you will work across the systems that enable large-scale AI training and serving: high-performance networks, GPU clusters, storage, scheduling, and the operational tooling that keeps them reliable. This is a hands-on role for someone who enjoys taking complex infrastructure from “it works” to dependable, observable, and scalable. You do not need to be an expert in every layer of the stack. We are looking for deep strength in at least one area—networking, cluster scheduling, storage, GPU systems, or AI infrastructure— and the curiosity and judgment to collaborate across the rest. ## Responsibilities - Design, operate, and improve reliable infrastructure for AI training and inference workloads - Own and automate operational workflows across one or more core areas: networking, compute allocation, storage, GPU/server configuration, or AI platforms - Build monitoring, alerting, runbooks, and incident-response practices that make systems easier to operate - Diagnose performance, capacity, and reliability issues across hardware, operating systems, networks, schedulers, and distributed workloads - Partner closely with ML, research, and platform teams to translate workload needs into practical infrastructure improvements - Improve provisioning, configuration management, testing, and deployment automation - Help plan cluster growth, capacity allocation, upgrades, and lifecycle management - Contribute to a thoughtful reliability culture through documentation, post-incident learning, and pragmatic engineering standards ## Minimum Qualifications - 4+ years of experience in site reliability engineering, infrastructure engineering, systems engineering, or a related production-operations role - Strong hands-on expertise in at least one of the following: - Networking, including firewalls, switching, routing, ASN/BGP configuration, or InfiniBand - Cluster and systems allocation with Kubernetes, SLURM, MAAS, or similar platforms - Distributed storage, particularly Ceph - GPU and server administration, including CUDA drivers, firmware, BIOS, and hardware troubleshooting - AI training or model-serving infrastructure - Experience operating production systems with a focus on availability, performance, security, and automation - Strong Linux administration and scripting skills - A systematic approach to troubleshooting across multiple layers of a complex system - Clear written and verbal communication skills, including the ability to work effectively with a distributed team ## Preferred Qualifications - ## Experience supporting GPU-intensive AI or HPC environments - Experience with NVIDIA GPUs, CUDA, NCCL, and high-performance interconnects - Experience with InfiniBand, RDMA, RoCE, or 100Gb+ Ethernet - Familiarity with Kubernetes, SLURM, MAAS, Terraform, Ansible, or similar infrastructure tooling - ## Experience operating or tuning Ceph clusters - Familiarity with observability tooling such as Prometheus, Grafana, and centralized logging systems - Experience with hardware provisioning, firmware management, and bare-metal automation - Experience running large-scale distributed training or high-throughput inference workloads - Familiarity with cloud and hybrid infrastructure across AWS, GCP, or Azure Boson AI is building AI systems for real-world, business-critical use. If you enjoy solving difficult infrastructure problems and want your work to directly enable the next generation of AI products, we’d love to hear from you. ## About Boson AI ## Company Overview - **One-liner**: Boson AI builds real-time, audio-native AI models and agentic systems for natural, conversational voice interactions with machines. - **Entity Type**: Private (funding not disclosed) - **Headquarters**: Santa Clara, California, United States (also Toronto, Canada office) - **Founded**: 2023 - **Founders**: Dr. Alex Smola (Co-Founder & CEO/CSO), Dr. Mu Li (Co-Founder & CEO/CTO), Yi Zhu (Co-Founder, departed May 2026), Yizhi Liu (Co-Founder) ## Core Business - **Primary industries**: Conversational AI, Voice AI, Multimodal AI, Enterprise AI infrastructure - **Target customers**: B2B – enterprises deploying voice agents, real-time customer service, and AI-powered communication workflows - **Mission**: “Make communication with machines as easy, natural and fun as talking to a human.” ## Products & Services - **Higgs Realtime**: An end-to-end, audio-native real-time speech-to-speech model for enterprise voice agents. Supports interruptions, code-switching (100+ languages), ~700ms speech-in to speech-out latency, and API compatibility with OpenAI Realtime (change three lines of code). Pricing: $0.0023/min audio in, $0.014/min audio out. [LinkedIn post](https://linkedin.com/company/boson-ai) - **Higgs Avatar v1**: Real-time conversational avatar foundation model that generates 480p video at 16 FPS from a single image and streaming audio. Designed for real-time digital presence in voice agents. Private preview announced May 2026. [LinkedIn post](https://linkedin.com/company/boson-ai) - **Agent Platform (Agent OS)**: An internal platform for building, deploying, and orchestrating AI agents (referenced in job postings as “Member of Technical Staff - Agent Platform”). ## Market Standing - **Valuation**: Not publicly available - **Key Metric**: Total funding not disclosed; headcount of 32 employees (as of mid-2026) - **Notable Investors/Partners**: Not publicly listed. Talent sources include Amazon Web Services (AWS), Google, University of Toronto, Vector Institute, indicating strong ties to top AI research and engineering communities. - **Growth Signals**: Launch of Higgs Realtime (August 2026) and Higgs Avatar v1 (May 2026); active hiring for senior engineering and ML roles; 10,000+ LinkedIn followers with +2.5% monthly growth; technical team comprises 68% of workforce. ## Competitive Advantages - **Full-stack AI ownership**: Builds own foundation models (audio, avatar, agent orchestration) rather than stitching external components, enabling deep co-design for low latency and natural interaction. - **Real-time expertise**: Achieves ~700ms speech-in to speech-out and ~125ms barge-in yield, with top benchmarks in tool-calling and interruption recovery. - **Cost efficiency**: Higgs Realtime pricing ($0.0023/min audio in) is significantly lower than many competitors, targeting enterprise-scale deployment. - **Research pedigree**: Founders have “almost a century of expertise in AI” (Alex Smola and Mu Li are well-known ML researchers; Smola co-authored seminal work on kernel methods and scalable ML). ## Strategic Focus - **Real-time conversational AI for enterprise**: Prioritizing production-ready voice agents that can handle interruptions, code-switching, and emotional alignment. - **Multimodal expansion**: Adding visual presence (avatars) to voice agents to make interactions more natural and human-compatible. - **Enterprise deployment**: Offering self-hosting options and private previews for large-scale customers. - **Building the full agentic pipeline**: From data and modeling to training, tuning, and serving – all in-house. ## Why Work Here - **Culture**: Described as “a diverse group of researchers, engineers, and industry specialists united by a passion for innovation.” Emphasis on building scalable AI that serves millions. - **Work policy**: Most roles are on-site at Santa Clara HQ or Toronto office. A Site Reliability Engineer role was listed as remote (Toronto). Likely hybrid/on-site for core engineering. - **Engineering environment**: Deep technical stack (PyTorch, TensorFlow, Kubernetes, NVIDIA, Supermicro, etc.). Opportunity to work on frontier AI models and real-time systems. - **Growth stage**: Small team (~32 people) with strong research roots; employees have previously worked at AWS, Google, Robinhood, and alumni go to OpenAI, Anthropic, xAI – indicating high-caliber talent and career mobility. - **Hiring process**: Uses AI tools to assist with resume review and analysis, but final decisions made by humans. ## Sources 1. [boson.ai/about](https://www.boson.ai/about) 2. [boson.ai/about/team](https://www.boson.ai/about/team) 3. [linkedin.com/company/boson-ai](https://linkedin.com/company/boson-ai) 4. [jobs.lever.co/bosonai](https://jobs.lever.co/bosonai) ## Other roles at Boson AI - [Datacenter Technician](https://feeny.ai/job/datacenter-technician-boson-ai-barrie-11tetdz9z3vr) — Barrie, Canada - [Software Engineer - Platform & Application](https://feeny.ai/job/software-engineer-platform-application-boson-ai-santa-clara-zct1esjbfzcn) — Santa Clara, CA - [Senior Software Engineer - Systems](https://feeny.ai/job/senior-software-engineer-systems-boson-ai-santa-clara-94h0qsa26h84) — Santa Clara, CA - [Frontend Engineer](https://feeny.ai/job/frontend-engineer-boson-ai-santa-clara-nbsv99agxh8b) — Santa Clara, CA - [Machine Learning Engineer - Enterprise](https://feeny.ai/job/machine-learning-engineer-enterprise-boson-ai-toronto-cm87w800xvcq) — Toronto, Canada - [Member of Technical Staff - Agent Platform (Agent OS)](https://feeny.ai/job/member-of-technical-staff-agent-platform-agent-os-boson-ai-santa-clara-95vzvadf5n6c) — Santa Clara, CA - [Site Reliability Engineer](https://feeny.ai/job/site-reliability-engineer-singlestore-linkedin-portugal-r7dtrfaa17rn) — Portugal - [Site Reliability Engineer](https://feeny.ai/job/site-reliability-engineer-genesis-london-cpq05vt4mpzk) — London, United Kingdom - [Site Reliability Engineer](https://feeny.ai/job/site-reliability-engineer-instrumental-inc-palo-alto-rzgc6dd10n8w) — Palo Alto, CA - [Site Reliability Engineer](https://feeny.ai/job/site-reliability-engineer-delinea-home-office-0kk2x5z01xfd) — Home Office, United Kingdom