--- title: 'Senior Site Reliability Engineer, Database Infrastructure at Zello' canonical: 'https://feeny.ai/job/senior-site-reliability-engineer-database-infrastructure-zello-austin-e77sr6pg4xta' type: 'job' last_seen: '2026-09-10' --- # Senior Site Reliability Engineer, Database Infrastructure at Zello - **Company:** Zello - **Location:** Austin, TX - **Employment:** full-time - **Work type:** hybrid - **Posted:** 2026-05-13 - **Last confirmed live:** 2026-09-10 - **Apply:** https://jobs.ashbyhq.com/Zello/5b5e82e4-93c9-4dd2-a339-a81b0d528a44/application **Skills:** MySQL, MongoDB, Google Cloud, Linux, Docker, Kubernetes, Prometheus, Loki, Tempo, Python, Go, Bash, OpenTelemetry, AI tooling, ScyllaDB, Cassandra, Elasticsearch, AWS, Azure > The Senior Site Reliability Engineer owns the reliability of the data tier, including MySQL and MongoDB, across Google Cloud. Responsibilities include designing highly available clusters, optimizing query performance, extending observability with Prometheus and Loki, and leading on-call rotations and incident response. ## Job description IMPORTANT: Please be aware, scammers may try to impersonate Zello by reaching out regarding job opportunities. We will never ask you for bank account information, checks, or other sensitive information as part of our hiring process. All correspondence will come from the zello.com email domain. If you’re unsure, please email recruiting@zello.com with questions. ## About Zello Zello is a voice-first communication platform, powered by our industry-leading push-to-talk technology, to improve collaboration and productivity for desk-less workers. With over 175+ million users, we’re the #1 rated push-to-talk app in the world, delivering 9 billion (yes, with a B) messages a month. At Zello, our company values are at the heart of what we do everyday. We’re proud to serve the frontline, we’re privileged to connect people in times of crisis across the globe, and we’re honored to support first responders. And this is where you come in. We're seeking a Senior Site Reliability Engineer who can own our data tier at high availability while also pulling weight across the broader platform. As Zello scales, the line between "database problem" and "platform problem" keeps blurring. We want someone who can sit on either side of it. This hire owns our data tier reliability (MySQL, MongoDB, ScyllaDB, Elasticsearch, Redis) and contributes to monitoring, on-call, and our ongoing cloud modernization efforts. ## About Zello Zello is the leading push-to-talk communication platform, enabling instant voice communication for frontline workers across hospitality, logistics, transportation, construction, and public safety. When a hotel manager radios housekeeping or a trucker calls dispatch, they're on Zello — and they need it to work every time. The Platform team builds and operates the infrastructure that makes that possible. Databases sit at the center of that promise: every channel, every message, every login depends on them. ## The Role You'll join the Platform team and report to the Director of Platform Engineering. You'll own the reliability of our MySQL and MongoDB footprint across Google Cloud, work alongside application engineers on performance and schema decisions, and contribute to the broader platform, observability with Prometheus, Loki, and Tempo; on-call; incident response;. This role suits someone who likes operating real production systems, doesn't get stage fright in incidents, and writes the runbook for the next person who hits the same problem. We're investing in AI to compress incident response, build agents and tooling that speed up root-cause analysis, and lift developer productivity across engineering. We want someone curious about what that looks like for an SRE and excited to help shape it. After a Successful First Year, You Will Have: - Operated Zello's MySQL and MongoDB clusters to documented availability targets, with automated backups, regularly tested restores, and failover the on-call team trusts under real incident pressure. - Cut latency or capacity cost on at least one critical database workload through measurable performance work — index strategy, query tuning, schema changes, or sharding. - Extended our Observability coverage so incidents are diagnosed in minutes rather than hours, with dashboards and alerts the team actually uses. - Owned a slice of the Platform on-call rotation and led postmortems that turned recurring incidents into permanent fixes. ## What You'll Do - Design, deploy, and operate highly available MySQL and MongoDB clusters across our cloud environments; replication, sharding, backups, point-in-time recovery, upgrades, and disaster recovery. - Tune query performance, schema, and index strategy in partnership with application engineers and push fixes upstream into the application when that's the right answer. - Extend our observability stack — Prometheus, Loki, and Tempo — so the data tier is as well instrumented as the application tier, and traces actually reach the root cause. - Participate in the Platform on-call rotation, lead incident response for data-tier issues, and write postmortems that drive durable change. - Improve disaster recovery, security posture, and compliance for our database footprint — encryption, access control, audit logging, backup integrity. - Evaluate and operate ScyllaDB/Cassandra and Elasticsearch where they fit the workload, and bring an opinion on when they don't. - Write the automation, tooling, and operators that take repetitive work off the team's plate. - Use AI to compress incident response and root-cause analysis; building agents, automation, and developer-enablement tooling that scale the team's reliability work ## Who You Are - You've operated highly available MySQL and MongoDB in production at scale; replication, sharding, backups, point-in-time recovery, and failover drills you've actually run, not just designed on paper. - You diagnose database performance end-to-end; query plan, indexes, locking, OS, storage, network — and can point to specific incidents where you found and fixed root cause that others had missed. - You've shipped meaningful work on at least two of bare metal Linux, containerized workloads (Docker, Kubernetes, or similar), and a major cloud (GCP preferred; AWS or Azure equivalent is fine). - You instrument what you build. You've used Prometheus, OpenTelemetry, or comparable systems to close real incidents, and you've written the dashboard the next on-call engineer will actually open. - You write code that runs in production: Python, Go, Bash, or similar for automation, tooling, or operators. You don't hand off scripting to someone else. - You communicate clearly under pressure and after the fact. Your postmortems are blameless, specific, and lead to changes that stick — and the people you've worked with describe collaborating with you as straightforward. - You bring an opinion on managed vs. self-managed databases, and can defend the trade-off based on availability, cost, and operational burden. - 7+ years in SRE, DevOps, platform, infrastructure, or database reliability roles, with at least 3 years owning production databases. - BSc in Computer Science or equivalent practical experience. - ScyllaDB/Cassandra or Elasticsearch experience is a plus - You've used AI tooling: copilots, agents, or custom automation to expedite incident response, root-cause analysis, or developer workflows. We hire for potential, passion for our mission, and a knack for solving difficult problems over checking every qualification box. We have competitive pay, equity with significant upside, and intentionally design our benefits to encourage healthy and well-balanced employees, flexible schedules and time off. We even offer a sabbatical after every five years of service so you’re able to pursue and enjoy what matters most to you. And of course, we wouldn’t be a technology company without a ping-pong table and free snacks in our break room. Join us! Zello provides equal employment opportunities to all employees and applicants for employment and prohibits discrimination and harassment of any type without regard to race, color, religion, age, sex, national origin, disability status, genetics, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state or local laws. All Zello personnel are required to comply with defined security, privacy, and compliance requirements applicable to their role along with requirements that are applicable to all Zello personnel. ## About Zello ## Company Overview - **One-liner**: Zello provides a push-to-talk walkie-talkie app for frontline workers, augmented with AI-powered transcription, translation, and operational insights. - **Entity Type**: Private (self-funded and profitable) - **Headquarters**: Austin, Texas, USA - **Founded**: 2011 - **Founders**: Alex Gavrilov (CEO) ## Core Business - **Primary industry**: Communication technology / Frontline operations software - **Target customers**: B2B – enterprises with deskless workers in transportation, retail, construction, hospitality, healthcare, and first responder organizations - **Mission**: “We give voice to everyday heroes” – to lift up the voices of the frontline and bring them the same level of technical investment as desk workers ## Products & Services - **[Zello App](https://zello.com/)**: Mobile and desktop push-to-talk walkie-talkie app with 99.99% uptime, supporting iOS, Android, Windows, and Mac. Offers unlimited range over Wi-Fi/cellular and on-premise deployment via Zello Enterprise Server. - **[Zello AI](https://zello.com/)**: AI assistant “Ella” that answers questions using company knowledge; live voice translation to break language barriers; conversation replay, transcription, and offline message recovery. - **[Zello Mobile SDK](https://zello.com/)**: Embed push-to-talk natively into third-party apps (iOS, Android, React Native). - **[Radio Gateways](https://zello.com/)**: Bridge LMR (land mobile radio) networks with Zello for seamless cross-platform communication. ## Market Standing - **Valuation/Market Cap**: Not disclosed (self-funded, no external funding rounds reported) - **Key Metric**: 5 million monthly active users, 10 billion messages per month (source: [zello.com](https://zello.com/)); conflicting report from Built In cites 8 billion messages and 170 million total users ([builtin.com](https://builtin.com/company/zello)) - **Notable Investors/Partners**: None publicly listed; self-funded and profitable - **Growth Signals**: Rapid hiring across engineering, AI, sales, and customer success; expanding AI capabilities (live translation, AI assistant); over 3,000 emergency organizations use Zello free for first responders ## Competitive Advantages - **99.99% uptime** – highly reliable for mission-critical frontline communication - **Only frontline AI powered by real conversations** – Zello AI listens to actual push-to-talk traffic and transforms it into operational intelligence - **Enterprise readiness** – on-premise deployment option, radio gateway integration, and IT-friendly management tools - **Large installed base** – millions of active users across industries creates network effects and deep integration into workflows ## Strategic Focus - **AI-first innovation** – building AI assistant “Ella,” live translation, and transcription to turn voice chatter into actionable insights - **Deepening frontline integration** – SDK and radio gateways to embed Zello into existing ecosystems - **Scaling enterprise sales** – hiring Director of Global Partner & Channel Strategy, Enterprise Account Executives, and Director of Customer Success ## Why Work Here - **Self-funded and profitable** – described as a “build-up” rather than a startup, offering stability with growth potential and stock options - **Hybrid workplace** – Austin employees work in-office Tuesdays, Wednesdays, and Thursdays; remote flexibility on other days - **Generous benefits** – 100% paid medical, vision, and dental; 4% 401k matching; $2,000 annual professional development allowance; open PTO; paid sabbatical after 5 years; generous parental leave - **Engineering culture** – yearly hackathon at an incredible destination; small team (80 total employees, 36 in product+tech) with a focus on impact, excellence, teamwork, transparency, and caring - **Growth opportunities** – roles in applied AI, back end, site reliability, and design; the company is actively scaling and investing in new technology ## Sources 1. [zello.com/careers](https://zello.com/careers/) 2. [zello.com/about-us](https://zello.com/about-us/) 3. [zello.com](https://zello.com/) 4. [builtin.com/company/zello](https://builtin.com/company/zello) 5. [builtin.com/company/zello/faq/workplace-perception](https://builtin.com/company/zello/faq/workplace-perception) ## Other roles at Zello - [Growth Marketing Manager](https://feeny.ai/job/growth-marketing-manager-zello-austin-tb506qmv0b8e) — Austin, TX - [Senior Software Engineer, AI Systems](https://feeny.ai/job/senior-software-engineer-ai-systems-zello-austin-ba3kfnn74x2m) — Austin, TX - [Contracts Manager](https://feeny.ai/job/contracts-manager-zello-austin-5ckhyngt72fn) — Austin, TX - [Escalation Engineer](https://feeny.ai/job/escalation-engineer-zello-austin-vkt9wcewq63s) — Austin, TX - [Analytics Intern](https://feeny.ai/job/analytics-intern-zello-austin-1548zj88wzwr) — Austin, TX - [Product Marketing Manager](https://feeny.ai/job/product-marketing-manager-zello-austin-d17f91xtkyrr) — Austin, TX - [Product Data Analyst](https://feeny.ai/job/product-data-analyst-zello-austin-1h35msq44fq3) — Austin, TX - [Enterprise Account Executive](https://feeny.ai/job/enterprise-account-executive-zello-austin-3g2dcjr60m9y) — Austin, TX