--- title: 'Senior Site Reliability Engineer at Rootly' canonical: 'https://feeny.ai/job/senior-site-reliability-engineer-rootly-office-8xxtjk804g8w' type: 'job' last_seen: '2026-09-10' --- # Senior Site Reliability Engineer at Rootly - **Company:** Rootly - **Location:** Office - **Employment:** full-time - **Work type:** hybrid - **Posted:** 2023-04-20 - **Last confirmed live:** 2026-09-10 - **Apply:** https://jobs.gem.com/rootly/4016398007 ## Job description ## About Rootly At [Rootly](https://rootly.com/), we are on a mission to be the go-to way companies respond when things go wrong, helping every organization be more reliable. We do this by building an industry-leading incident management platform that allows companies around the world to consistently and quickly resolve incidents. We are not simply transforming an industry, we are carving an entirely new +$B segment ourselves and need incredible talent to achieve this ambitious goal together. Customers love Rootly. Some of the fastest growing companies around the world such as NVIDIA, Figma, Canva, Tripadvisor, Squarespace and more rely on Rootly to power their critical incident management process. They obsess over our delightful enterprise-ready platform and unique partnership model. See why our customers have reviewed us [5 stars on G2](https://www.g2.com/products/rootly-manage-incidents-on-slack/reviews). Investors love Rootly. We are backed by some of the most respected funds in the world from Y Combinator to operators like the CTO of Dropbox and GitHub. We'd be happy to disclose our entire funding and profitability picture live during the interview. As a culture we relentlessly put transparency first. We conduct monthly financial reviews as a team so everyone has a pulse on the health of the business and publish what we are building in our [weekly changelog](https://rootly.com/changelog). ## About the Role This is an opportunity to join Rootly as an early SRE leader and shape our technical foundation. You will experience the balance of being scrappy and operating at scale. What you’ll be doing one day could look very different the next. You will be empowered to identify opportunities that will help us grow and own it. In short, this role is designed for individuals that crave ownership, stimulating technical challenges, love shipping fast, and are mission driven. We won’t sugarcoat it, the work will be challenging, but it will also be one of the most rewarding learning experiences of your career. - Embed with product teams to enhance observability, reliability, and performance of their services. - Own our CI/CD pipelines, observability tooling, monitoring systems, and incident response processes. - Build tools and automation to eliminate manual toil, improve engineering velocity and developer experience, and improve system reliability. - Collaborate deeply across engineering to understand systems at the code level and surface cross-cutting reliability, performance, and scaling concerns. - Architect and scale our infrastructure, ensuring best-in-class performance, availability, and operational excellence. - Drive capacity planning efforts to ensure our infrastructure is resilient and scalable as we grow. - Define and manage SLOs and error budgets in partnership with Engineering teams who own production services. - Be vocal - act as a strong voice and force of reliability, quality, performance, and scalability. About You'll Need: - 5+ years of experience in an SRE, Platform, or Infrastructure Engineering role. - 5+ years of experience writing software in a production environment. - Strong technical knowledge of cloud infrastructure, distributed systems, and reliability practices. - Strong understanding of observability, performance tuning, and scaling strategies. - Deep familiarity with incident response, monitoring, and CI/CD systems. - Hands-on experience supporting web or RPC services at meaningful scale. - You write code to solve infrastructure problems; not shell scripts alone, but production-grade software. ## Preferred Qualifications - You have a big-picture systems mindset and a proactive approach to reliability. - You’ve embedded with product teams and influenced design and architecture decisions. - You’re comfortable taking ownership of complex problems—and seeing them through. - Experience with Ruby and Go is a plus. Why Rootly? We’re not just another startup. We’re building something category-defining and want teammates who crave ownership, love solving hard problems, and thrive in a high-bar, high-impact environment. Here’s what you can expect when you join Rootly: - Competitive compensation and early equity in a fast-growing, venture-backed company. - Comprehensive medical, dental, and vision coverage. - 3 weeks of vacation, plus unlimited sick and mental health days, and a company-wide end-of-year shutdown to recharge. - $500 stipend for home office setup. - Unlimited token usage and access to AI tools - A fast-moving, high-impact environment where your leadership and ideas directly shape the future of the company. If this sounds like the kind of challenge and opportunity you’re looking for, apply now and let’s build something great together. Rootly is an equal opportunity employer. We aim to create an environment where every team member at Rootly feels like they belong so they can have a greater impact on our business and customers. We do not discriminate on the basis of race, religion, colour, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status. ## About Rootly ## Company Overview - **One-liner**: Rootly is an AI-native incident management and on-call platform that helps engineering teams resolve incidents faster, learn continuously, and sleep better. - **Entity Type**: Private (Series A, raised $15.3M total) - **Headquarters**: Toronto, Ontario, Canada (with an office in San Francisco) - **Founded**: 2021 - **Founders**: JJ Tang (CEO) and Quentin Rousseau (CTO/CISO) ## Core Business - **Primary industry**: Software Development / Incident Management / DevOps - **Target customers**: B2B, serving engineering teams at companies ranging from startups to large enterprises (e.g., Dropbox, Lattice, Webflow, Figma, LinkedIn, NVIDIA) - **Mission**: To help every organization become more reliable by bringing clarity and structure to incident response. ## Products & Services - **Incident Response**: End-to-end incident lifecycle management – declare incidents, spin up Slack channels, coordinate responders, track tasks, and automate operational busywork – all within Slack or Microsoft Teams. - **On-Call**: AI-powered scheduling, paging, escalation, and coverage requests with a mobile-first experience designed to reduce alert fatigue. - **AI SRE Agents**: Autonomous agents that connect signals (code changes, telemetry, past incidents) to provide probable root cause analysis, suggested fixes, and reasoning – helping responders move faster. - **Status Pages**: Public and private status pages to keep customers and stakeholders updated during incidents. - **Retrospectives**: Collaborative post-incident reviews with AI-generated timelines, summaries, and action-item tracking to turn incidents into learning. ## Market Standing - **Valuation/Market Cap**: Not publicly disclosed (private company) - **Key Metric**: Annual revenue of $1.2M (per LinkedIn); total funding of $15.3M (Seed + Series A) - **Notable Investors/Partners**: Y Combinator (S21), XYZ Venture Capital (Seed lead), Renegade Partners (Series A lead) - **Growth Signals**: 58 employees (+33.9% YoY); workforce distributed across 8 countries; active job postings up +1300% year-over-year; used by 100+ companies including Figma, LinkedIn, and NVIDIA. ## Competitive Advantages - **AI-native platform**: Built with AI at the core – automated RCA, AI meeting bot scribing, instant catch-up summaries, and AI-generated retros reduce cognitive load during incidents. - **Deep integrations**: Works seamlessly with Slack, Teams, Jira, Zoom, Terraform, API, and MCP server – allowing teams to operate without leaving their existing tools. - **Purpose-built for modern teams**: Codified from millions of real-world incidents, designed for teams that can’t afford downtime yet simple enough for anyone to use. ## Strategic Focus - **AI-driven reliability**: Continuing to embed AI across the incident lifecycle to automate diagnosis, remediation, and learning. - **Global expansion**: Hiring across US, Canada, Europe, and Singapore; building a distributed workforce to serve a global customer base. - **Product-led growth**: Offering a modular platform (incident response, on-call, AI SRE, status pages, retrospectives) that teams can adopt individually or as a whole. ## Why Work Here - **Culture**: Fast-moving, high-impact environment where employees own big problems from day one. Transparent interview process including a paid work trial to ensure mutual fit. - **Work policy**: Hybrid – hires across US, Canada, Europe, and Singapore; free lunches at HQ (Toronto); remote-friendly with flexible timezones. - **Perks**: Competitive salary, meaningful equity, 401(k) matching (US employees), top-tier medical/dental/vision/life insurance, flexible PTO, company holidays, team off-sites, year-end shutdown, annual home office stipend, and top-tier hardware. - **Engineering culture**: Obsession with building simple but powerful design, reliability-first infrastructure, and customer-driven iteration. AI is central to how the team builds and evaluates candidates. ## Sources 1. [rootly.com/careers](https://rootly.com/careers) 2. [rootly.com](https://rootly.com/) 3. [ca.linkedin.com/company/rootlyhq](https://ca.linkedin.com/company/rootlyhq) 4. [ycombinator.com/companies/rootly](https://www.ycombinator.com/companies/rootly) 5. [jobs.gem.com/rootly](https://jobs.gem.com/rootly) ## Other roles at Rootly - [Forward Deployed Technical Product Manager](https://feeny.ai/job/forward-deployed-technical-product-manager-rootly-office-42517371gs8f) — Office - [Member of Technical Staff, Intern](https://feeny.ai/job/member-of-technical-staff-intern-rootly-toronto-8q76y5dm7q4b) — Toronto, Canada - [Strategic Customer Success Manager (Bay Area)](https://feeny.ai/job/strategic-customer-success-manager-bay-area-rootly-bay-area-p4w9j18zvsn9) — Bay Area - [Enterprise Account Executive (Bay Area)](https://feeny.ai/job/enterprise-account-executive-bay-area-rootly-office-aw4hwphz4qnv) — Office - [AI Quality Engineer](https://feeny.ai/job/ai-quality-engineer-rootly-office-k4e9bcfc4h82) — Office - [Manager, Mid Market Sales](https://feeny.ai/job/manager-mid-market-sales-rootly-office-ajz6qs5dyarf) — Office - [Technical Support Engineer - APAC](https://feeny.ai/job/technical-support-engineer-apac-rootly-japan-7788299p06cv) — Japan - [Business Development Representative](https://feeny.ai/job/business-development-representative-rootly-office-50sqtwzy96k0) — Office - [Technical Support Engineer](https://feeny.ai/job/technical-support-engineer-rootly-office-w1yya8b3dqsg) — Office - [Senior Platform Engineer](https://feeny.ai/job/senior-platform-engineer-rootly-office-q8dx19585e5g) — Office