--- title: 'Senior Backend Engineer at Gremlin' canonical: 'https://feeny.ai/job/senior-backend-engineer-gremlin-based-in-the-efqcpk1r255r' type: 'job' last_seen: '2026-09-05' --- # Senior Backend Engineer at Gremlin - **Company:** Gremlin - **Location:** Based in the, United States - **Work type:** remote - **Posted:** 2025-03-28 - **Last confirmed live:** 2026-09-05 - **Apply:** https://job-boards.greenhouse.io/gremlin/jobs/7934212002 ## Job description Today’s complex, fast-paced systems have become a minefield of reliability risks—any of which could cause an outage that costs millions and destroys customer confidence. That’s why high-availability teams use the Gremlin to find and fix ‌reliability risks before they become incidents. Gremlin Reliability Platform helps software teams proactively monitor and test their systems for common reliability risks, build and enforce reliability standards, and automate their reliability practices organization-wide. As the industry leader in Chaos Engineering and reliability testing, we work with hundreds of the world’s largest organizations where high availability is non-negotiable. About the Role of the Senior Software Engineer As a Software Engineer at Gremlin, you will have the opportunity to improve the reliability of the internet at large by developing Chaos Engineering tooling. You will be able to leverage your engineering experience to inform product design as well as solve complex technical problems that directly impact our customers (which range from the Fortune 500 to smaller organizations). You will work closely with a small, talented team focused on quality, delivery, and predictability. In this role, you’ll get to: - Work closely with engineers, product managers, and other stakeholders to design and build the latest and greatest in Chaos Engineering tooling - Leverage strong collaboration and communication skills to deliver new features within a remote culture - Partner with product and other business units to understand business problems and present technical solutions and tradeoffs - Actively mentor and grow your teammates - Care deeply about the customer experience We'll expect you to have: - 5+ years professional Java software engineering experience - Experience in Go & Systems Level Programming - Experience in cloud technologies: e.g AWS, Lambda, Serverless. Experience with other cloud technologies like Google, Oracle also considered - Experience in DynamoDB and/or other no-sql DB or experience in any major relational databases - Experience in infrastructure & systems level technologies: e.g., Linux, Docker, Kubernetes, OpenShiftExperience in architecting complex distributed systems and integrating with external systems - Strong advocate and practitioner of automated testing, CI/CD, and engineering best practices Bonus Experience: - Has been on-call and participated in an incident management program - Familiarity with modern JavaScript frameworks & web development practices: e.g., React, TypeScript, etc. - Experience taking features from concept to full production release *The role does not offer sponsorship employment benefits. **If you don't think you meet all of the criteria below but still are interested in the job, please apply. Nobody checks every box—we’re looking for candidates that are particularly strong in a few areas, and have some interest and capabilities in others. ## Compensation We expect the salary range for this role to be $220,000 - $290,000. We recognize that salary varies from person to person depending on level of experience and we welcome direct conversations about it. The final offer will vary based on assessment of a candidate's skills and ability and our budget and market data. Gremlin offers competitive total compensation packages including 401k Matching, Equity and other benefits such as flexible time off and paid company holidays. About Gremlin: Gremlin is a team of industry veterans and people eager to learn from one another. We set the standard for reliability and equip leading organizations with the mindset and expertise needed to drive reliability improvements that move the world forward. We’re backed by top-tier investors Index Ventures, Amplify Partners, and Redpoint Ventures. Our customers love us, and we’re thrilled to be a partner in their success. What Do We Care About: - We Care about our People People are our critical differentiators. The company strives to treat our people with respect, empathy, and dignity. We expect that our people will treat each other similarly. In both cases, we will assume good intent. All are welcome at Gremlin. We know our differences make us stronger and that our best ideas and contributions can come from anyone at any level. - We Care about Collaboration Gremlin is strongest when we come together as one team with shared goals. Be the glue, not the glitter. But as a remote company, teamwork and collaboration won’t happen by accident. We approach every challenge as a shared challenge. We rely on each other for diverse perspectives and creative ideas. We celebrate our wins as a team. - We Care about Results Be high productivity, low drama. Results matter. To keep our pace, everyone owns the outcomes of their actions and takes action when needed. We reward speed over perfection. We empower each other to iterate and experiment.You are welcome at Gremlin for who you are. The more voices and ideas we have represented in our business, the more we will all flourish, contribute, and build a more reliable internet. Gremlin is a place where everyone can grow and is encouraged. However you identify and whatever background you bring with you, please apply if this sounds like a role that would make you excited to come into work everyday. It’s in our differences that we will find the power to keep building a more reliable internet by building and designing tools used by the best companies in the world. You are welcome at Gremlin for who you are. The more voices and ideas we have represented in our business, the more we will all flourish, contribute, and build a more reliable internet. Gremlin is a place where everyone can grow and is encouraged. However you identify and whatever background you bring with you, please apply if this sounds like a role that would make you excited to come into work everyday. It’s in our differences that we will find the power to keep building a more reliable internet by building and designing tools used by the best companies in the world. Visit our website to learn more - https://www.gremlin.com/about ## About Gremlin ## Company Overview - **One-liner**: Gremlin helps engineering teams proactively manage reliability at scale through resilience testing, risk detection, and disaster recovery validation. - **Entity Type**: Private (Series B – $28M total funding) - **Headquarters**: San Jose, California, United States - **Founded**: 2016 - **Founders**: Kolton Andrus (CEO) and co-founders who previously served as ‘Call Leaders’ at Amazon and Netflix (names not publicly disclosed) ## Core Business - **Primary industry/industries**: Software Development / Reliability Engineering (Chaos Engineering & Enterprise Reliability Management) - **Target customers**: B2B – primarily large enterprises (used by 4 of the top 5 US banks, plus major insurers, SaaS, retail, media) - **Mission or purpose statement**: “Helping every business build more reliable software.” (from careers page) ## Products & Services - **Gremlin Platform (Enterprise Reliability Management)**: Combined passive risk detection, dependency discovery, and active resilience/chaos testing. Provides a forward-looking view of service and application resilience. - **Resilience Testing**: Pre-built reliability test suites, customizable test suites, automated standardized testing across all services with reliability scores and trend tracking. - **Disaster Recovery Validation**: Simulate zone outages, region evacuations, and catastrophic failures from a central console; turn weeks of manual preparation into repeatable, auditable exercises. - **Risk Detection & Dependency Mapping**: Passively detect configuration drift, status errors, and deviations from standards; map hidden dependencies without requiring active tests. - **Reliability Intelligence (AI-powered)**: Custom-tailored experiment analysis, recommended remediations, and reliability insights. - **Failure Flags**: A new approach to testing application reliability on serverless platforms (e.g., AWS Lambda) and service meshes (Istio/Envoy). - **Private Edition**: Fully isolated version of Gremlin that can run within a customer’s private network. - **Reliability Reporting & Governance**: Executive dashboards, benchmarking across teams, and data-backed evidence for reliability investments. - **Gremlin Enterprise Chaos Engineering Certification (GECEC)**: Training and certification program (issued over 2K credentials in its first year). ## Market Standing - **Valuation/Market Cap**: Not disclosed - **Key Metric**: Annual Revenue – $12M (LinkedIn estimate); Total Funding – $28M (Seed $1.2M, Series A $8.8M, Series B $18M) - **Notable Investors/Partners**: Amplify Partners (Seed lead), Index Ventures (Series A lead), Redpoint (Series B lead). Partners include Dynatrace, AWS. - **Growth Signals**: 43 employees (-3.1% YoY decline, per LinkedIn); maintains 99.999% availability by using its own product; launched multiple new features in 2024-2026 (Disaster Recovery Testing, Reliability Intelligence, Detected Risks, Gremlin for AWS, Failure Flags for Istio); used by top global banks and major insurers; 90% reduction in downtime reported by customers. ## Competitive Advantages - **100% focused on reliability** – Not a side project; every line of code, every hire, every roadmap decision is dedicated to making customers more reliable. - **Proven at the largest enterprises** – Used by 4 of the top 5 US banks and other demanding organizations. - **Complete infrastructure coverage** – Supports bare metal, on-prem, multi-cloud, and serverless environments. - **AI-powered expert recommendations** combined with automated testing and reliability tracking. - **Safety-first design** – Blast radius management, halt conditions, and safety controls for testing in live production. - **Dogfooding** – Gremlin uses its own product to maintain 99.999% availability. ## Strategic Focus - **Current priorities**: Moving beyond Chaos Engineering to become the standard for Enterprise Reliability Management. Key areas include AI-driven reliability intelligence, automated risk detection, disaster recovery validation, and governance reporting to make reliability “fundable” with real metrics. - **Expansion**: Deepening partnerships (e.g., Dynatrace, AWS), launching new product modules (Failure Flags, Private Edition), and scaling the certification program. - **Customer segment**: Continuing to serve large enterprises while making reliability accessible to smaller teams through pre-built tests and automated discovery. ## Why Work Here - **Culture**: Remote-first company with a close-knit team; values “high productivity, low drama,” results over perfection, collaboration, and diversity. “Be the glue, not the glitter.” - **Remote/hybrid policy**: Fully remote-first; regular in-person get-togethers to build connections. - **Perks & benefits**: Generous benefits package supporting health and well-being of employees and their families. - **Engineering culture**: Built by reliability experts from Amazon, Netflix, Google, Salesforce, New Relic, etc. Team has deep incident management and on-call experience. Engineers use the product themselves, which fosters empathy and ownership. - **Growth opportunities**: Work on cutting-edge reliability challenges at scale; opportunities to contribute to the broader reliability community (certification, thought leadership). ## Sources 1. [gremlin.com](https://www.gremlin.com/) – Platform overview, use cases, customer results 2. [gremlin.com/about](https://www.gremlin.com/about) – Company timeline, founding, product milestones 3. [gremlin.com/team](https://www.gremlin.com/team) – Careers page, culture, values, benefits 4. [linkedin.com/company/gremlin-inc](https://www.linkedin.com/company/gremlin-inc.) – Revenue, funding, headcount, employee demographics, open roles ## Other roles at Gremlin - [Data Scientist, AI/ML](https://feeny.ai/job/data-scientist-ai-ml-gremlin-based-in-the-9mpav8tcgyga) — Based in the, United States - [Pre/Post Sales Solutions Architect](https://feeny.ai/job/pre-post-sales-solutions-architect-gremlin-based-in-the-vr0aem86thkw) — Based in the, United States - [Senior Backend Engineer](https://feeny.ai/job/senior-backend-engineer-verkada-san-mateo-ca-8cbw2zg7ym0t) — San Mateo CA, United States - [Senior Backend Engineer](https://feeny.ai/job/senior-backend-engineer-feverup-spain-15qsx6yx5r1d) — Spain - [Senior Backend Engineer](https://feeny.ai/job/senior-backend-engineer-sumup-london-england-9r9bgnz6t1sm) — London England, United Kingdom - [Senior Backend Engineer](https://feeny.ai/job/senior-backend-engineer-hyperexponential-warsaw-szqvq3t6aqem) — Warsaw, Poland - [Senior Backend Engineer](https://feeny.ai/job/senior-backend-engineer-peec-ai-berlin-1b3070vztc1y) — Berlin, Germany - [Senior Backend Engineer](https://feeny.ai/job/senior-backend-engineer-gitlab-bengaluru-wp9zybp29sx3) — Bengaluru, India - [Senior Backend Engineer](https://feeny.ai/job/senior-backend-engineer-revenuecat-americas-x4rjp8m2v339) — Americas - [Senior Backend Engineer](https://feeny.ai/job/senior-backend-engineer-partnerize-tel-aviv-1wa0yjjwm4h7) — Tel Aviv, Israel