
Senior Site Reliability Engineer at Lodgify (Spain)
Lodgify· Spain·
Role details
Job description
⭐ Who we are
Lodgify is a fast-growing scale-up company leading the vacation rental industry. Backed by $30M in funding, our platform empowers property owners and managers worldwide to efficiently manage and grow their business through technology.
Headquartered in sunny Barcelona, we're now a team of 380+ people representing over 60 nationalities, united by a passion for transforming the future of short-term rentals.
⭐ How will you make an impact?
- Define meaningful SLIs, SLOs, and reliability targets for the platform.
- Collaborate with the software engineering teams to define and achieve the best practices for software observability, SLIs, SLOs and reliability.
- Strengthen production readiness by improving service ownership, observability, alerting, runbooks, scaling assumptions, rollback paths, and failure-mode preparedness.
- Improve the reliability, scalability, and performance of cloud, Kubernetes, and shared infrastructure, including how systems scale during growth, traffic spikes, and dependency failures.
- Build actionable observability using metrics, logs, traces, and golden signals, with tools such as Datadog, Prometheus, and Grafana.
- Implement operational and security best practices through guidelines, policies and automation.
- Reduce alert noise and improve signal quality so teams can detect, understand, and resolve issues quickly.
- Automate repetitive operational work using Python or other languages, turning recurring manual work into safer automation and clearer runbooks.
- Implement self-service Internal Developer Platform features via APIs and Kubernetes operators.
- Improve deployment safety, rollbackability, and release observability.
- Improve reliability of critical stateful systems such as databases, caches, queues, and streaming platforms.
- Participate in on-call, troubleshoot, and coordinate incident response, and facilitate blameless post-incident reviews that turn into concrete improvements.
- Execute disaster recovery drills and analyse cloud/platform usage to identify cost and resource-efficiency gains without compromising reliability.
⭐ What makes you a great fit?
- You have 7+ years of production experience operating Kubernetes-based platforms and cloud infrastructure.
- You understand and apply SRE practices: SLIs, SLOs, error budgets, production readiness, incident response, post-incident learning, toil reduction, scalability, capacity planning, high availability, backups, and disaster recovery.
- You can design and improve observability and alerting for critical systems using metrics, logs, traces, and golden signals, and are comfortable troubleshooting complex distributed systems to identify systemic reliability improvements.
- You can write maintainable software to automate operational tasks and reduce manual intervention.
- You have experience with stateful production systems such as relational databases, caches, queues, or streaming platforms.
- You know how to balance reliability, performance, cost, and delivery speed pragmatically.
- You are comfortable working in a transitional environment where SRE practices are being introduced while critical infrastructure and delivery systems still need hands-on reliability support.
- You collaborate effectively with Engineering, Platform, Security, and Product stakeholders.
- You communicate clearly, document well, and enjoy coaching teams toward stronger production ownership.
- You model initiative and accountability, raising risks early and driving improvements through to completion.
⭐ What does success look like?
- Critical services have clear owners, meaningful SLIs/SLOs, actionable alerts, dashboards, runbooks, and production readiness coverage.
- Reliability targets are consistently met across critical infrastructure and services.
- Operational toil and manual intervention are measurably reduced through automation and safer workflows.
- MTTR improves through reduced alert noise, better signal quality, stronger observability, and clear incident response playbooks and escalation paths.
- Post-incident actions are tracked, completed, and used to reduce repeat incidents.
- Disaster recovery exercises validate that critical services and infrastructure can recover within agreed expectations.
- Cloud and infrastructure resources are optimised without sacrificing performance, elasticity, or resilience.
Why you’ll love us: You’ll be part of a growing, dynamic company with a truly international team. At Lodgify, we are full of contagious energy, hard work, and passion for what we do. We celebrate diversity and are proud to acknowledge a variety of backgrounds, perspectives and skills in our team; committed to creating a workplace where everyone is heard and feels a sense of belonging.
What's in it for you?*
🏠 Remote Flexibility: The freedom to work from home any day that works for you. 🌴 Time to Recharge: 25 working days of paid vacation and Jornada Intensiva in August.. 💊 Alan Health Insurance: Premium health, dental, and mental health support via Alan. Pre-existing conditions are covered. 😋 Meal Perk:€150/month allowance on your Alan card + 50% off Ametller Origen prepared dishes at the office. 💸 Tax-Free Savings: Increase your take-home pay by using Flexible Remuneration for extra meal costs (up to €70/mo) and public transport (up to €136/mo). 🖥️ Home Office Gear: We provide a table, ergonomic chair, and monitor for your home setup. 🇪🇸 Language Learning: Free Spanish classes. 🤑 Referrals: Cash rewards for bringing in new talent. 🌟 Social Life: Daily office breakfast and monthly team events 🎯 Dynamic Hub: A high-energy, inclusive environment designed for collaboration and connection with a team that represents over 60 countries. *Benefits offered may differ based on the type of contract that is issued
So, what are you waiting for? Apply now! All applications and CVs must be submitted in English 😉
Why work at Lodgify
- Flexible work: Fully remote or hybrid options; choose the schedule that fits your life. lodgify.com/careers
- Time off: 25 working days of paid vacation plus public holidays, and Jornada Intensiva (shorter hours) in August.
- Health & wellness: Private health insurance via Alan (covers dental and mental health support, pre-existing conditions included).
- Perks: Daily breakfast, company offsites, beach volleyball, ping pong, and team social events.
- Growth: Access to internal training and external resources like Reforge and Udemy.
- Culture: High-energy, inclusive environment with 60+ nationalities; English is the working language. jobs.lever.co/lodgify
- Tax benefits: Flexible remuneration for meals (up to €220/mo) and public transport (up to €136/mo).