
Site Reliability Engineer (m/f/d) at Codesphere (Germany)
Codesphere· Germany·
Role details
Employee Equity Participation · 30+ Vacation Days · Meal Allowance · Hybrid Work Setup · Bike Lease (Job-Rad) · Gym Access · Employee Events · Company Pension Scheme
Summary
The Site Reliability Engineer defines and enforces SLOs, monitors system health, and automates deployments and infrastructure provisioning. This role involves diagnosing production incidents, leading post-mortems, and managing cloud infrastructure via IaC to ensure scalability, fault tolerance, and security compliance.
Job description
ABOUT CODESPHERE
Codesphere is a Virtual Cloud Provider from Germany building the future of sovereign cloud infrastructure. Our platform gives enterprises and governments full sovereignty without giving up modern cloud capability – a vision recently validated by a series of multi-million European government tenders.
Since our founding in Karlsruhe in 2020, we’ve expanded into an international team of 60+ experts. Based in Karlsruhe and Munich and backed by top-tier investors, we are chasing a bold vision.
We’re scaling fast and would love for you to join us and grow alongside us 🚀
WHAT YOU'LL DRIVE
- You define and enforce SLOs, SLIs, and SLAs across production
- You monitor system health, plan capacity, and automate deployments, patching, and infrastructure provisioning
- You diagnose and resolve production incidents fast – including 24/7 on-call participation
- You lead post-mortems and turn findings into prevention; maintain runbooks and escalation procedures
- You manage cloud infrastructure via IaC and own CI/CD pipeline design and maintenance
- You drive scalability, fault tolerance, disaster recovery, and security compliance
- You partner with Dev teams on production readiness, Shift Left practices, and error budget management
WHAT MAKES YOU A GREAT FIT
- Proven experience in an SRE, DevOps, or platform engineering role with hands-on production ownership
- Strong knowledge of Kubernetes, Terraform, and Ansible
- Familiarity with Ceph or comparable distributed storage systems
- Experience with SLOs, SLIs, error budgets, and CI/CD pipeline design
- Degree in a relevant field or comparable qualification
- Calm, structured, and fast under pressure – strong debugging and incident response skills
- Good communicator, able to translate operational concerns into guidance for Dev teams
- Go development experience is a plus
WHAT'S IN IT FOR YOU
- 32 days of paid time off – 30 regular vacation days plus Christmas Eve and New Year's Eve off
- Meal allowance – up to 15 digital vouchers per month, adding up to over €100 net for you
- Flexibility – hybrid work setup with mobile work options and flexibility around core hours
- Steep learning curve – fast-moving environment, real ownership, and a front-row seat to scaling a company
- Job-Rad – lease a bike through us, tax-free
- Gym access – stay active on site (Karlsruhe office only)
- Employee events – from team offsites to regular get-togethers
- Company pension scheme – company-supported pension to set you up for later
- Great public transport links – both offices are within walking distance of tram and metro stops
Why work at Codesphere
- Culture & Mission: Purpose-driven work focused on solving infrastructure dependency and digital sovereignty — a high-impact, meaningful challenge in European tech.
- Growth Trajectory: Rapidly scaling company (23% YoY headcount growth, 900% increase in job postings) offering career acceleration and ownership.
- Engineering Excellence: Strong engineering pedigree rooted in Baden-Württemberg, Germany's high-tech hub. Technical team makes up 31% of headcount. Focus on patented, novel technology (not just "cloud wrapping").
- International & Diverse: Team of 60+ people from 10+ nationalities across 9 countries. Offices in Karlsruhe (HQ), Munich, and New York.
- Hybrid/Remote: Multiple office locations with flexibility; strong German engineering culture with global reach.
- Notable Perks: Work on a patented deployment technology that genuinely differentiates in the market; opportunity to shape a product used by 50k+ users and large enterprises.