--- title: 'Senior Cloud Infrastructure Engineer at Langfuse' canonical: 'https://feeny.ai/job/senior-cloud-infrastructure-engineer-langfuse-europe-1mfv027gsv08' type: 'job' last_seen: '2026-09-07' --- # Senior Cloud Infrastructure Engineer at Langfuse - **Company:** Langfuse - **Location:** Europe - **Compensation:** €90k–€160k - **Employment:** full-time - **Work type:** hybrid - **Posted:** 2026-05-27 - **Last confirmed live:** 2026-09-07 - **Apply:** https://jobs.ashbyhq.com/langfuse/1bc2e248-89e7-41d7-b32f-08e9320eb5d0 ## Job description ## ABOUT LANGFUSE Open Source LLM Engineering Platform that helps teams build useful AI applications via tracing, evaluation, and prompt management (mission https://tracking.us.nylas.com/l/6d586a21a6fc4e1a8aacc7eb75882b72/0/82383757e54352130f65066e1b2fc4708aacab7897561bcb8000fe4c8a9c6a21?cache_buster=1761124921, product https://tracking.us.nylas.com/l/6d586a21a6fc4e1a8aacc7eb75882b72/1/b9fba3a93b6ffcc0f99ecda62767a17cc437fe8fe0b16181d1c43c1391212e3d?cache_buster=1761124921). We are now part of ClickHouse. We're building the "Datadog" of this category; model capabilities continue to improve, but building useful applications is really hard, both in startups and enterprises. Largest open source solution in this category: trusted by 19 of the Fortune 50, >2k customers, >26M monthly SDK downloads, >6M Docker pulls. We joined ClickHouse in January 2026 because LLM observability is fundamentally a data problem and Langfuse already ran on ClickHouse. Together we can move faster on product while staying true to open source and self-hosting, and join forces on GTM and sales to accelerate revenue. Previously backed by Y Combinator, Lightspeed, and General Catalyst. We're a small, engineering-heavy, and experienced team in Berlin and San Francisco. We are also hiring for engineering in EU timezones and expect one week per month in our Berlin office (how we work https://langfuse.com/handbook/how-we-work/principles). ## WHY CLOUD INFRASTRUCTURE AT LANGFUSE Your work will keep Langfuse running — everywhere. Langfuse processes over a billion trace events per month. When a Fortune 50 company relies on Langfuse in production, they're relying on the infrastructure you operate. You'll own uptime, performance, and cost efficiency across our entire cloud footprint — and you'll make sure every self-hosted deployment runs just as smoothly. You'll operate Langfuse Cloud on AWS ECS Fargate and ClickHouse Cloud, with Datadog as the observability backbone. You'll also own our public self-hosted infrastructure — including our Helm chart, Docker Compose setup, and everything in between — so that teams from startups to enterprises can run Langfuse on their own terms. This isn't a "maintain what exists" role. We're scaling fast, and you'll be the person who makes sure the infrastructure grows ahead of demand — not behind it. Langfuse is now part of ClickHouse, which means the team behind the database at the core of our stack is one channel away. Few infrastructure roles give you that kind of direct access to the people who build your most critical dependency. ## YOU WILL GROW AT LANGFUSE BY Own Langfuse Cloud operations: You'll run our production environments on AWS ECS Fargate and ClickHouse Cloud. You'll manage deployments, autoscaling, capacity planning, and cost optimization — making sure we stay fast and affordable as traffic scales. Build world-class observability: You'll own our Datadog setup end to end — dashboards, alerts, and SLOs. When something degrades, you'll ensure we know before our customers do. You'll build the monitoring culture that lets the whole team ship with confidence. Make self-hosting effortless: Thousands of teams run Langfuse on their own infrastructure. You'll own and evolve our Helm chart, Docker Compose configuration, and deployment documentation. You'll turn "works on my machine" into "works on every machine" — from a single-node setup to a multi-region enterprise deployment. Automate everything: CI/CD pipelines, infrastructure-as-code, automated scaling, zero-downtime deployments. You'll replace manual processes with automation that makes the team faster and the platform more reliable. Scale for what's next: We're growing fast and new product directions — like complex long-running agent observability and real-time evaluation — push the infrastructure in new ways. You'll be thinking ahead about what breaks at 10x scale and building the foundation before we get there. 10x is always just one quarter away here at Langfuse. Harden security and compliance: As more enterprises adopt Langfuse, you'll help ensure our cloud and self-hosted deployments meet the security and compliance bar that large organizations require. ## WHAT WE'RE LOOKING FOR - Strong infrastructure or SRE engineer who gets excited about running systems at scale and making them better every day - Experience operating production workloads on AWS (ECS/Fargate, networking, IAM, S3, etc.) or on comparable hyperscale vendors. - Comfortable with container orchestration — Kubernetes and/or ECS, Helm charts, Docker - Experience with infrastructure-as-code (Terraform, Pulumi, CloudFormation, or similar) - Strong monitoring and observability instincts — you've built dashboards and alerts that actually caught problems (Datadog experience is a plus) - You organize yourself. You have strong opinions about reliability, automation, and how to ship infrastructure changes safely - Interest in open source software and genuine enjoyment helping users debug their self-hosted deployments - Thrives in a small, accountable team where your output is visible and matters - CS or quantitative degree preferred Bonus points: - Experience with ClickHouse Cloud or other managed analytical databases - Background in operating high-throughput event processing or observability infrastructure - Contributions to open source infrastructure tooling (Helm charts, Terraform modules, etc.) - Former founder ## PROCESS We can run the full process to your offer letter in less than 7 days (hiring process https://langfuse.com/handbook/how-we-hire/hiring-process). ## TECH STACK We run a TypeScript monorepo: Next.js on the frontend, Express workers for background jobs, PostgreSQL for transactional data, ClickHouse for tracing at scale, S3 for file storage, and Redis for queues and caching. You should be familiar with a good chunk of this, but we trust you'll pick up the rest quickly (Stack https://langfuse.com/handbook/product-engineering/tech-stack, Architecture https://langfuse.com/handbook/product-engineering/architecture). ## HOW WE SHIP Link to handbook https://langfuse.com/handbook/how-we-work/principles - We trust you to take ownership (ownership overview https://langfuse.com/handbook/how-we-work/ownership) for your area. You identify what to build, propose solutions (RFCs), and ship them. Everyone here thinks about the user experience and the technical implementation at the same time. Everyone manages their own Linear. - You're never alone. Anyone from the team is happy to go into a whiteboard session with you. 15 minutes of shared discussion can very much improve the overall output. - We implement maker schedule and communication. There are two recurring meetings a week: Monday check-in on priorities (15 min) and a demo session on Fridays (60 min). - Code reviews are mentorship. New joiners get all PRs reviewed to learn the codebase, patterns, and how the systems work (onboarding guide https://langfuse.com/handbook/product-engineering/how-we-work/onboarding). - We use AI as much as possible in our workflows to make our users happy. We encourage everyone to experiment with new tooling and AI workflows. ## WHY LANGFUSE (NOW PART OF CLICKHOUSE) - This role puts you at the forefront of the AI revolution, partnering with engineering teams who are building the technology that will define the next decade(s). - This is an open-source devtools company. We ship daily, talk to customers constantly, and fight for great DX. Reliability and performance are central requirements. - Your work ships under your name. You'll appear on changelog posts https://langfuse.com/changelog for the features you build, and during launch weeks, you'll produce videos https://langfuse.com/blog/2025-10-29-launch-week-4 to announce what you've shipped to the community. You’ll own the full delivery end to end. - We're solving hard engineering problems: figuring out which features actually help users improve AI product performance, building SDKs developers love, visualizing data-rich traces, rendering massive LLM prompts and completions efficiently in the UI, and processing terabytes of data per day through our ingestion pipeline. - You'll work closely with the ClickHouse team and learn how they build a world-class infrastructure company. We're in a period of strong growth: Langfuse is growing organically and accelerating through ClickHouse's GTM. (Why we joined ClickHouse https://langfuse.com/blog/joining-clickhouse) - If you wonder what to build next, our users are a Slack message or a Github discussions post away. - You’re on a continuous learning journey. The AI space develops at breakneck speed and our customers are at the forefront. We need to be ready to meet them where they are and deliver the tools they need just-in-time. ## About Langfuse ## Company Overview - **One-liner**: Langfuse is an open-source AI engineering platform that provides observability, prompt management, evaluations, and analytics for teams building LLM applications. - **Entity Type**: Private (Acquired by ClickHouse in January 2026) - **Headquarters**: San Francisco, USA & Berlin, Germany - **Founded**: 2022 - **Founders**: Marc Klingen (CEO), Max Deichmann (CTO), Clemens Rawert (COO) ## Core Business - **Primary industry**: AI Engineering / LLM Observability & Evaluation - **Target customers**: B2B, Enterprise, and Developer teams building LLM applications - **Mission**: To help teams build, monitor, and improve LLM applications across the full development lifecycle, enabling more AI applications to reach production ## Products & Services - **Langfuse Platform**: An integrated open-source (MIT licensed) platform combining observability (hierarchical traces), prompt management, evaluations (LLM-as-a-judge, heuristics, human review), experiments, and analytics dashboards. It covers the full LLM engineering loop from prototype to production. - **Self-Hosted Deployment**: Available via Docker Compose, Kubernetes (Helm), and Terraform for AWS, GCP, and Azure. Scales to billions of monthly events. - **Cloud / Managed Service**: Free tier (50k observations/month) and paid enterprise tiers with 99.9% uptime SLA. - **SDKs & Integrations**: Native SDKs for Python, TypeScript, Go, Java, .NET, Ruby, PHP, Swift; OTel-native; 100+ integrations with frameworks (LangChain, Vercel AI SDK, CrewAI) and model providers (OpenAI, Anthropic, Google Gemini). - **MCP Server & CLI**: Developer tools for interacting with Langfuse from IDEs and CI/CD pipelines. ## Market Standing - **Valuation/Market Cap**: Not disclosed (acquired by ClickHouse in January 2026) - **Key Metric**: Processes 10+ billion observations per month; 50M+ SDK installs/month; 2,300+ customers - **Notable Investors/Partners**: Backed by Lightspeed, General Catalyst, Y Combinator, and angel investors. Acquired by ClickHouse (January 2026). Used by 19 of the Fortune 50, including Canva, Samsara, Twilio, Khan Academy, and Intuit. - **Growth Signals**: Acquired by ClickHouse in January 2026; 22,000+ GitHub stars; 5,000+ Discord members; 100,000+ engineers building on the platform; team grew to 19 people (as of YC profile); 50M+ SDK installs/month. ## Competitive Advantages - **Open Source (MIT License)**: Full product features are MIT licensed, allowing self-hosting, forking, and modification — no vendor lock-in. - **OTel Native & Async by Default**: Works with standard OpenTelemetry instrumentation and never blocks application performance. - **Built for Scale**: ClickHouse OLAP backend enables querying millions of traces in milliseconds; processes 10+ billion observations/month. - **Full Lifecycle Platform**: Combines observability, prompt management, evals, experiments, and human annotation in one integrated workflow — not just a point solution. - **Enterprise Trust**: Used by 19 of the Fortune 50, with strong security and compliance posture. ## Strategic Focus - **Post-Acquisition Growth**: Joining ClickHouse to accelerate growth, likely focusing on deeper integration with ClickHouse's OLAP database and expanding enterprise reach. - **Developer Experience**: Continued investment in MCP servers, CLI tools, and coding agent integrations (Claude Code, Cursor, Codex) to make the platform "made for developers, loved by agents." - **Open Source Community**: Maintaining the largest OSS community in the LLM observability category, with weekly releases and community hours. ## Why Work Here - **Strong Traction & Impact**: Work on a platform used by 19 of the Fortune 50, processing 10+ billion observations/month — your work directly impacts how the world's leading AI teams ship production LLM applications. - **Remote-First with In-Person Connection**: European roles are remote-first with one week per month in Berlin (travel covered); SF and Berlin locals use the office when useful. Flexible, async-friendly culture. - **Transparent Culture**: Everything is public — the company handbook, core principles, processes, and metrics. Radical transparency for alignment and community trust. - **Acquired by ClickHouse**: Joining a larger, well-funded infrastructure company (ClickHouse) provides stability, resources, and a broader platform for growth while maintaining Langfuse's product identity. - **Small, High-Impact Team**: ~19 people as of acquisition, meaning significant ownership and influence for new hires. - **Open Source Ethos**: MIT-licensed codebase, active community (22k+ GitHub stars, 5k Discord members), and a culture of contribution. ## Sources 1. [langfuse.com](https://langfuse.com/) 2. [langfuse.com/careers](https://langfuse.com/careers) 3. [jobs.ashbyhq.com/langfuse](https://jobs.ashbyhq.com/langfuse) 4. [ycombinator.com/companies/langfuse](https://www.ycombinator.com/companies/langfuse) 5. [github.com/langfuse/langfuse](https://github.com/langfuse/langfuse) ## Other roles at Langfuse - [Senior Software Engineer (SDK)](https://feeny.ai/job/senior-software-engineer-sdk-langfuse-europe-ays8et8hzgp1) — Europe - [Product Marketing Engineer](https://feeny.ai/job/product-marketing-engineer-langfuse-europe-305tp2579dxk) — Europe - [Developer Relations Engineer (Events & Community)](https://feeny.ai/job/developer-relations-engineer-events-community-langfuse-europe-or-qkkfh55y9w0k) — Europe OR, United States - [Product Marketing Manager](https://feeny.ai/job/product-marketing-manager-langfuse-europe-dqfgsg616ckn) — Europe - [Senior Product Engineer](https://feeny.ai/job/senior-product-engineer-langfuse-europe-yc9z6j0zgcy6) — Europe - [Senior Backend Engineer (Data Infrastructure)](https://feeny.ai/job/senior-backend-engineer-data-infrastructure-langfuse-europe-bt2mztjz0tr4) — Europe - [Senior Cloud Infrastructure Engineer](https://feeny.ai/job/senior-cloud-infrastructure-engineer-platform-science-brazil-44qj9rkp7bx5) — Brazil - [Senior Cloud Infrastructure Engineer](https://feeny.ai/job/senior-cloud-infrastructure-engineer-applied-intuition-london-vcz5xr3k4b1t) — London, United Kingdom - [Senior Cloud Infrastructure Engineer](https://feeny.ai/job/senior-cloud-infrastructure-engineer-trener-san-jose-24qjfesc8wsz) — San Jose, CA - [Senior Cloud Infrastructure Engineer](https://feeny.ai/job/senior-cloud-infrastructure-engineer-anduril-industries-washington-district-of-bkcarb7spxex) — Washington District of Columbia, United States