--- title: 'Senior Network & Site Reliability Engineer at Alembic' canonical: 'https://feeny.ai/job/senior-network-site-reliability-engineer-alembic-san-francisco-mfdnj1j8nb1n' type: 'job' last_seen: '2026-09-07' --- # Senior Network & Site Reliability Engineer at Alembic - **Company:** Alembic - **Location:** San Francisco, CA - **Compensation:** $210k–$240k - **Employment:** full-time - **Work type:** onsite - **Posted:** 2026-06-12 - **Last confirmed live:** 2026-09-07 - **Apply:** https://jobs.ashbyhq.com/alembic/a56b5e4f-73c5-49e3-aee7-8c32f1bbf41f ## Job description ## ABOUT US Alembic is the pioneering Causal AI platform. We help the world's largest enterprises move past correlation to prove what actually drives business outcomes — the question marketing and growth teams have never been able to answer with confidence. Fortune 100 companies including Nvidia, Delta Air Lines, and Mars use Alembic to make multimillion-dollar decisions on trusted, causal evidence. We're backed by a $145M Series B from WndrCo (founded by Jeffrey Katzenberg), Jensen Huang, Joe Montana, Prysm Capital, and Accenture. Our models run on our own NVIDIA DGX SuperPOD built on Grace Blackwell infrastructure — one of the fastest private supercomputers in the world. (We've melted GPUs getting here.) ## ABOUT THE ROLE We're building infrastructure that has to perform under real-world scale, reliability, and security demands — and we're looking for an engineer who wants to own the foundation it runs on. This isn't a traditional "keep the lights on" role. You'll design and operate the global network and reliability layer behind one of the world's fastest private supercomputers — the fabric powering distributed compute, ML workloads, real-time analytics, and mission-critical enterprise systems. You'll work across networking, systems, automation, observability, and reliability engineering to scale a platform where performance genuinely matters, with real influence over architecture decisions. It's a strong fit if you like solving deep infrastructure problems, building resilient systems, automating everything repetitive, and owning architecture rather than just maintaining it. ## WHAT YOU'LL DO - Architect and operate scalable, secure network architecture for high-security requirements and large-scale machine learning workloads. - Own network device configuration management end to end, ensuring consistency and reliability across the fleet. - Improve system and network reliability and performance through automation, observability, and proactive capacity planning. - Implement and manage complex network protocols and connectivity, including BGP, VPNs, and WAN circuits and external peering. - Build and maintain comprehensive monitoring, alerting, and incident response — SLOs, runbooks, and on-call rotations — and drive post-incident analysis and continuous improvement. - Ensure security, compliance, and operational readiness across our network and cloud infrastructure. - Partner across engineering and data science to drive a culture of performance and reliability. ## WHAT WILL HELP YOU SUCCEED - 8+ years in network or infrastructure engineering, including 5+ years in datacenter operations and/or systems and network administration. - A strong background in network security, architecture, design, and operations. - Extensive hands-on experience with network devices (firewalls, switches, load balancers) and large-scale architectures and protocols — BGP, QoS, MPLS, and IPsec VPNs. - Experience designing and operating modern datacenter network fabrics (spine-leaf, EVPN/VXLAN, ECMP). - Network automation and IaC tooling (Ansible, Terraform, Nornir, or similar), plus IPAM/DCIM platforms (NetBox, Infoblox, or similar). - WAN engineering — carrier circuit provisioning and external network peering. - Familiarity with Kubernetes networking (CNI plugins, ingress, service networking, network policy) and strong operational experience with Linux-based production infrastructure. - Experience with monitoring and observability stacks (Prometheus, Grafana, Datadog, ELK, OpenTelemetry). - Solid scripting (Python, Bash) to debug complex network and system issues and automate solutions, plus excellent cross-functional communication. ## ALSO HELPFUL - NVIDIA networking technologies — Cumulus Linux, InfiniBand, Spectrum-X, and BlueField DPUs (this is the fabric behind our SuperPOD). - Familiarity with data-intensive platforms (Spark, Airflow, Kafka) and storage network protocols (NFS, LustreFS, iSCSI). - Security practices for applications and infrastructure, and experience in high-compliance or SOC 2 environments. ## THE ROLE IS RIGHT FOR YOU IF - You want to own mission-critical network and infrastructure end to end — from architecture to incident management — not just keep it running. - You'd rather build and automate than direct from a distance, and you want meaningful influence over how a high-performance platform scales. ## WHY YOU MIGHT BE EXCITED ABOUT ALEMBIC - Hard problems with real impact: You'll own the network and reliability layer behind systems that influence multimillion-dollar decisions at Fortune 100 companies. - Cutting-edge technology: Operate our own NVIDIA DGX SuperPOD on Grace Blackwell — one of the fastest private supercomputers in the world — and run a fabric (InfiniBand, Spectrum-X, BlueField) almost no company has in-house. - Technical autonomy: Ownership over architecture decisions and the freedom to solve hard infrastructure problems your way. - Elite team: Join top engineers who thrive on hard problems and high-impact work. - Series B momentum, real ownership: Meaningful equity at a Series B company that's raised $145M, with proven product-market fit and Fortune 100 traction. ## WHY YOU MIGHT NOT BE EXCITED - If you only want to tell people what to build instead of building and automating alongside them, this isn't the environment for you. - You prefer companies with 100% built-out process for every detail. - You prefer static over dynamic — projects and priorities adapt as we grow. We have real paying customers and a playbook, and we still move at startup speed at Series B scale. ## About Alembic ## Company Overview - **One-liner**: Alembic provides a real-time Causal AI platform that enables enterprises to measure, simulate, and optimize the true incremental impact of marketing investments on revenue. - **Entity Type**: Private (Series B) - **Headquarters**: San Francisco, California, United States - **Founded**: 2018 - **Founders**: Not publicly available ## Core Business - **Primary industry/industries**: Marketing analytics / Enterprise AI / Causal inference software - **Target customers**: B2B – Fortune 500 enterprises, large B2C brands, and sophisticated marketing organizations - **Mission or purpose statement**: To make blind, guess-based analysis obsolete by helping global enterprises act with clarity and confidence in the intelligence era. ## Products & Services - **Alembic Causal AI Platform (3.0)**: A SaaS platform that continuously recomputes large-scale causal graphs across billions of signals (media exposure, pricing, macroeconomics, consumer behavior, operational inputs). Enables real-time budget simulation, marginal ROI analysis, incremental lift measurement, and scenario testing without batch processing. - **Marketing Measurement & Attribution**: Replaces traditional MMM, MTA, and dashboards with dollar-for-dollar attribution across the full marketing funnel, isolating true causal impact per channel and tactic. - **Enterprise Intelligence**: Extends causal inference beyond marketing to areas like pricing, sponsorship evaluation, and go-to-market strategy for cross-functional decision-making. ## Market Standing - **Valuation/Market Cap**: $645 million (as of March 2026, per Techmeme) - **Key Metric**: Total Funding – $145M Series B led by Prysm Capital and Accenture (March 2026); prior rounds include $16M Series A (WndrCo-led) and $2.6M Seed (KB Partners-led). Total disclosed funding approximately $128M–$145M. - **Notable Investors/Partners**: Prysm Capital, Accenture, NextEquity, WndrCo, SLW, KB Partners. NVIDIA is the founding enterprise customer and exclusive supercomputing partner (first Causal AI company to secure NVIDIA DGX SuperPOD). - **Growth Signals**: Headcount grew 76.6% year-over-year (to ~57 employees). Active job postings for engineering and applied AI roles. Multiple Fortune 500 clients (e.g., major airline, B2B tech firm) with measurable pipeline and revenue impact. ## Competitive Advantages - **Causal AI vs. correlation**: Models cause-and-effect at scale, not just patterns; can simulate downstream impact of strategic decisions in real time. - **Real-time graph architecture**: Continuously recomputes causal graphs instead of running batch reports, giving live views of incremental revenue drivers. - **NVIDIA partnership**: Exclusive access to DGX NVL72 supercomputing infrastructure for enterprise decision-making, creating a hardware/software moat. - **Proven enterprise traction**: Demonstrated pipeline growth of 37% for a Fortune 500 B2B client and precise sponsorship measurement for a major airline. ## Strategic Focus - **Current priorities and direction for growth**: Expand causal AI applications beyond marketing into broader enterprise intelligence (pricing, operations, strategy). Scale platform adoption with Fortune 500 companies and deepen integrations with NVIDIA’s AI infrastructure. Continue hiring top AI research and engineering talent to advance causal inference capabilities. ## Why Work Here - **Culture highlights**: Values include “embrace curiosity,” “be data driven,” “build stuff & ship it,” “maintain low egos,” “don’t blame, solve,” “everyone teaches,” and “brevity is strength.” Emphasis on experimentation, cross-department learning, and shipping without overanalysis. - **Remote/hybrid/office policy**: In-person at San Francisco HQ (all open positions listed as “In Person Full Time”). - **Notable perks or engineering culture**: Work on frontier Causal AI models with access to NVIDIA DGX SuperPOD. Opportunity to solve complex, high-impact problems for Fortune 500 clients. Flat, low-ego environment with a focus on shipping and data-driven decisions. Active open roles: Research Engineer – Causal AI, Senior Applied AI Engineer, Senior Site Reliability Engineer, Patent Counsel, Senior Field Scientist. ## Sources 1. [Alembic Website – Platform Overview](https://alembic.com/) 2. [Alembic – About Us & Culture](https://alembic.com/company) 3. [Alembic – Careers Page](https://alembic.com/careers) 4. [LinkedIn – Alembic Technologies Company Profile](https://www.linkedin.com/company/getalembic) 5. [Techmeme – Alembic raises $145M Series B at $645M valuation](https://www.techmeme.com/) 6. [Accenture Newsroom – Accenture Invests in Alembic](https://newsroom.accenture.com/) 7. [Alembic – Job Board (Ashby)](https://jobs.ashbyhq.com/alembic) ## Other roles at Alembic - [Corporate IT Engineer](https://feeny.ai/job/corporate-it-engineer-alembic-san-francisco-r23xef67272p) — San Francisco, CA - [Personal Trainer](https://feeny.ai/job/personal-trainer-alembic-san-francisco-5z7ydpcvkh3x) — San Francisco, CA - [Technical Recruiter](https://feeny.ai/job/technical-recruiter-alembic-san-francisco-4n2tm5wcztb1) — San Francisco, CA - [Technical Sourcer](https://feeny.ai/job/technical-sourcer-alembic-san-francisco-nn6tpm5b3zch) — San Francisco, CA - [Patent Counsel](https://feeny.ai/job/patent-counsel-alembic-san-francisco-za6p9s63wmk8) — San Francisco, CA - [Senior Field Scientist](https://feeny.ai/job/senior-field-scientist-alembic-san-francisco-5rspg70tzw49) — San Francisco, CA - [Senior Applied AI Engineer](https://feeny.ai/job/senior-applied-ai-engineer-alembic-san-francisco-jh4y886w59at) — San Francisco, CA - [Senior Site Reliability Engineer](https://feeny.ai/job/senior-site-reliability-engineer-alembic-san-francisco-dmmytp9nn6x1) — San Francisco, CA - [Research Engineer - Causal AI](https://feeny.ai/job/research-engineer-causal-ai-alembic-san-francisco-ems0psx1gntj) — San Francisco, CA