--- title: 'Staff AI Infrastructure Engineer at Biohub' canonical: 'https://feeny.ai/job/staff-ai-infrastructure-engineer-biohub-redwood-city-qvk5j1g1mxxh' type: 'job' last_seen: '2026-09-06' --- # Staff AI Infrastructure Engineer at Biohub - **Company:** Biohub - **Location:** Redwood City, CA - **Work type:** hybrid - **Posted:** 2026-04-02 - **Last confirmed live:** 2026-09-06 - **Apply:** https://job-boards.greenhouse.io/biohub/jobs/7775820 ## Job description Biohub is the first large-scale initiative bringing frontier AI models, massive compute, and frontier experimental capabilities under one roof. We're building a general-purpose system to accelerate scientific discovery, integrating frontier AI models, biological foundation models, and lab capabilities, with the ultimate goal of curing disease. Our technology powers scientists around the world, translating AI capabilities into tools that accelerate research everywhere. The Team The AI Cluster Production Engineering team is part of the AI Compute Platform organization at Biohub, a non-profit research lab committed to open science and open-source AI. We own the design, operation, and reliability of large-scale multi-GPU AI clusters that power frontier AI biology research: protein language models, genomic foundation models, and scientific reasoning systems built to be shared, not monetized. Our clusters run Slurm on Kubernetes infrastructure and support everything from day-to-day AI researcher workflows to multi-node hero training runs at thousands of GPUs. The team works at the intersection of AI tooling, distributed systems, HPC, and frontier AI, debugging deep AI infrastructure problems and building AI systems critical to the entire AI organization. The Opportunity CZ Biohub's mission is to cure or prevent all human disease. Achieving that requires training frontier-scale AI biology models, and that demands reliable, high-performance compute infrastructure. This is production engineering work at a frontier AI lab, with the twist that the mission is biology and the science is open. You'll keep GPU clusters running at high utilization, debug the toughest distributed systems failures, and build the operational foundations for scaling to multi-thousand GPU hero runs.  The technical problems are genuinely hard (e.g., multi-node distributed training, InfiniBand fabrics, large-scale storage, Slurm at scale) inside an organization where the work is aimed at helping people, not optimizing ad revenue. ## What You'll Do - Own reliability, observability, and incident response for multi-site GPU clusters running Slurm on Kubernetes. Build the systems, automation, and processes that keep clusters healthy,  and that enable fast, efficient recovery when things break. - Debug and resolve deep infrastructure failures across storage, networking, scheduling, and GPU compute layers. Build the tooling and operational patterns that make these failures easier to detect, diagnose, and prevent. - Design and execute GPU cluster scaling plans, systematically validating storage, networking, interconnect, and scheduler behavior as clusters grow to support larger training runs. - Build automation and tooling to manage cluster operations at scale: capacity planning, GPU utilization monitoring workload manager policy management, and pod lifecycle automation. - Drive configuration-as-code practices, ensuring cluster state is reproducible and auditable, and managed through version-controlled pipelines. - Collaborate directly with AI researchers and hero run leads to understand training workload patterns and design infrastructure that meets frontier-scale requirements. - Own the vendor relationship on technical issues — escalating SEV1s, coordinating across multiple partners and network backbone teams, holding them accountable to root/proximate cause analysis and SLAs. - Contribute to capacity planning: projecting GPU demand, managing cluster expansion across GPU generations, and coordinating multi-cluster strategy. - Improve operational resilience, reducing mean time to detect and resolve incidents, reducing toil through automation, and developing runbooks that scale the team's operational knowledge beyond any individual. ## What You'll Bring - 8+ years of AI/ML infrastructure engineering experience, with deep expertise in at least one of: HPC/Slurm cluster operations, Kubernetes at scale, distributed systems debugging, or GPU compute infrastructure. - Strong Linux systems fundamentals — networking (TCP/IP, InfiniBand, RDMA, MTU/MSS/PMTUD), storage (NFS, VAST, WEKA, POSIX semantics), kernel internals (cgroups, namespaces, eBPF, sysctls). - Hands-on experience with Kubernetes and cloud-native infrastructure — pod lifecycle, CNI plugins (Cilium preferred), StatefulSets, Helm, ArgoCD, or equivalent GitOps tooling. - Experience with HPC workload managers — Slurm strongly preferred (QoS, partitions, preemption, accounting, Sunk/CoreWeave patterns a plus). - Debugging instinct: ability to form hypotheses quickly, design controlled experiments, and root cause complex multi-system failures under pressure. You enjoy finding the hard bugs. - Proficiency in Python and Bash for automation and tooling. Go, Rust, or C/C++ a plus. - Experience with observability stacks — Prometheus/VictoriaMetrics, Grafana, DCGM metrics, distributed tracing. You know how to instrument systems you don't control. - Excellent communication — you can write a crisp incident summary for researchers, a technical escalation to a vendor CTO, and a system design doc for teammates, all in the same day. - Bonus: experience with distributed AI training infrastructure (NCCL, PyTorch DDP, multi-node job debugging, checkpoint/restart patterns, container environments for large-scale training). ## Compensation The Redwood City, CA base pay range for a new hire in this role is $241,000 - $331,000. New hires are typically hired into the lower portion of the range, enabling employee growth in the range over time. Actual placement in range is based on job-related skills and experience, as evaluated throughout the interview process. Better Together As we grow, we’re excited to strengthen in-person connections and cultivate a collaborative, team-oriented environment. This role is a hybrid position requiring you to be onsite for at least 60% of the working month, approximately 3 days a week, with specific in-office days determined by the team’s manager. The exact schedule will be at the hiring manager's discretion and communicated during the interview process. ## Benefits for the Whole You We’re thankful to have an incredible team behind our work. To honor their commitment, we offer a wide range of benefits to support the people who make all we do possible. - Provides a generous employer match on employee 401(k) contributions to support planning for the future. - Paid time off to volunteer at an organization of your choice. - Funding for select family-forming benefits. - Relocation support for employees who need assistance moving If you’re interested in a role but your previous experience doesn’t perfectly align with each qualification in the job description, we still encourage you to apply as you may be the perfect fit for this or another role. #LI-Hybrid ## About Biohub ## Company Overview - **One-liner**: Biohub (Chan Zuckerberg Biohub) is a nonprofit research organization building AI-powered tools and engineered cells to detect, treat, and ultimately cure age-related diseases. - **Entity Type**: Nonprofit (Private) – parent institution is the Chan Zuckerberg Initiative; funded by a $600 million endowment from Mark Zuckerberg and Priscilla Chan. - **Headquarters**: Redwood City, California, United States (with additional offices in San Francisco, Chicago, New York) - **Founded**: 2016 - **Founders**: Priscilla Chan and Mark Zuckerberg (co-founded as part of the Chan Zuckerberg Initiative; initial scientific leadership by Stephen Quake and Joseph DeRisi) ## Core Business - **Primary industry/industries**: AI-powered biology, biotechnology research, healthcare - **Target customers**: Scientists and researchers in academia and industry; indirectly benefits biotech and pharmaceutical companies through open-source models and foundational research; also invests in early‑stage companies via Science Ventures. - **Mission or purpose statement**: “To cure or prevent all disease” by combining frontier artificial intelligence with frontier biology to understand why disease happens and how to correct it. ## Products & Services - **AI Models for Biology**: Frontier AI models trained on large‑scale biological datasets (e.g., a world model of protein biology) that allow scientists to generate, test, and refine new protein designs. These are released as open discovery engines. - **Multi‑Dimensional Imaging Platforms**: Advanced tools (e.g., laser phase plate microscopy) that capture life from single proteins to whole organisms, revealing how cells function and communicate. Enables new AI models to predict cellular behavior. - **High‑Throughput Data Generation Engines**: Proprietary platforms for measuring, imaging, and programming biology at unprecedented scale, powering AI model training and hypothesis generation. - **Science Ventures**: An investment arm that supports early‑stage companies aligned with Biohub’s Grand Challenges (e.g., Somite AI, Adaptyx Biosciences). - **Grants & Collaborations**: Targeted grantmaking and open competitions to expand scientific research, including the Investigator Program for external scientists. ## Market Standing - **Valuation/Market Cap**: Not applicable (nonprofit); endowment of US$600 million (initial funding in 2016). - **Key Metric**: ~367 employees (as of mid‑2025); operating in 5 countries (US, UK, France, Canada, Netherlands); active job postings: 35+. - **Notable Investors/Partners**: Chan Zuckerberg Initiative (primary funder); academic partners include UC Berkeley, UCSF, Stanford; recent investments in Somite AI and Adaptyx Biosciences. - **Growth Signals**: Headcount growing 1.5% monthly; quarterly job posting increase of 133%; expanding into new locations (Chicago, New York); launching open‑source biological AI models. ## Competitive Advantages - **Scale of Compute & Data**: Unprecedented access to frontier AI compute clusters combined with unique, large‑scale biological datasets (genomics, proteomics, imaging). - **Open Science Model**: Results and models are shared openly, creating a virtuous feedback loop with the global research community. - **Talent Density**: Strong hiring pipeline from top institutions (Stanford, UC Berkeley, UCSF, Caltech) and a collaborative culture across AI, engineering, and biology. - **Founder Backing**: Deep, sustained funding from Mark Zuckerberg and Priscilla Chan ensures long‑term, risk‑taking research without commercial pressure. ## Strategic Focus - **Grand Challenges**: Decoding inflammation, early detection of age‑related diseases (cancer, Alzheimer’s, Parkinson’s), rare disease research, and building a “world model” of protein biology. - **AI‑First Biology**: Developing AI models that can predict cellular behavior and guide targeted treatment “only when and where needed.” - **Platform Scaling**: Expanding high‑throughput data generation engines and imaging tools to break through sparsity of biological data. - **Open Collaboration**: Accelerating translation from basic discovery to patient benefit through partnerships, grants, and open‑source releases. ## Why Work Here - **Mission‑Driven**: Opportunity to work at the intersection of AI and biology on problems that aim to eliminate the most significant causes of death worldwide. - **Culture & Team**: A collaborative team of scientists, engineers, and ML experts from top universities and companies (Genentech, EvolutionaryScale, etc.). Emphasis on “audacious, important scientific challenges.” - **Work Policy**: Hybrid model for many roles in Redwood City and New York; some positions are onsite in Chicago or San Francisco. - **Compensation & Perks**: Competitive salaries (e.g., Computational Biologist II ~$99k/yr, Clinical Lab Specialist ~$58k/yr); strong benefits typical of a well‑funded nonprofit; no equity, but meaningful mission impact. - **Growth**: Rapid hiring expansion across AI research, engineering, and biology; opportunity to work on frontier AI compute infrastructure and foundational biology models. ## Sources 1. [biohub.org](https://biohub.org/) 2. [Careers page – Greenhouse](https://job-boards.greenhouse.io/biohub) 3. [LinkedIn – CZ Biohub](https://www.linkedin.com/company/cz-biohub) 4. [Wikipedia – Chan Zuckerberg Biohub](https://en.wikipedia.org/wiki/Chan_Zuckerberg_Biohub) (referenced via search result) ## Other roles at Biohub - [Lab Manager, Aquaculture](https://feeny.ai/job/lab-manager-aquaculture-biohub-san-francisco-jbb8r64afpz6) — San Francisco, CA - [Postdoctoral Fellow/Scientist I, Synthetic Spatial Omics (Imaging Technology)](https://feeny.ai/job/postdoctoral-fellow-scientist-i-synthetic-spatial-omics-imaging-technology-1nyrb04bmy85) — New York, NY - [Scientist II, Scaling Lead](https://feeny.ai/job/scientist-ii-scaling-lead-biohub-san-francisco-9xk1n38bp90g) — San Francisco, CA - [Computational Biologist, Immune Cell Repolarization](https://feeny.ai/job/computational-biologist-immune-cell-repolarization-biohub-new-york-dt3na0g615eh) — New York, NY - [Research Associate II, Structural & Cellular Biology (Proteomics & CryoET)](https://feeny.ai/job/research-associate-ii-structural-cellular-biology-proteomics-cryoet-biohub-ge4pjxz6hq0k) — Redwood City, CA - [Staff Research Scientist, AI Safety](https://feeny.ai/job/staff-research-scientist-ai-safety-biohub-new-york-68fy1gymrhmg) — New York, NY - [Computational Biologist, Synthetic Spatial Omics](https://feeny.ai/job/computational-biologist-synthetic-spatial-omics-biohub-new-york-srwtkdew50sf) — New York, NY - [Principal Technical Program Manager, Science, Imaging](https://feeny.ai/job/principal-technical-program-manager-science-imaging-biohub-redwood-city-v9dbd6a743s4) — Redwood City, CA - [Director, Virtual Biology Initiative](https://feeny.ai/job/director-virtual-biology-initiative-biohub-redwood-city-pwf0p3w6k56w) — Redwood City, CA - [Staff Data Scientist, Imaging](https://feeny.ai/job/staff-data-scientist-imaging-biohub-redwood-city-20s3vpyx8gb5) — Redwood City, CA