--- title: 'Member of Technical Staff, GPU / ML Systems at SkyPilot' canonical: 'https://feeny.ai/job/member-of-technical-staff-gpu-ml-systems-skypilot-san-mateo-cjcq4ysp96ka' type: 'job' last_seen: '2026-09-11' --- # Member of Technical Staff, GPU / ML Systems at SkyPilot - **Company:** SkyPilot - **Location:** San Mateo, CA - **Employment:** full-time - **Work type:** onsite - **Posted:** 2026-07-20 - **Last confirmed live:** 2026-09-11 - **Apply:** https://jobs.gem.com/skypilot/am9icG9zdDqzLbsInR-OJIk0z4fMcX5z ## Job description ## About SkyPilot SkyPilot accelerates the world's most ambitious AI teams. Every hour they spend fighting infrastructure is an hour the frontier doesn't move — so SkyPilot turns fragmented compute across clusters into one optimized, highly available and easy-to-use pool: a single "AI supercomputer." SkyPilot ([10k+ GitHub stars](https://github.com/skypilot-org/skypilot), 14M+ downloads) is deployed at 100s of companies — from Fortune 500s to top AI-natives like Abridge, Applied Compute, Mistral, Unconventional AI, H Company, and Nubank — with usage growing exponentially. Born in the UC Berkeley lab behind Spark and Databricks, our growing team includes top-tier talent from Databricks, Google, Berkeley, MIT, CMU, and Cornell. ## The role SkyPilot exists because GPUs are scarce, expensive, and scattered — and the workloads that need them (pre-training, post-training, RL, high-throughput inference) push hardware to its limits. We're looking for an engineer to own the GPU and ML-systems layer that frontier AI teams run on: accelerator scheduling and utilization, the serving path, and the integrations that make SkyPilot the fastest, most cost-efficient place to run demanding AI workloads. A few points of GPU utilization here can save a team millions in compute and days on every training run. ## What you'll do - Own GPU scheduling, utilization and health: how SkyPilot discovers, packs, and binpacks accelerator capacity across clouds and Kubernetes, with real-time GPU health monitoring and automatic failure recovery. - Build optimizations for training and serving: Enable large scale pre-training with node hot-swapping, design storage systems for fast model checkpointing, container migration, inference autoscaling and multi-cluster serving, preemption handling, and sandboxes for training, RL rollouts, and evals. - Make the AI stack run great out of the box: deepen integrations with vLLM, PyTorch, Slime,  and the frameworks teams use for pre-training and high-throughput inference. ## What we're looking for - Hands-on experience with GPU or accelerator systems and with ML training or inference infrastructure. - Strongly preferred: Familiarity with the modern ML ecosystem (e.g. vLLM, PyTorch, CUDA, verl/slime) and workload-orchestration frameworks (e.g. Kueue, KAI, KServe). - You've done real ML-systems performance work - tell us about a bottleneck you hunted down (a stalled data pipeline, GPUs idling on a scheduling gap, communication you overlapped with compute) and what you measured before and after. - Strong Python, and comfort reaching into systems-level and GPU-adjacent details. - You care about squeezing most from the compute available to you - Experience operating large-scale training or high-throughput inference in production ## What we offer - Competitive compensation and equity - Comprehensive medical, dental, vision coverage for you and your dependents - The chance to work with some of the best minds in cloud, distributed, and AI systems — with significant autonomy and ownership. - A front-row seat at the latest open-source infra startup from Berkeley (lineage: Databricks, Anyscale). - Gourmet lunch & dinner for the team to do their best work Location: San Mateo, CA. Remote will be considered for exceptional candidates. ## About SkyPilot ## Company Overview - **One-liner**: SkyPilot provides a unified AI compute platform that turns fragmented GPU clusters across hyperscalers, neoclouds, and on-premises resources into a single, manageable “AI supercomputer.” - **Entity Type**: Private (Seed stage – raised $20M) - **Headquarters**: San Francisco, California, USA - **Founded**: Not publicly available (company launched from stealth in July 2026; technology originated from UC Berkeley’s Sky Computing Lab) - **Founders**: Zongheng Yang (CEO), Zhanghao Wu, Romil Bhardwaj, alongside Databricks co-founder Ion Stoica and networking pioneer Scott Shenker (both Berkeley professors) ## Core Business - **Primary industry**: AI infrastructure / cloud computing orchestration - **Target customers**: Frontier AI teams – including research labs, foundation model builders, healthcare, financial services, and large enterprises – that need to manage large-scale GPU fleets across multiple providers. - **Mission / Purpose**: “Turn fragmented compute into one AI supercomputer” so that AI teams can spend less time managing infrastructure and more time building intelligence. ## Products & Services - **[SkyPilot Platform (Managed)](https://skypilot.ai)**: A managed control plane that abstracts heterogeneous compute (Kubernetes, Slurm, VMs, hyperscalers, neoclouds) into a single pool. Supports interactive development, pre‑training at scale, sandboxes for RL/agents, multi‑cluster serving, and production workloads. Includes GPU health monitoring, auto‑remediation, HA, quota management, SSO, RBAC, and SOC 2 compliance. Customers have seen up to 20x performance improvements over the open‑source version. - **[SkyPilot Open Source](https://skypilot.ai)**: Free, community‑driven project with 14M+ downloads and 280+ contributors. Used by hundreds of organizations to unify compute across 20+ clouds. Provides CLI, intelligent scheduler, GPU monitoring, and multi‑cluster support. ## Market Standing - **Valuation / Market Cap**: Not disclosed (seed round) - **Key Metric (Private)**: **Total Funding** – $20M seed (July 2026) - **Notable Investors & Partners**: - Lead: Lux Capital - Participants: Amplify Partners, Coatue Management, Foundation Capital, Race Capital, The House Fund - Angel operators: Ali Ghodsi (Databricks), Jeff Dean (Google), Guillermo Rauch (Vercel), Amjad Masad (Replit), Clem Delangue (Hugging Face), Tristan Handy (dbt Labs) - Cloud partners: Nebius, CoreWeave, Lambda, AWS (EKS, EC2, S3, SageMaker Hyperpod) - **Growth Signals**: - 14M+ open‑source downloads (6M in last 3 months) - GPU hours consumed grew 35% MoM (6x in last 6 months) - Top deployments: 1,000+ nodes and 10,000+ GPUs - 280+ open‑source contributors - Adopted by Nubank, Abridge, Applied Compute, H Company, Archer, Hippocratic AI ## Competitive Advantages - **Unified abstraction** across 20+ clouds and any accelerator – customers are not locked into a single provider. - **AI‑native scheduling** (topology‑aware, priority queueing, GPU sharing, preemption) that maximizes fleet utilization. - **Proactive GPU health checks** with auto‑remediation – reduces downtime in large training runs. - **Enterprise‑grade security** (BYOC, BYOK, private VPCs, air‑gapped, SOC 2) combined with the agility of an open‑source ecosystem. - **World‑class founding team** – includes a Databricks co‑founder and UC Berkeley systems researchers. ## Strategic Focus - **Product development** – continue building out the managed platform (already 20x faster than OSS). - **Engineering expansion** – actively hiring in engineering and go‑to‑market roles. - **Go‑to‑market growth** – opening platform access to select new customers; scaling sales. - **Open‑source ecosystem** – deepen contributions and community engagement. ## Why Work Here - **Cutting‑edge mission**: Building the foundational infrastructure for the AI industry alongside top hyperscalers and neoclouds. - **Strong team**: Founded by renowned systems researchers and backed by elite investors and AI leaders. - **Traction & growth**: Explosive adoption (6x growth in GPU hours) and a clear product‑market fit with frontier AI teams. - **Autonomy & impact**: Small, fast‑moving startup where engineers can directly shape a control plane used by thousands of GPUs. - **Work policy** – not publicly specified, but typical for a stealth‑stage startup; likely offers flexibility. The team is headquartered in San Francisco. - **Careers**: Open roles in Engineering and GTM (see [SkyPilot on Gem](https://jobs.gem.com/skypilot)). ## Sources 1. [skypilot.ai (homepage)](https://skypilot.ai) 2. [skypilot.ai (about)](https://skypilot.ai/about) 3. [skypilot.ai (funding announcement blog)](https://skypilot.ai/blog/skypilot-the-company) 4. [fortune.com](https://fortune.com/2026/07/21/skypilot-from-databricks-cofounder-raises-20m-to-be-the-switzerland-of-ai-compute/) 5. [prnewswire.com](https://www.prnewswire.com/news-releases/skypilot-launches-with-20m-to-accelerate-custom-intelligence-for-frontier-ai-teams-302830808.html) ## Other roles at SkyPilot - [GTM Specialist](https://feeny.ai/job/gtm-specialist-skypilot-san-mateo-c1js372pj0j0) — San Mateo, CA - [Product Marketing Manager](https://feeny.ai/job/product-marketing-manager-skypilot-san-mateo-ac5fset49yp9) — San Mateo, CA - [Member of Technical Staff, Product](https://feeny.ai/job/member-of-technical-staff-product-skypilot-san-mateo-3fvgs8gy6hqw) — San Mateo, CA - [Member of Technical Staff, Platform](https://feeny.ai/job/member-of-technical-staff-platform-skypilot-san-mateo-4efm1nkm0rsp) — San Mateo, CA - [General Application](https://feeny.ai/job/general-application-skypilot-san-mateo-jht04k9gw4xr) — San Mateo, CA - [Founding Account Executive](https://feeny.ai/job/founding-account-executive-skypilot-san-mateo-jxym4wpcxwb5) — San Mateo, CA - [Forward Deployed Engineer](https://feeny.ai/job/forward-deployed-engineer-skypilot-san-mateo-eyv68cyw4ghc) — San Mateo, CA - [Developer Relations Engineer](https://feeny.ai/job/developer-relations-engineer-skypilot-san-mateo-sw1asm63sxv5) — San Mateo, CA - [Member of Technical Staff, Distributed Systems](https://feeny.ai/job/member-of-technical-staff-distributed-systems-skypilot-san-mateo-g32j2zaakscz) — San Mateo, CA