SkyPilot

Member of Technical Staff, GPU / ML Systems at SkyPilot (San Mateo, CA)

SkyPilot· San Mateo, CA·

Role details

Work type
Onsite
Employment
Full-Time

Job description

About SkyPilot

SkyPilot accelerates the world's most ambitious AI teams. Every hour they spend fighting infrastructure is an hour the frontier doesn't move — so SkyPilot turns fragmented compute across clusters into one optimized, highly available and easy-to-use pool: a single "AI supercomputer."

SkyPilot (10k+ GitHub stars, 14M+ downloads) is deployed at 100s of companies — from Fortune 500s to top AI-natives like Abridge, Applied Compute, Mistral, Unconventional AI, H Company, and Nubank — with usage growing exponentially. Born in the UC Berkeley lab behind Spark and Databricks, our growing team includes top-tier talent from Databricks, Google, Berkeley, MIT, CMU, and Cornell.

The role

SkyPilot exists because GPUs are scarce, expensive, and scattered — and the workloads that need them (pre-training, post-training, RL, high-throughput inference) push hardware to its limits. We're looking for an engineer to own the GPU and ML-systems layer that frontier AI teams run on: accelerator scheduling and utilization, the serving path, and the integrations that make SkyPilot the fastest, most cost-efficient place to run demanding AI workloads. A few points of GPU utilization here can save a team millions in compute and days on every training run.

What you'll do

  • Own GPU scheduling, utilization and health: how SkyPilot discovers, packs, and binpacks accelerator capacity across clouds and Kubernetes, with real-time GPU health monitoring and automatic failure recovery.
  • Build optimizations for training and serving: Enable large scale pre-training with node hot-swapping, design storage systems for fast model checkpointing, container migration, inference autoscaling and multi-cluster serving, preemption handling, and sandboxes for training, RL rollouts, and evals.
  • Make the AI stack run great out of the box: deepen integrations with vLLM, PyTorch, Slime,  and the frameworks teams use for pre-training and high-throughput inference.

What we're looking for

  • Hands-on experience with GPU or accelerator systems and with ML training or inference infrastructure.
  • Strongly preferred: Familiarity with the modern ML ecosystem (e.g. vLLM, PyTorch, CUDA, verl/slime) and workload-orchestration frameworks (e.g. Kueue, KAI, KServe).
  • You've done real ML-systems performance work - tell us about a bottleneck you hunted down (a stalled data pipeline, GPUs idling on a scheduling gap, communication you overlapped with compute) and what you measured before and after.
  • Strong Python, and comfort reaching into systems-level and GPU-adjacent details.
  • You care about squeezing most from the compute available to you
  • Experience operating large-scale training or high-throughput inference in production

What we offer

  • Competitive compensation and equity
  • Comprehensive medical, dental, vision coverage for you and your dependents
  • The chance to work with some of the best minds in cloud, distributed, and AI systems — with significant autonomy and ownership.
  • A front-row seat at the latest open-source infra startup from Berkeley (lineage: Databricks, Anyscale).
  • Gourmet lunch & dinner for the team to do their best work

Location: San Mateo, CA. Remote will be considered for exceptional candidates.

Why work at SkyPilot

  • Cutting‑edge mission: Building the foundational infrastructure for the AI industry alongside top hyperscalers and neoclouds.
  • Strong team: Founded by renowned systems researchers and backed by elite investors and AI leaders.
  • Traction & growth: Explosive adoption (6x growth in GPU hours) and a clear product‑market fit with frontier AI teams.
  • Autonomy & impact: Small, fast‑moving startup where engineers can directly shape a control plane used by thousands of GPUs.
  • Work policy – not publicly specified, but typical for a stealth‑stage startup; likely offers flexibility. The team is headquartered in San Francisco.
  • Careers: Open roles in Engineering and GTM (see SkyPilot on Gem).

Application questions