--- title: 'Technical Product Manager, Infrastructure at Vast.ai' canonical: 'https://feeny.ai/job/technical-product-manager-infrastructure-vast-ai-los-angeles-b0h8y0t8d4fy' type: 'job' last_seen: '2026-09-10' --- # Technical Product Manager, Infrastructure at Vast.ai - **Company:** Vast.ai - **Location:** Los Angeles, CA - **Employment:** full-time - **Work type:** onsite - **Posted:** 2026-07-20 - **Last confirmed live:** 2026-09-10 - **Apply:** https://jobs.ashbyhq.com/vastai/8f32618b-c391-4987-84bd-d55f5623c457 ## Job description ## About Us Vast.ai's cloud powers AI projects and businesses all over the world. We are democratizing and decentralizing AI computing—reshaping our future for the benefit of humanity. We are a growing and highly motivated team dedicated to an ambitious technical plan. Our structure is flat, our ambitions are out-sized, and leadership is earned by shipping excellence. ## About the Role Vast.ai is seeking a Technical Product Manager to drive the backend our GPU cloud marketplace runs on. This is the software behind every live GPU rental (over 700k transactions a month): the daemon on every host's machines, the orchestration that powers each instance, and the infrastructure and test systems that let us ship it all at high velocity. The Scale: With over 20k GPUs, our AI cloud platform powers thousands of bleeding-edge training runs and critical production workloads for 120k+ developers all over the planet. This is a product with real scale, real data, and real users from day one. The work you ship moves tangible revenue in weeks, not quarters. The Challenge: We don't own the GPUs. Thousands of independent hosts price and operate them, and no two machines are alike. Your job is to make that heterogeneous, decentralized supply behave like the top-tier cloud our customers expect: fast, secure, and reliable. ## What You'll Own You'll drive the infrastructure, security, and reliability roadmap: sequencing competing priorities, turning non-functional requirements into specs engineering can build, and seeing them through until they ship. You'll own the metrics that prove it worked: platform uptime, fleet reliability, cost-to-serve. Initial focus areas: - Scalability & performance. Keep the core systems fast and ahead of demand as the marketplace grows, from database performance to end-to-end latency. - Security, trust & compliance. Harden the platform against attacks and abuse, and build the compliance roadmap enterprise customers need. - Observability & infra tooling. Give every engineer a clear view of platform health, and every host a clear view of their fleet. Own the metrics, tracing, and tooling behind it. Ideal Experience - 3+ years as a backend engineer AND 3+ years in product management - Hands-on experience with several of: security, scalability, reliability at scale, infrastructure tooling, distributed systems, observability, abuse/fraud prevention, compliance - Industry experience in one or more of the following areas: software infrastructure, developer tooling (APIs, SDKs, CLIs), AI/ML, cloud computing, GPUs, or two-sided marketplaces - Experience at a fast-paced startup or rapid-growth team You Are - AI-native. You work with AI agents every day, and you want to build the compute layer the AI era runs on. - Deeply technical. You came up as an engineer, and you reason about scalability, security, and efficiency as product concerns, not afterthoughts. - Persuasive across the stack. You can make a technical tradeoff legible to c-suite and a thankless migration compelling to the team doing it. - Biased toward action. You'd rather ship the fix today than present the plan next week. Tech Stack Python, C++, SQL/PostgreSQL, Redis, Linux, Docker, AWS, Terraform, REST APIs, LLMs/AI agents Interview Process - Initial screening (virtual) - Deep dive on the role and your experience (virtual) - Live product and technical assessment (virtual) - In person panel interview + product and technical assessment (on-site) Annual Salary Range $170,000–$240,000 base depending on experience, plus equity. We're profitable, venture-backed, and growing, so the equity is a stake in a real company with a real valuation, not a lottery ticket. ## Benefits - Comprehensive health, dental, vision, and life insurance - 401(k) with company match - Meaningful early-stage equity - Onsite meals, snacks, and close collaboration with founders/tech leaders - Ambitious, fast-paced startup culture where initiative is rewarded - Ample AI agent budget LOCATION: On-site at our office in Los Angeles (Westwood). ## About Vast.ai ## Company Overview - **One-liner**: Vast.ai operates a decentralized GPU marketplace that enables developers and AI agents to provision and manage compute power across a global network of hardware, offering a cost-effective alternative to hyperscaler clouds. - **Entity Type**: Private (raised $30M in funding) - **Headquarters**: Los Angeles, CA (1100 Glendon Ave, STE 1840) and San Francisco, CA (100 1st Street, STE 2250) - **Founded**: 2016 (incorporated June 28, 2016) - **Founders**: Jake Cannell (CEO) and Christian Horne ## Core Business - **Primary industry**: Cloud infrastructure / GPU compute / AI infrastructure - **Target customers**: AI researchers, ML engineers, AI agent developers, enterprises needing GPU compute for training and inference - **Mission**: "To organize, optimize, and orient the world's computation." - **Vision**: "To make life substrate-independent through Vast Artificial Intelligence." ## Products & Services - **GPU Cloud**: On-demand instances across 20,000+ GPUs in 40+ data centers. Deploy via CLI, SDK, Python API. Per-second billing. Real-time pricing set by supply and demand. [vast.ai](https://vast.ai/) - **Serverless Inference**: Deploy models as endpoints with automatic GPU optimization, auto-scaling to zero, pay-per-compute-time. [vast.ai](https://vast.ai/) - **GPU Clusters**: Dedicated multi-node clusters with InfiniBand networking for large-scale training. [vast.ai](https://vast.ai/) - **API & SDK**: REST API, Python SDK, CLI for programmatic compute provisioning. Agent-native interface for autonomous compute procurement. [vast.ai](https://vast.ai/) ## Market Standing - **Valuation**: Not publicly available - **Key Metric**: Total funding $30M (as of 2026, per jobsbyculture.com) [jobsbyculture.com](https://jobsbyculture.com/blog/working-at-vast-2026) - **Notable Investors/Partners**: Not disclosed in available sources - **Growth Signals**: - 310% year-over-year growth (2024–2025) - 20,000+ GPUs, 350+ hosts, 700K+ transactions/month - SOC 2 Type I certification achieved in 2024 - Enterprise and Secure Cloud offerings launched; customers include professional data center partners - Headcount grew from fully distributed team to 40+ employees across two offices (LA & SF) [vast.ai/about](https://vast.ai/about) - Conflicting reports: jobsbyculture.com cites ~30 employees, while vast.ai/about states "40+ employees" ## Competitive Advantages - **Decentralized marketplace model**: Taps into underutilized GPU hardware (gaming rigs, mining farms, research labs, small data centers) to offer compute at 3–5x cheaper than AWS, no contracts required. - **Agent-ready infrastructure**: The same API used by developers is designed for AI agents to autonomously procure and optimize compute – a moat for the upcoming agentic economy. - **Real-time, transparent pricing**: Prices set by supply and demand, programmatically queryable via API. No hidden costs or enterprise sales friction. - **Heterogeneous hardware support**: 68+ GPU types across 40+ data centers, enabling users to compare and switch between hardware types easily. - **SOC 2 certified**: Security and compliance for enterprise workloads. ## Strategic Focus - **Agentic compute**: Building an "infrastructure layer where AI agents design, procure, and optimize their own compute" [vast.ai](https://vast.ai/) - **Enterprise expansion**: Growing Secure Cloud (certified data centers) and dedicated cluster products for professional customers. - **Scaling the network**: Increasing GPU count, host partnerships, and geographic diversity while maintaining real-time pricing and low latency. - **Open infrastructure**: Keeping compute distributed and independent, countering hyperscaler concentration. ## Why Work Here - **Culture**: "High level of rigor, precision, and professionalism" – employees are stakeholders who own the impact of their work. Fast-paced startup environment where initiative is rewarded. [vast.ai/about](https://vast.ai/about) - **Work location**: On-site in Los Angeles (Westwood) or San Francisco (SOMA). Not remote – all current openings require on-site presence. - **Perks**: Comprehensive health/dental/vision insurance, 401(k) with company match, meaningful early-stage equity, onsite meals and snacks, close collaboration with founders and tech leaders. [vast.ai/jobs](https://vast.ai/jobs) - **Interview process**: Screened by technical team; stages include a 15-min screening, 45-min deep dive, 1-hour LLM-assisted coding assessment, and 2-hour on-site meet-and-greet. Aim to complete in about one week. [vast.ai/jobs](https://vast.ai/jobs) - **Salary ranges**: For roles like Senior Infrastructure Engineer and GPU Systems Engineer, posted salary $120K–$180K (may vary by role). [vast.ai/jobs](https://vast.ai/jobs) - **Engineering culture**: Reports directly to CEO/founder Jake Cannell, a prolific writer on AI and compute scaling theory. Emphasis on systems engineering, GPU optimization, and cutting-edge research. ## Sources 1. [vast.ai/about](https://vast.ai/about) 2. [vast.ai](https://vast.ai/) 3. [vast.ai/jobs](https://vast.ai/jobs) 4. [jobsbyculture.com](https://jobsbyculture.com/blog/working-at-vast-2026) 5. [vast.ai/jobs/apply/systems-gpu-research-engineer](https://vast.ai/jobs/apply/systems-gpu-research-engineer) ## Other roles at Vast.ai - [Head of Sales](https://feeny.ai/job/head-of-sales-vast-ai-los-angeles-yrkth8v89kkh) — Los Angeles, CA - [Developer Relations Engineer](https://feeny.ai/job/developer-relations-engineer-vast-ai-san-francisco-19dx882ce5x1) — San Francisco, CA - [Controller](https://feeny.ai/job/controller-vast-ai-los-angeles-0d4jvgnj6v48) — Los Angeles, CA - [AI, HPC & GPU Infrastructure Support Engineer](https://feeny.ai/job/ai-hpc-gpu-infrastructure-support-engineer-vast-ai-los-angeles-40e39w8ek9p1) — Los Angeles, CA - [Systems Operations Support Engineer — Linux](https://feeny.ai/job/systems-operations-support-engineer-linux-vast-ai-los-angeles-v3hjj8k7jj4v) — Los Angeles, CA - [Technical Support Engineer II (Linux)](https://feeny.ai/job/technical-support-engineer-ii-linux-vast-ai-los-angeles-0h1q79ajj6kc) — Los Angeles, CA - [Director of Engineering](https://feeny.ai/job/director-of-engineering-vast-ai-san-francisco-q3avgemppyhv) — San Francisco, CA - [Senior Infrastructure Engineer](https://feeny.ai/job/senior-infrastructure-engineer-vast-ai-los-angeles-ejanf64s8d34) — Los Angeles, CA - [Security Engineer](https://feeny.ai/job/security-engineer-vast-ai-los-angeles-5p7ef9m994ce) — Los Angeles, CA - [GPU Systems Engineer – HPC / Parallel Computing](https://feeny.ai/job/gpu-systems-engineer-hpc-parallel-computing-vast-ai-san-francisco-e7dn3x08y2sq) — San Francisco, CA