--- title: 'Technical Support Engineer II (Linux) at Vast.ai' canonical: 'https://feeny.ai/job/technical-support-engineer-ii-linux-vast-ai-los-angeles-0h1q79ajj6kc' type: 'job' last_seen: '2026-09-10' --- # Technical Support Engineer II (Linux) at Vast.ai - **Company:** Vast.ai - **Location:** Los Angeles, CA - **Compensation:** $90k–$130k - **Employment:** full-time - **Work type:** onsite - **Posted:** 2026-07-17 - **Last confirmed live:** 2026-09-10 - **Apply:** https://jobs.ashbyhq.com/vastai/4bad56d6-a3e3-4b25-8800-f37410791955 ## Job description ## About Us [Vast.ai](http://Vast.ai)'s cloud powers AI projects and businesses all over the world. We are democratizing and decentralizing AI computing — reshaping our future for the benefit of humanity. Our mission is to organize, optimize, and orient the world's computation. We value elegance, ownership, integrity, and continuous learning. You'll have the opportunity to dive into state-of-the-art AI systems while collaborating with a globally distributed team. ## About the Role This is a technical support role focused on escalated infrastructure issues that go beyond frontline triage. You'll be the engineering resource our L1 support team leans on when tickets get complex: diagnosing and resolving issues across the full stack — hardware/BIOS/firmware, networking, Ubuntu, Docker, NVIDIA CUDA/GPU, and virtualization (KVM). You'll handle higher-complexity issues, own escalation resolution end-to-end, and contribute to internal documentation and runbooks. The best engineers in this role don't just resolve tickets — they build the tooling and runbooks that eliminate recurring ones. You'll collaborate directly with the engineering team and host support team on systemic issues. Strong technical depth and support experience are the primary requirements. You should be comfortable working autonomously across Ubuntu environments, diagnosing container and GPU issues, and communicating findings clearly to both technical and non-technical audiences. [Vast.ai](http://Vast.ai) users or hosts strongly preferred. This role is full-time and onsite in our office in Westwood (LA) Schedule: Sunday - Thursday. ## Key Responsibilities - Handle escalated support tickets, including GPU workload failures, container issues, networking problems, account infrastructure, and host-side configuration - Diagnose and resolve issues across Docker, NVIDIA CUDA/GPU drivers, and virtualization environments (KVM) - Troubleshoot network-layer issues: VLAN, DNS, DHCP, VPN, NAT, firewall rules, and connectivity failures on host machines - Investigate performance issues on GPU utilization, container resource constraints, thermal throttling, driver conflicts, disk I/O bottlenecks - Advise suppliers (hosts) on installation best practices — hardware setup, driver configuration, BIOS/firmware settings, and network configuration for optimal performance - Provide managed support for supplier onboarding and ongoing machine management, acting as a technical resource through installation, configuration, and post-setup troubleshooting - Write and maintain internal runbooks, escalation guides, and knowledge base articles to reduce repeat escalations - Build diagnostic and automation tooling in Python and Bash to reduce manual triage overhead - Collaborate with the engineering team and infrastructure support team to flag and document systemic or recurring platform issues - Assist clients and infrastructure suppliers working with AI frameworks (TensorFlow, PyTorch) and GPU-accelerated workloads - Provide coverage for L1 support team overflow during peak periods or incidents, per a defined on-call rotation You Are - Fluent in Linux — you navigate systems, read logs, and solve problems from the command line without hesitation - Methodical and thorough: you gather data, dig into root causes, and don't settle for surface-level fixes - A self-starter who can manage a queue of complex tickets with minimal supervision - Adaptable to a defined on-call rotation which may include weekend coverage - A clear written communicator: able to explain technical findings and write useful internal documentation - Genuinely curious about AI infrastructure, GPU computing, and distributed systems Must-Haves - Solid Linux SysOps experience: Ubuntu Server, RHEL/CentOS, Debian; comfortable with systems, networking, storage, and permissions - Proficiency with Docker: container debugging, Docker Compose, image management, cgroup resource limits, Docker storage/filesystem management - Experience with virtualization: Proxmox VE, VMware, or similar hypervisors; provisioning and troubleshooting VMs - Networking fundamentals: VLAN, DNS, DHCP, NAT, VPN, firewall rules, and general L2/L3 troubleshooting - Hands-on experience with NVIDIA GPU drivers, CUDA, and GPU workload troubleshooting (essential) - Scripting in Python and Bash for automation and diagnostic tooling - Strong English written communication: clear, professional, and technically precise - Experience providing technical support in a customer-facing or internal helpdesk context - Ability to prioritize across a concurrent queue of escalated tickets, triaging by severity and customer impact, balancing reactive resolution against proactive documentation and tooling work, and making clear judgment calls on when to escalate versus own resolution end-to-end Nice-to-Haves - Familiarity with AI/ML frameworks (TensorFlow, PyTorch) and running GPU-accelerated containers - Monitoring and observability experience (Prometheus, Grafana) - Relevant certifications: RHCSA, CompTIA Linux+, or similar - Knowledge of the [Vast.ai](http://Vast.ai) platform as a client or infrastructure supplier Annual Salary Range $90,000 – $150,000 + equity + benefits [Vast.ai](http://Vast.ai) is hiring across all experience levels with compensation commensurate with background, experience and potential. ## Benefits - Comprehensive health, dental, vision, and life insurance - 401(k) with company match - Meaningful early-stage equity - Onsite meals, snacks, and close collaboration with founders/tech leaders - Ambitious, fast-paced startup culture where initiative is rewarded ## About Vast.ai ## Company Overview - **One-liner**: Vast.ai operates a decentralized GPU marketplace that enables developers and AI agents to provision and manage compute power across a global network of hardware, offering a cost-effective alternative to hyperscaler clouds. - **Entity Type**: Private (raised $30M in funding) - **Headquarters**: Los Angeles, CA (1100 Glendon Ave, STE 1840) and San Francisco, CA (100 1st Street, STE 2250) - **Founded**: 2016 (incorporated June 28, 2016) - **Founders**: Jake Cannell (CEO) and Christian Horne ## Core Business - **Primary industry**: Cloud infrastructure / GPU compute / AI infrastructure - **Target customers**: AI researchers, ML engineers, AI agent developers, enterprises needing GPU compute for training and inference - **Mission**: "To organize, optimize, and orient the world's computation." - **Vision**: "To make life substrate-independent through Vast Artificial Intelligence." ## Products & Services - **GPU Cloud**: On-demand instances across 20,000+ GPUs in 40+ data centers. Deploy via CLI, SDK, Python API. Per-second billing. Real-time pricing set by supply and demand. [vast.ai](https://vast.ai/) - **Serverless Inference**: Deploy models as endpoints with automatic GPU optimization, auto-scaling to zero, pay-per-compute-time. [vast.ai](https://vast.ai/) - **GPU Clusters**: Dedicated multi-node clusters with InfiniBand networking for large-scale training. [vast.ai](https://vast.ai/) - **API & SDK**: REST API, Python SDK, CLI for programmatic compute provisioning. Agent-native interface for autonomous compute procurement. [vast.ai](https://vast.ai/) ## Market Standing - **Valuation**: Not publicly available - **Key Metric**: Total funding $30M (as of 2026, per jobsbyculture.com) [jobsbyculture.com](https://jobsbyculture.com/blog/working-at-vast-2026) - **Notable Investors/Partners**: Not disclosed in available sources - **Growth Signals**: - 310% year-over-year growth (2024–2025) - 20,000+ GPUs, 350+ hosts, 700K+ transactions/month - SOC 2 Type I certification achieved in 2024 - Enterprise and Secure Cloud offerings launched; customers include professional data center partners - Headcount grew from fully distributed team to 40+ employees across two offices (LA & SF) [vast.ai/about](https://vast.ai/about) - Conflicting reports: jobsbyculture.com cites ~30 employees, while vast.ai/about states "40+ employees" ## Competitive Advantages - **Decentralized marketplace model**: Taps into underutilized GPU hardware (gaming rigs, mining farms, research labs, small data centers) to offer compute at 3–5x cheaper than AWS, no contracts required. - **Agent-ready infrastructure**: The same API used by developers is designed for AI agents to autonomously procure and optimize compute – a moat for the upcoming agentic economy. - **Real-time, transparent pricing**: Prices set by supply and demand, programmatically queryable via API. No hidden costs or enterprise sales friction. - **Heterogeneous hardware support**: 68+ GPU types across 40+ data centers, enabling users to compare and switch between hardware types easily. - **SOC 2 certified**: Security and compliance for enterprise workloads. ## Strategic Focus - **Agentic compute**: Building an "infrastructure layer where AI agents design, procure, and optimize their own compute" [vast.ai](https://vast.ai/) - **Enterprise expansion**: Growing Secure Cloud (certified data centers) and dedicated cluster products for professional customers. - **Scaling the network**: Increasing GPU count, host partnerships, and geographic diversity while maintaining real-time pricing and low latency. - **Open infrastructure**: Keeping compute distributed and independent, countering hyperscaler concentration. ## Why Work Here - **Culture**: "High level of rigor, precision, and professionalism" – employees are stakeholders who own the impact of their work. Fast-paced startup environment where initiative is rewarded. [vast.ai/about](https://vast.ai/about) - **Work location**: On-site in Los Angeles (Westwood) or San Francisco (SOMA). Not remote – all current openings require on-site presence. - **Perks**: Comprehensive health/dental/vision insurance, 401(k) with company match, meaningful early-stage equity, onsite meals and snacks, close collaboration with founders and tech leaders. [vast.ai/jobs](https://vast.ai/jobs) - **Interview process**: Screened by technical team; stages include a 15-min screening, 45-min deep dive, 1-hour LLM-assisted coding assessment, and 2-hour on-site meet-and-greet. Aim to complete in about one week. [vast.ai/jobs](https://vast.ai/jobs) - **Salary ranges**: For roles like Senior Infrastructure Engineer and GPU Systems Engineer, posted salary $120K–$180K (may vary by role). [vast.ai/jobs](https://vast.ai/jobs) - **Engineering culture**: Reports directly to CEO/founder Jake Cannell, a prolific writer on AI and compute scaling theory. Emphasis on systems engineering, GPU optimization, and cutting-edge research. ## Sources 1. [vast.ai/about](https://vast.ai/about) 2. [vast.ai](https://vast.ai/) 3. [vast.ai/jobs](https://vast.ai/jobs) 4. [jobsbyculture.com](https://jobsbyculture.com/blog/working-at-vast-2026) 5. [vast.ai/jobs/apply/systems-gpu-research-engineer](https://vast.ai/jobs/apply/systems-gpu-research-engineer) ## Other roles at Vast.ai - [Head of Sales](https://feeny.ai/job/head-of-sales-vast-ai-los-angeles-yrkth8v89kkh) — Los Angeles, CA - [Developer Relations Engineer](https://feeny.ai/job/developer-relations-engineer-vast-ai-san-francisco-19dx882ce5x1) — San Francisco, CA - [Controller](https://feeny.ai/job/controller-vast-ai-los-angeles-0d4jvgnj6v48) — Los Angeles, CA - [AI, HPC & GPU Infrastructure Support Engineer](https://feeny.ai/job/ai-hpc-gpu-infrastructure-support-engineer-vast-ai-los-angeles-40e39w8ek9p1) — Los Angeles, CA - [Systems Operations Support Engineer — Linux](https://feeny.ai/job/systems-operations-support-engineer-linux-vast-ai-los-angeles-v3hjj8k7jj4v) — Los Angeles, CA - [Technical Product Manager, Infrastructure](https://feeny.ai/job/technical-product-manager-infrastructure-vast-ai-los-angeles-b0h8y0t8d4fy) — Los Angeles, CA - [Director of Engineering](https://feeny.ai/job/director-of-engineering-vast-ai-san-francisco-q3avgemppyhv) — San Francisco, CA - [Senior Infrastructure Engineer](https://feeny.ai/job/senior-infrastructure-engineer-vast-ai-los-angeles-ejanf64s8d34) — Los Angeles, CA - [Security Engineer](https://feeny.ai/job/security-engineer-vast-ai-los-angeles-5p7ef9m994ce) — Los Angeles, CA - [GPU Systems Engineer – HPC / Parallel Computing](https://feeny.ai/job/gpu-systems-engineer-hpc-parallel-computing-vast-ai-san-francisco-e7dn3x08y2sq) — San Francisco, CA