--- title: 'Staff Software Engineer - AI Compute, Together Cloud at Together AI' canonical: 'https://feeny.ai/job/staff-software-engineer-ai-compute-together-cloud-together-ai-san-francisco-egf3h6r021mj' type: 'job' last_seen: '2026-09-25' --- # Staff Software Engineer - AI Compute, Together Cloud at Together AI - **Company:** Together AI - **Location:** San Francisco, CA - **Posted:** 2026-09-23 - **Last confirmed live:** 2026-09-25 - **Apply:** https://job-boards.greenhouse.io/togetherai/jobs/5224090007 ## Job description ## About the Role Together AI is building the AI Native Cloud, an end-to-end platform for the full generative AI lifecycle, combining the fastest LLM inference engine with state-of-the-art GPU cloud infrastructure. The Together Cloud team builds the [Together GPU Clusters](https://www.together.ai/gpu-clusters) flagship IaaS product that provides high-performance, AI-ready GPU clusters through a self-serve cloud console, along with the virtualized infrastructure layer powering Together's inference, RL, and fine-tuning products. As a Staff Software Engineer focusing on AI Compute in the Together Cloud org, you will set technical direction for and build major components of the next generation AI cloud platform – a highly available, global cloud infrastructure with cutting-edge virtualization of the latest ML hardware: GB300s/VRs, BlueField DPUs, InfiniBand and dual/quad-plane RoCEv2 fabrics. That virtualized computing platform powers our own SaaS products – inference, RL, and fine-tuning – and serves external cloud customers through self-serve offerings such as on-demand/reserved Kubernetes/Slurm clusters, across dozens of data centers and hundreds of thousands of GPUs. This is an architect-and-build role. Fully automated bootstrapping of GPU data centers, high-performance virtualization of GPU compute and DC networking without compromising isolation or portability, and fault-tolerant decentralized control planes — you'll set the architecture for these across our global and in-DC services, and be a key owner of the hardest parts, in the code as well as the design. Your designs will span the IaaS layer of a greenfield Vera Rubin data center up to the global management plane that schedules capacity across all of them. At this level the job is as much leverage as code: the standards you set and the engineers you grow decide how fast the rest of Together Cloud ships. ## Responsibilities - Own the GPU and network virtualization stack: the hypervisor, kernel, and SDN work that keeps GPU compute and DC networking high-performance, portable, and strongly isolated across heterogeneous hardware. - Own the in-DC IaaS layer: architect and roadmap the services, Kubernetes operators, and libraries that provision and manage compute, storage, and networks in our data centers — VMs, parallel filesystems, VPCs, and InfiniBand partitions; lead its build-out for a new Vera Rubin data center with thousands of GPUs, from hardware bring-up to customer-facing API. - Design the GPU scheduling and global management plane: the distributed control plane behind on-demand and reserved clusters across dozens of data centers, including the systems that scale per-cluster limits and automate the onboarding of new capacity. - Architect monitoring and automated remediation for fault tolerance: the strategy for automated detection, isolation, and recovery of failed nodes that keeps distributed pretraining and large-scale inference running through hardware failures. - Set technical direction across teams: lead design reviews, resolve cross-cutting architectural disagreements, unblock cross-team dependencies and integration risks, and define the standards other engineers build against — measured in cluster reliability, time-to-first-GPU on new capacity, and quality at scale. - Grow the team: mentor senior and junior engineers, deepen the team's expertise in virtualization, DC networking, and GPU infrastructure, and help raise the hiring bar for Together Cloud. - Set the engineering bar: create the testing frameworks, tools, and developer documentation that make our systems robust and usable by other teams, and shape the core, open-source Together AI platform. To be successful you'll need to be deeply technical and an excellent communicator — expert software development fundamentals, deep systems knowledge and troubleshooting instincts, and the leadership and diplomacy skills to align teams that don't report to you. Much of this work starts ambiguous, and we expect you to define the scope yourself and drive it to production. ## Requirements - 7+ years of professional software development experience, with expert-level proficiency in at least one backend language (Golang desired), writing high-performance, well-tested, production-quality code. - Track record of owning the architecture of large distributed systems from blank page to production at scale, including the judgment calls that could not be reversed cheaply. - Deep experience building and operating globally distributed, high-performance microservice architectures across one or more cloud providers (AWS, Azure, GCP). - Expert systems knowledge across compute, networking, and storage — including concurrency, memory management, performant I/O, and scale at a global level. - Demonstrated technical leadership beyond your own commits: mentoring senior engineers, leading design reviews, and driving alignment across teams that do not report to you. - Excellent communication and diplomacy skills — able to write design docs that settle arguments, and to work effectively with technical and non-technical stakeholders. - Experience building and operating reliable, customer-facing production systems at scale, and owning the infrastructure automation (Terraform, Ansible), observability (Prometheus, Grafana), and CI/CD (GitHub Actions, ArgoCD) that keep them healthy. ## Preferred Qualifications (not must haves) - Deep Kubernetes internals experience, such as implementing non-trivial Kubernetes operators, device/storage/network plugins, custom schedulers, or patches to Kubernetes itself - Deep experience with VMs/hypervisors, such as QEMU/KVM, cloud-hypervisor, VFIO, virtio, PCIE passthrough, Kubevirt, SR-IOV - Deep experience with DC networking tech + solutions, such as VLAN, VXLAN, VPN, VPC, OVS/OVN - Experience with Cluster API or similar - Experience working on high-performance compute, networking, and/or storage - Experience virtualizing GPUs and/or InfiniBand - Experience building IaaS or PaaS systems at scale - Experience with DPUs/SmartNICs - GPU programming, NCCL, CUDA knowledge ## About Together AI Together AI is a research-driven artificial intelligence company. We believe open and transparent AI systems will drive innovation and create the best outcomes for society, and together we are on a mission to significantly lower the cost of modern AI systems by co-designing software, hardware, algorithms, and models. We have contributed to leading open-source research, models, and datasets to advance the frontier of AI, and our team has been behind technological advancement such as FlashAttention, Hyena, FlexGen, and RedPajama. We invite you to join a passionate group of researchers in our journey in building the next generation AI infrastructure. ## Compensation We offer competitive compensation, startup equity, health insurance, and other benefits, as well as flexibility in terms of remote work. The US base salary range for this full-time position is: $260,000 - $300,000 + equity + benefits. Our salary ranges are determined by location, level and role. Individual compensation will be determined by experience, skills, and job-related knowledge. ## Equal Opportunity Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more. Please see our privacy policy at https://www.together.ai/privacy ## About Together AI ## Company Overview - **One-liner**: Together AI is the AI Native Cloud, a full-stack platform for production AI that provides inference, fine-tuning, and GPU compute powered by cutting-edge systems research. - **Entity Type**: Private (Series C) - **Headquarters**: San Francisco, California, United States - **Founded**: 2022 - **Founders**: Vipul Ved Prakash (CEO), Ce Zhang (CTO), Tri Dao (Chief Scientist), Chris Ré, Percy Liang ## Core Business - **Primary industry**: AI Infrastructure / Cloud Computing - **Target customers**: B2B – AI-native startups (e.g., Cursor, Decagon, Eleven Labs), enterprise SaaS (Salesforce, Zoom, Zomato), and AI researchers. - **Mission**: “Significantly lower the cost of modern AI systems by co-designing software, hardware, algorithms, and models” — with a commitment to open and transparent AI development. ## Products & Services - **Serverless Inference**: Fast on-demand inference for open-source models with no infrastructure management. - **Batch Inference**: Asynchronous processing of massive workloads (up to 30B tokens per model) at reduced cost. - **Dedicated Model Inference**: Deploy models on dedicated GPU infrastructure for speed and control. - **Dedicated Container Inference**: GPU infrastructure optimized for generative media (video, audio, image). - **Fine-Tuning**: Fine-tune open-source models using latest research techniques (improves accuracy, reduces hallucinations). - **Accelerated Compute**: Self-serve instant clusters to thousands of GPUs, optimized with Together Kernel Collection. - **Model Shaping**: Tools to accelerate inference, model shaping, and pre-training. - **Managed Storage**: High-performance object storage and parallel filesystems with zero egress fees. - **Sandbox**: Fast, secure code sandboxes for AI app/agent development environments. ## Market Standing - **Valuation/Market Cap**: Not publicly disclosed; recent Series C reportedly doubled valuation (amount undisclosed) [cbinsights.com](https://www.cbinsights.com/company/together-4) - **Key Metric**: **Total Funding** – $1.334B (including $800M Series C announced in mid-2026) [cbinsights.com](https://www.cbinsights.com/company/together-4) - **Notable Investors/Partners**: Kleiner Perkins, NVIDIA, Salesforce Ventures, Prosperity7 Ventures (Aramco), and 29+ others across rounds. - **Growth Signals**: - Headcount grew 91% YoY to ~280 employees [linkedin.com](https://www.linkedin.com/company/togethercomputer) - Acquired CodeSandbox (Dec 2024) and Refuel.AI (May 2025) - Customers include Cursor, Decagon, Eleven Labs, AI21, Hedra, Cartesia, Salesforce, Zoom, Zomato - Published nine papers at ICML 2026; open-sourced ParallelKernelBench ## Competitive Advantages - **Full-stack AI platform** covering inference, fine-tuning, compute, storage, and sandboxes — all integrated. - **World-class research team** co-designing algorithms, kernels, and hardware (e.g., Tri Dao, co-inventor of FlashAttention). - **Open-source commitment** and contributions (models, datasets, research) build developer trust and ecosystem lock-in. - **Cost performance**: Claims 2x faster inference and 60% lower cost vs. alternatives through research-driven optimizations. ## Strategic Focus - Accelerate the shift to open-source AI by making it economically viable at scale. - Continue investing in systems research (GPU programming, kernel optimization, model shaping) to improve performance and cost. - Expand enterprise adoption through dedicated deployments and partnerships. - Grow the platform’s capabilities (e.g., managed storage, sandboxes) to reduce friction for AI-native builders. ## Why Work Here - **Culture**: Values include “Do more with less,” “Open and responsible development,” “Empower innovation,” and “Optimizers” — a research-driven, mission-oriented environment. - **Remote/Hybrid/Office**: Primarily office-based in San Francisco (251 Rhode Island Street) with employees across 16 countries. Office perks include lunch, dinner, snacks, parking/transit stipends, and relocation support. - **Benefits**: Competitive salary and equity, 401K matching, health/dental/vision insurance, flexible PTO, generous parental leave, life and disability protection. - **Engineering Culture**: Deep focus on systems research and open-source; engineers work alongside leading AI researchers (e.g., Tri Dao, Ce Zhang). Talent sourced from Apple, Google, Amazon, NVIDIA, Stanford AI Lab. - **Career Growth**: 53 active job postings (as of mid-2026); roles span ML engineering, platform engineering, research, and go-to-market. High growth trajectory offers rapid advancement. ## Sources 1. [together.ai/about-us](https://www.together.ai/about-us) 2. [together.ai/](https://www.together.ai/) 3. [together.ai/careers](https://www.together.ai/careers) 4. [linkedin.com/company/togethercomputer](https://www.linkedin.com/company/togethercomputer) 5. [cbinsights.com/company/together-4](https://www.cbinsights.com/company/together-4) ## Other roles at Together AI - [Senior Recruiter, GTM & Business](https://feeny.ai/job/senior-recruiter-gtm-business-together-ai-san-francisco-wp8ngndvhfdy) — San Francisco, CA - [Research Intern, Frontier Agents (Summer 2027)](https://feeny.ai/job/research-intern-frontier-agents-summer-2027-together-ai-san-francisco-vrgyvzdr6fkn) — San Francisco, CA - [Research Intern, Frontier Agents (Winter 2027)](https://feeny.ai/job/research-intern-frontier-agents-winter-2027-together-ai-san-francisco-kbpsevmdj074) — San Francisco, CA - [Systems Research Engineer Intern - GPU Programming (Winter 2027)](https://feeny.ai/job/systems-research-engineer-intern-gpu-programming-winter-2027-together-ai-san-n6158j2xrf4w) — San Francisco, CA - [Systems Research Engineer Intern - GPU Programming (Summer 2027)](https://feeny.ai/job/systems-research-engineer-intern-gpu-programming-summer-2027-together-ai-san-5azr883avc6d) — San Francisco, CA - [Software Engineer, New Grad (2027)](https://feeny.ai/job/software-engineer-new-grad-2027-together-ai-san-francisco-dkq06899jtp8) — San Francisco, CA - [Software Engineer Intern (Winter 2027)](https://feeny.ai/job/software-engineer-intern-winter-2027-together-ai-san-francisco-4gt34sy5whsw) — San Francisco, CA - [Software Engineer Intern (Summer 2027)](https://feeny.ai/job/software-engineer-intern-summer-2027-together-ai-san-francisco-md84jhmgfnwk) — San Francisco, CA - [Software Development In Test Intern (Summer 2027)](https://feeny.ai/job/software-development-in-test-intern-summer-2027-together-ai-san-francisco-myx5b6h3ta75) — San Francisco, CA - [Research Intern, Model Shaping (Winter 2027)](https://feeny.ai/job/research-intern-model-shaping-winter-2027-together-ai-san-francisco-b4j4s6ttt9p3) — San Francisco, CA