--- title: 'Forward Deployed Engineer - SRE at Andromeda Cluster' canonical: 'https://feeny.ai/job/forward-deployed-engineer-sre-andromeda-cluster-north-gq659csdwgjm' type: 'job' last_seen: '2026-09-03' --- # Forward Deployed Engineer - SRE at Andromeda Cluster - **Company:** Andromeda Cluster - **Location:** North, United States / San Francisco, CA - **Employment:** full-time - **Work type:** remote - **Posted:** 2026-07-31 - **Last confirmed live:** 2026-09-03 - **Apply:** https://jobs.ashbyhq.com/andromeda/9a8cac0b-aec6-4df1-a03e-2cac4b3f9813 ## Job description ## FORWARD DEPLOYED ENGINEER - SRE ## LOCATION: NORTH AMERICA REMOTE/SF-HYBRID · FULL-TIME ## ABOUT ANDROMEDA Andromeda is a market and infrastructure platform to buy, sell, and operate compute. We believe demand for compute will grow exponentially. So fast that a handful of vertically integrated providers won't be able to scale across operations, capital, supply chains, and politics to serve it. The result is a massive wave of fragmentation, with AI factories of every shape and size coming to market to fill this demand. Our job is to enable all of that fragmented compute to flow through one platform, delivering reliable capacity to model builders, research labs, and inference providers when they need it. We believe every spare electron should be made productive for AI and we're building the platform that makes that possible. We sit at the center of three forces: - Companies that need reliable, high-performance compute fast - A fragmented global supply of GPUs across hyperscalers, neoclouds, and independent data centers - Capital, risk, and operational complexity that most teams are not equipped to manage When we succeed, trillions of dollars of compute will flow through Andromeda. Builders get capacity when they need it. Providers get a reliable way to monetize, operate, and finance infrastructure at scale. Capital gets an easy way to deploy, hedge, and underwrite. In five years, Andromeda won't just participate in the AI infrastructure market. We will shape it. ## The Role This is not a generalist SRE role, and it is not a support role. You will embed directly with the teams running large-scale training and inference on our clusters. You are responsible for onboarding them, tuning their jobs, and debugging their failures alongside them, while owning the infrastructure and automation that makes those clusters reliable in the first place. Forward deployed means you spend real time inside customer environments: reading their training code, sitting in their Slack channels, watching their runs, and shipping fixes that land in our platform. When a multi-hundred-GPU run stalls, you are the person who figures out whether it's the fabric, the driver, the scheduler, or their dataloader, then you make sure it can't happen the same way twice. We're looking for engineers who have personally run GPU clusters in production, understand the failure modes of distributed training, and can reason about performance from network fabric → kernel → framework. Equally important: you can explain what you found to someone else's engineering team without condescension, and turn that conversation into a product improvement. ## What You’ll Do - Serve as the primary technical point of contact for teams running large-scale training and inference workloads. Own onboarding end to end; environment setup, orchestration choice (Slurm, Kubernetes, or direct SSH), storage layout, first successful run at scale. You will continue to stay engaged as their workloads grow. - Work inside customer environments to diagnose real failures: NCCL timeouts, stragglers, checkpoint I/O stalls, degraded links, OOM patterns, container and driver mismatches. Read their code when you need to. Reproduce, isolate, fix, and write it down. - Profile and improve distributed training performance on live workloads. Improving MFU, cutting idle GPU time, and reducing time-to-first-successful-run for new deployments. - Own reliability outcomes for the accounts you're deployed on. - Ensure the health and performance of high-speed interconnects (InfiniBand, RoCE, NVLink) that underpin distributed training. Diagnose and resolve fabric-level issues that degrade collective operations. - Build deep visibility into GPU utilization, memory pressure, interconnect throughput, job performance, and hardware health. - Turn every repeated deployment problem into automation: cluster provisioning, GPU health checks and burn-in, preflight validation, self-healing, firmware/driver lifecycle management, and reusable reference configurations for common training and serving stacks. - Lead incident response for complex, multi-layer failures spanning hardware, networking, orchestration, and ML frameworks. Own the customer-facing communication during the incident and the blameless postmortem and systemic fix after it. - You will see our rough edges before anyone else does. Bring that signal back to influence the roadmap, file the hard bugs, and build the missing pieces yourself when that's the fastest path. ## What We’re Looking For - Hands-on experience operating GPU clusters in production (NVIDIA A100/H100/H200/B200 or equivalent). You understand GPU memory hierarchies, ECC behavior, thermal throttling, and hardware failure modes from direct experience. - Production experience with InfiniBand, RoCE, or NVLink fabrics in the context of distributed training. You can diagnose why an all-reduce is slow, identify a degraded link in a fat-tree topology, and reason about congestion control at scale. - Working knowledge of how large training and inference jobs actually run. You don't need to design models, but you need to understand what's happening at the systems level when a large run stalls. - Expert-level Linux experience: kernel tuning, driver management (NVIDIA drivers, CUDA toolkit), cgroup/namespace internals, container runtimes, and performance profiling at the syscall and hardware level. - Strong experience running Kubernetes in production with GPU workloads. Experience with device plugins, topology-aware scheduling, multi-cluster, custom operators. Experience with Slurm or other HPC schedulers is equally valued. - Strong engineering skills in Python, Go, or Bash. You build production-grade tools and services, not just scripts. - Infrastructure-as-Code proficiency (Terraform, Helm, Ansible, or equivalent). - Hands-on experience building monitoring and alerting for GPU-specific telemetry (DCGM, nvidia-smi, fabric manager metrics) integrated into actionable dashboards. - You can go deep on architecture with a customer's infra team and clearly articulate tradeoffs to their leadership. You're comfortable being the only one in the room who knows the answer, and equally comfortable saying you don't yet. - Proven track record leading incident response for complex distributed systems. Strong Candidates May Have - Experience with high-performance parallel file systems (VAST, WEKA, Lustre, GPFS) and the checkpoint I/O and data-loading bottlenecks that come with large training runs. - Time spent embedded with external engineering teams, i.e solutions architecture, professional services, deployed SRE, or technical account ownership at an infrastructure company. - Experience operating production inference. Hands on experience with autoscaling, batching, KV cache behavior, cold starts, and multi-tenant GPU sharing. - Contributions to relevant OSS projects, or benchmarks, postmortems, and deep-dives you've published. - Experience working across heterogeneous providers and regions rather than a single hyperscaler. What Success Looks Like By the end of your first year: - You know each of your accounts' actual technical goals including what they're training, what their scaling roadmap looks like over the next two quarters, what their real constraints are (budget, deadline, headcount, data), and you've written that down somewhere the rest of us can read it. - You are the person your accounts' engineers message first, before they file a ticket, because you've earned it. - Their reliability and throughput numbers are visibly better than at onboarding, and you can point to the specific changes that did it. - Recurring problems you found in the field exist as automation, preflight checks, or documentation, not as tribal knowledge in your head. - You've advocated internally for at least one roadmap change on behalf of a strategic customer, and it shipped. ## Why You’ll Love It Here - High-growth environment: Get in early at a company at the center of the AI infrastructure boom - Ownership: First FDE for the solutions engineering team, you’ll get to build this function from the ground up - Competitive compensation: + meaningful equity - Comprehensive benefits: for you and your dependents, including healthcare, dental, and vision coverage, 401(k), and unlimited PTO Andromeda Cluster is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. We do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status. ## About Andromeda Cluster ## Company Overview - **One-liner**: Andromeda is building the liquidity layer for global AI compute, connecting AI teams with high-performance GPU capacity through a trusted marketplace that standardizes and matches supply with demand. - **Entity Type**: Private (Funding not disclosed) - **Headquarters**: San Francisco, California, United States - **Founded**: 2023 (inferred from public launch in March 2026 after three years of operation) - **Founders**: Nat Friedman, Daniel Gross ## Core Business - **Primary industry**: AI infrastructure and compute marketplace - **Target customers**: B2B – AI labs, data centers, cloud providers, and any organization building large-scale AI workloads - **Mission/purpose**: To make high-performance compute available to every team building at the frontier, so breakthroughs are driven by ideas, not hardware access. ## Products & Services - **[Andromeda Marketplace (Buy Side)]**: A single platform where AI teams can define their workload (GPU type, quantity, region, timeline) and access real-time pricing across 100+ providers. The platform benchmarks, prices, routes, and deploys compute – reserved, on-demand, or spot – with standardized SLAs, contracts, and full observability. - **[Andromeda Marketplace (Sell Side)]**: Enables compute providers (telcos, crypto data centers, MSPs, sovereign clouds) to onboard their infrastructure, get it certified against enterprise-grade benchmarks, receive qualified demand, and start earning revenue with standardised terms and one invoice. - **Managed Clusters**: Began as a single multi-thousand GPU cluster for early-stage startups; now operates a constellation of clusters across many providers, supporting both training and inference workloads. ## Market Standing - **Valuation/Market Cap**: Not publicly available - **Key metric**: 1 billion+ GPU-hours of compute managed; over 1000 transactions completed; 100+ compute providers on the platform - **Notable Investors/Partners**: Founders are well-known tech figures (Nat Friedman, former GitHub CEO; Daniel Gross, investor/entrepreneur). Specific investors not disclosed. Works with leading AI labs, data centers, and cloud providers. - **Growth Signals**: Public launch in March 2026; 14 employees with +15.8% monthly headcount growth; 10 active job postings; +11.1% monthly job posting trend; global remote hiring with San Francisco office. ## Competitive Advantages - **Standardized certification**: Every cluster on the network must meet uniform benchmarks for GPU, CPU, storage, network fabric, throughput, and security – removing trust barriers. - **Matching engine**: Prices, routes, and deploys the right infrastructure to the right workload automatically, with options for reserved, on-demand, and spot. - **Full-stack market infrastructure**: Sourcing, certification, contracting, matching, operations, and observability – all under one relationship. - **Founder credibility**: Built by Nat Friedman and Daniel Gross, giving the platform instant legitimacy and access to early AI labs. ## Strategic Focus - Build the liquidity layer for global AI compute: expand the network of providers and buyers, deepen the orchestration layer, and enable compute to flow as freely as other commodities. - Grow the team across engineering, sales, partnerships, and operations – with a strong emphasis on AI infrastructure roles. ## Why Work Here - **Culture & environment**: Small, high-impact team (14 people) working on one of the most critical bottlenecks in AI. Flat structure with direct access to leadership. - **Remote/hybrid policy**: Global remote with a San Francisco office. Many roles listed as "Global Remote / San Francisco, CA". - **Engineering culture**: Deep technical challenges – GPU orchestration, distributed systems, observability, kernel-level performance tuning. Tech stack includes Kubernetes, Slurm, PyTorch, CUDA, Weka, Prometheus, Grafana, and more. - **Roles available**: Compute Trader, Site Reliability Engineer, Software Engineer, Head of Partnerships, Solutions Engineer, Strategic Compute Finance Lead, and more – all focused on AI infrastructure. ## Sources 1. [Andromeda Homepage](https://www.andromeda.ai/) 2. [Andromeda Careers Page](https://andromeda.ai/careers) 3. [LinkedIn Company Page](https://www.linkedin.com/company/andromeda-cluster) 4. [Built In Profile](https://builtin.com/company/andromeda-andromeda-ai) 5. [Andromeda Insights Page](https://andromeda.ai/insights) ## Other roles at Andromeda Cluster - [Developer Relations Engineer](https://feeny.ai/job/developer-relations-engineer-andromeda-cluster-north-qq87xesk8z9a) — North, United States / San Francisco, CA - [Technical Recruiter](https://feeny.ai/job/technical-recruiter-andromeda-cluster-north-sr46j2svwq6d) — North, United States / San Francisco, CA - [Compute Procurement Lead](https://feeny.ai/job/compute-procurement-lead-andromeda-cluster-north-d47ra63emanm) — North, United States / San Francisco, CA - [Member of the Technical Staff - Platform](https://feeny.ai/job/member-of-the-technical-staff-platform-andromeda-cluster-north-5cvpct616zrg) — North, United States / San Francisco, CA - [Member of the Technical Staff - Systems](https://feeny.ai/job/member-of-the-technical-staff-systems-andromeda-cluster-north-v5rxr6ehp0n4) — North, United States / San Francisco, CA - [Member of the Technical Staff - Product](https://feeny.ai/job/member-of-the-technical-staff-product-andromeda-cluster-north-xr2e2vx5fvd1) — North, United States / San Francisco, CA - [Solutions Architect](https://feeny.ai/job/solutions-architect-andromeda-cluster-north-zr7mnep0c89s) — North, United States / San Francisco, CA - [Revenue Operations Lead](https://feeny.ai/job/revenue-operations-lead-andromeda-cluster-north-85ygfakh7akh) — North, United States / San Francisco, CA - [Technical Program Manager](https://feeny.ai/job/technical-program-manager-andromeda-cluster-global-san-francisco-ca-8nmssvsgsz33) — Global / San Francisco, CA - [Strategic Compute Finance Lead](https://feeny.ai/job/strategic-compute-finance-lead-andromeda-cluster-global-san-francisco-ca-a3qw3akn9rw7) — Global / San Francisco, CA