--- title: 'Member of Technical Staff, Supercomputing Platform & Infrastructure at Magic' canonical: 'https://feeny.ai/job/member-of-technical-staff-supercomputing-platform-infrastructure-magic-san-5h1ap46187xj' type: 'job' last_seen: '2026-09-09' --- # Member of Technical Staff, Supercomputing Platform & Infrastructure at Magic - **Company:** Magic - **Location:** San Francisco, CA - **Compensation:** $200k–$550k - **Employment:** full-time - **Work type:** onsite - **Posted:** 2024-01-25 - **Last confirmed live:** 2026-09-09 - **Apply:** https://jobs.ashbyhq.com/magic.dev/45d25f90-3be6-417d-810f-0d95f7704961 ## Job description Magic’s mission is to build safe AGI that accelerates humanity’s progress on the world’s most important problems. We believe the most promising path to safe AGI lies in automating research and code generation to improve models and solve alignment more reliably than humans can alone. Our approach combines frontier-scale pre-training, domain-specific RL, ultra-long context, and inference-time compute to achieve this goal. ## ABOUT THE ROLE As an engineer on the Supercomputing Platform & Infrastructure team, you will design, build, and operate the large-scale GPU infrastructure that powers Magic’s model training and inference workloads. A core part of this role is building and maintaining our infrastructure using Terraform-driven infrastructure-as-code practices, ensuring reproducibility, reliability, and operational clarity across clusters spanning thousands of GPUs. Magic’s long-context models create sustained pressure on compute, networking, and storage systems. Long-running distributed jobs, high-throughput data movement, and strict availability requirements demand infrastructure that is automated, observable, and resilient by design. You will own the systems and IaC foundations that make this possible, including the Kubernetes (K8s) environments that coordinate workloads across our GPU infrastructure. This role can evolve into broader ownership of supercomputing platform architecture, shaping how Magic scales GPU clusters and infrastructure reliability as model workloads grow. ## WHAT YOU’LL WORK ON - Design and operate large-scale GPU clusters for training and inference - Build and maintain infrastructure using Terraform across cloud and hybrid environments - Deploy, operate, and optimize K8s clusters used to schedule and manage AI workloads - Develop modular, scalable IaC patterns for compute, networking, and storage provisioning - Improve deployment reproducibility, environment consistency, and operational safety - Optimize networking and storage systems for high-throughput AI workloads - Automate fault detection and recovery across distributed clusters - Debug complex cross-layer issues spanning hardware, drivers, networking, storage, OS, and cloud - Improve observability, monitoring, and reliability of core platform systems ## WHAT WE’RE LOOKING FOR - Strong software engineering skills with experience building production infra systems - Deep, hands-on experience with Terraform, including module design, state management, environment isolation, and large-scale deployments - Experience operating production GPU infrastructure or high-performance distributed systems - Strong understanding of networking and storage systems - Experience with major cloud platforms (GCP, AWS, Azure, OCI, etc.) - Track record of owning production-critical infrastructure end-to-end ## OUR CULTURE - Integrity. Words and actions should be aligned - Hands-on. At Magic, everyone is building - Teamwork. We move as one team, not N individuals - Focus. Safely deploy AGI. Everything else is noise - Quality. Magic should feel like magic Magic strives to be the place where high-potential individuals can do their best work. We value quick learning and grit just as much as skill and experience. COMPENSATION, BENEFITS, AND PERKS (US): - Annual salary range between $200K - $550K depending on experience - Equity is a significant part of total compensation, in addition to salary - 401(k) plan with 6% salary matching - Generous health, dental and vision insurance for you and your dependents - Unlimited paid time off - Visa sponsorship and relocation stipend to bring you to SF, if possible - A small, fast-paced, highly focused team ## About Magic ## Company Overview - **One-liner**: Magic is building frontier-scale code models to automate software engineering and research, with the goal of achieving safe AGI. - **Entity Type**: Private (Series B) - **Headquarters**: San Francisco, California, United States - **Founded**: 2022 - **Founders**: Eric Steinberger, Sebastian De Ro ## Core Business - **Primary industry/industries**: Artificial Intelligence, Software Development - **Target customers**: B2B, Enterprise (via code generation models and APIs) - **Mission or purpose statement**: To safely deploy AGI by automating AI research and code generation, aligning models more reliably than humans can alone. ## Products & Services - **LTM (Long Term Memory) Models**: A family of frontier-scale LLMs with ultra-long context windows. LTM-1 had a 5,000,000 token context window, and LTM-2-Mini has a 100,000,000 token context window. These models are designed to understand and generate code across entire codebases. - **AI Software Engineer**: An internal and product-facing system aimed at autonomously performing software engineering tasks, from code generation to complex research. ## Market Standing - **Valuation/Market Cap**: Not publicly disclosed. - **Key Metric**: Total Funding of **$515 million** (across Seed, Series A, and Series B rounds). - **Notable Investors/Partners**: Nat Friedman, Daniel Gross, CapitalG (Alphabet’s independent growth fund), Elad Gil, Sequoia Capital, Jane Street, Eric Schmidt. - **Growth Signals**: - Rapid headcount growth: 79 employees in 2025 (up ~100% YoY from ~55 in 2024). - Partnership with Google Cloud (announced in conjunction with 100M token context windows and new funding). - Operates in 12 countries (including Austria, Switzerland). ## Competitive Advantages - **Ultra-Long Context Windows**: Their proprietary LTM architecture (100M token context) is a significant moat, enabling the model to process entire code repositories in a single pass. - **Infrastructure Scale**: Possesses "thousands of GB200s" (NVIDIA GPUs) for frontier-scale pre-training, a massive capital advantage. - **Small, High-Caliber Team**: A deliberate strategy of maintaining a small team of engineers and researchers to focus on a "short list of fundamental research problems." - **Direct Path to AGI Thesis**: Uniquely focused on using code generation as both the product and the core research path to AGI, differentiating from broader AI labs. ## Strategic Focus - **Automate AI Research**: Using code models to autonomously design, develop, and evaluate new alignment techniques. - **Safety & Alignment**: Actively publishing an "AGI Readiness Policy" and engaging in safety research, viewing alignment as a core technical problem to be solved by AI. - **Post-Training & Applied Team**: Recently announced an Applied Team focused on post-training to explore practical applications of their unreleased LTM2 models. - **Security Standards**: Advocating for and implementing cyber, physical, and information security standards comparable to the defense and nuclear industry. ## Why Work Here - **Mission-Driven**: A small team with a shared belief in the positive potential of responsibly deployed AGI. The work is directly tied to solving fundamental research problems. - **Culture & Values**: Emphasizes integrity, hands-on work ("everyone is building"), teamwork, focus, and quality. - **Compensation & Benefits**: - Competitive salary with significant equity compensation. - Unlimited paid time off. - 401K with 6% salary matching. - Health, dental, and vision insurance for employees and dependents. - Access to mental health, financial wellbeing, and telehealth programs. - Office catering (chef-prepared lunch and dinner). - **Work Environment**: In-person team meetings and office work are emphasized. The hiring process includes a take-home assessment and an in-person team meeting. - **Hiring Process**: Designed to be fair and focused on potential, not just resume credentials. Includes an application review, interview with a senior team member, take-home assessment (~5-8 hours), in-person team meeting, and reference check. ## Sources 1. [magic.dev](https://magic.dev/) 2. [magic.dev/careers](https://magic.dev/careers) 3. [magic.dev/safety](https://magic.dev/safety) 4. [linkedin.com/company/magicailabs](https://www.linkedin.com/company/magicailabs) 5. [theorg.com/org/magic-dev](https://theorg.com/org/magic-dev) ## Other roles at Magic - [Head of IT](https://feeny.ai/job/head-of-it-magic-san-francisco-sdsbb3qvy42e) — San Francisco, CA - [Member of Technical Staff, Security Engineer](https://feeny.ai/job/member-of-technical-staff-security-engineer-magic-san-francisco-rg197r1wt9nq) — San Francisco, CA - [Member of Technical Staff, Pre-training Systems](https://feeny.ai/job/member-of-technical-staff-pre-training-systems-magic-san-francisco-pkrp1989x0pb) — San Francisco, CA - [Member of Technical Staff, Inference & RL Systems](https://feeny.ai/job/member-of-technical-staff-inference-rl-systems-magic-san-francisco-m3ewreg7g147) — San Francisco, CA - [Member of Technical Staff, RL Research & Environments](https://feeny.ai/job/member-of-technical-staff-rl-research-environments-magic-san-francisco-b9dv41yxjrxm) — San Francisco, CA - [Principal Security Engineer/Head of Security](https://feeny.ai/job/principal-security-engineer-head-of-security-magic-san-francisco-0vvn5g9h1xxw) — San Francisco, CA - [](https://feeny.ai/job/insert-job-you-excel-at-magic-san-francisco-7n1rsssg65qn) — San Francisco, CA - [Research Engineer](https://feeny.ai/job/research-engineer-magic-san-francisco-kbt0xr1zv77g) — San Francisco, CA - [Member of Technical Staff, Kernels](https://feeny.ai/job/member-of-technical-staff-kernels-magic-san-francisco-j9gvfjz3n9ch) — San Francisco, CA