--- title: 'Network Engineer (Supercomputer Infrastructure) - Memphis at xAI' canonical: 'https://feeny.ai/job/network-engineer-supercomputer-infrastructure-memphis-xai-southaven-cv17szgyyymp' type: 'job' last_seen: '2026-09-07' --- # Network Engineer (Supercomputer Infrastructure) - Memphis at xAI - **Company:** xAI - **Location:** Southaven, MS / Memphis, TN - **Posted:** 2026-09-02 - **Last confirmed live:** 2026-09-07 - **Apply:** https://job-boards.greenhouse.io/xai/jobs/5229355007 ## Job description SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates. ## ABOUT THE ROLE: SpaceXAI is looking for an exceptional network engineer with experience in mission-critical, large-scale production environments to support the design, build-out, and operation of networks that power our AI supercomputer campuses. As a member of the Supercomputer Infrastructure / Network Engineering team, you will provide design and operational support for the fabrics used by GPU training and inference clusters, site operations, automation and controls, and facilities teams. The ideal candidate thrives in intense, high-flux environments, brings a strong sense of urgency balanced with operational excellence, communicates clearly, and demonstrates high technical acumen. ## RESPONSIBILITIES: - Design and implement highly available, low-latency, high-bandwidth networks, carefully balancing routing, congestion control, and redundancy technologies for AI training fabrics, inference front-ends, storage, and site/OT networks. - Design and maintain supercomputer data center and campus networks in accordance with company network standards. Collaborate with adjacent infrastructure, compute, storage, SiteOps, and enterprise teams. - Evaluate, procure, and deploy network hardware including data-center class switches, NICs, firewalls, optical multiplexers, and related appliances supporting 400G/800G and beyond. - Contribute to maturing network automation tooling; implement configuration analysis, linting, validation, and scalable deployment frameworks (GitOps / IaC). - Plan and coordinate network change windows with stakeholders to perform software updates, hardware refreshes, cluster expansions, and general maintenance (including evenings and weekends when required by compute schedules). - Troubleshoot and resolve network-related issues affecting cluster health and job performance; publish root cause analysis (RCA) documentation and host retrospectives. - Provide direct networking support during cluster bring-up, expansion, and production training/inference campaigns; serve as on-call or networking responsible engineer during operations. - Proactively tailor network monitoring and telemetry (fabric health, congestion, packet loss, NCCL/collective performance) so issues are detected before they impact training or inference. - Continuously create and update network documentation, including architecture overviews, design drawings, fiber/cable plant records, and operational procedures. - Collaborate with cross-functional teams to identify and resolve potential design issues, especially systemic or cascading failure modes and false redundancy in AI fabrics and site networks. - Perform job walks with customers, vendors, and contractors to gather requirements and produce implementation plans for new halls, rows, and campus interconnects. - Ensure networks are configured and maintained in compliance with industry and cybersecurity standards (e.g., ITAR, ISO, NIST), with particular attention to segmentation between compute fabrics, storage, OT/controls, and corporate networks. ## BASIC QUALIFICATIONS: - Bachelor’s degree in computer science, computer engineering, or other STEM discipline and 3+ years of professional network engineering experience; - OR 5+ years of professional network engineering experience in lieu of a degree. - Extensive hands-on experience designing, deploying, supporting, and troubleshooting Layer 2 and Layer 3 networks in latency-sensitive and/or industrial / data-center environments. - Functional experience with multiple network vendors in production or lab environments. - Experience with GitOps and Infrastructure as Code frameworks, both as a user and contributor. ## PREFERRED SKILLS AND EXPERIENCE: - Strong understanding of the OSI model and network standards. - Hands-on experience with Cisco, Arista, Juniper, and/or NVIDIA Spectrum-X data-center class switches. - Experience with RoCEv2 Ethernet AI/HPC fabrics; InfiniBand experience is a plus. - Working knowledge of AI training and inference traffic patterns and how they behave on the network (collectives, congestion, ECMP, adaptive routing). Familiarity with NCCL is a plus. - Experience with WDM and large-scale single-mode / multimode fiber plants, including OTDR and acceptance testing. - Experience with switch port security, network segmentation, QoS, multicast, and redundancy protocols. - Familiarity with network monitoring and Layer 1 test tools; experience building operational telemetry and dashboards. - Proficiency in scripting (Bash / PowerShell / Python) and automation frameworks (Terraform, Ansible, etc.). - Linux and Windows system administration experience professionally or from labs. - Industry-standard certifications such as CCNA or CCNP. - Experience supporting real-time systems, industrial control / OT networks, or high-reliability environments in data center, energy, aerospace, defense, or similar industries. - Excellent communication skills with internal and external customers, vendors, and management in both formal and informal settings. ## ADDITIONAL REQUIREMENTS: - Ability to pass applicable background checks for site access. - Ability to work in tight quarters; physical dexterity is necessary to perform job functions. - Availability for extended hours and/or weekends as the schedule varies with cluster build-out and operational needs; flexibility is required. - Ability to provide 24x7 on-call support in emergency situations and participate in an after-hours on-call rotation. - Willingness to travel (up to 20%) between supercomputer campuses and related sites. - Ability to lift 30 lbs. - Ability to work at heights. - Ability to drive (active valid driver’s license). SpaceXAI is an equal opportunity employer. For details on data processing, view our [Recruitment Privacy Notice](https://x.ai/legal/recruitment-privacy-notice). ## About xAI ## Company Overview - **One‑liner**: xAI builds artificial intelligence systems to accelerate human scientific discovery and “understand the universe,” with products spanning frontier reasoning models, real‑time voice, and generative media. - **Entity Type**: Private (last funding round: Series E, January 2026; acquired by SpaceX in February 2026) - **Headquarters**: Palo Alto, California, USA (additional offices in Seattle, WA; Memphis, TN; London, UK) - **Founded**: 2023 - **Founders**: Elon Musk ## Core Business - Primary industry: Artificial Intelligence (AI research, large language models, voice AI, generative media) - Target customers: B2B (developers via API), B2C (consumers via Grok on 𝕏, SuperGrok subscriptions), Enterprise & Government (xAI for Government) - Mission: “To understand the universe” – building AI that extends what humanity can know and do. ## Products & Services - **Grok (base models)** : A conversational AI assistant modeled after *The Hitchhiker’s Guide to the Galaxy*, available on the 𝕏 platform and via API. - **Grok 4 / Grok 4 Heavy**: The most intelligent reasoning models, featuring native tool use and real‑time search integration (SuperGrok and Premium+ tiers). - **Grok Voice Agent API / Voice Agent Builder**: Create custom voice agents in under two minutes (no‑code) – includes speech‑to‑text, text‑to‑speech, custom voice cloning, and multilingual support. - **Grok Imagine API**: State‑of‑the‑art video generation API (quality, cost, latency). - **Grok‑1.5 / Grok‑2 / Grok‑2 mini**: Earlier model releases with improved reasoning and 128K context length. - **PromptIDE**: An integrated development environment for prompt engineering and interpretability research. - **xAI for Government**: A suite of frontier AI products tailored for U.S. government customers. - **Colossus Supercomputer**: 200,000‑GPU cluster built in 122 days, used for training xAI models; now also leased to Anthropic (Colossus 1 agreement). All products are delivered as SaaS, API, and model weights (open‑source release of Grok‑1). ## Market Standing - **Valuation / Funding**: - Series B (May 2024): $6 B - Series C (Dec 2024): $6 B (led by A16Z, BlackRock, Fidelity, Kingdom Holdings, Lightspeed, MGX, Morgan Stanley, OIA, QIA, Sequoia, Valor, Vy Capital, others) - Series E (Jan 2026): $20 B - Acquisition by SpaceX announced Feb 2, 2026 (terms not disclosed). - **Key Metric**: Total disclosed funding ~$32 B; annual revenue not publicly available. - **Notable Investors/Partners**: A16Z, BlackRock, Fidelity, Sequoia Capital, Morgan Stanley, Kingdom Holdings, Lightspeed, MGX, Valor Equity Partners, Vy Capital, Anthropic (Colossus deal), U.S. Department of War (selected Dec 2025). - **Growth Signals**: - Rapid scaling from founding to 200K‑GPU cluster in 122 days. - Over 214 open roles listed on Greenhouse as of mid‑2026. - #1 on Big Bench Audio (Grok 4 audio performance). - Expanded into voice agents, video generation, and government verticals within 3 years. - Acquisition by SpaceX signals deep integration with space‑tech resources. ## Competitive Advantages - **Compute infrastructure**: Colossus supercomputer – one of the largest GPU clusters in existence, enabling faster training and larger models. - **Model performance**: Grok 4 claims “most intelligent model in the world” with top benchmarks (Big Bench Audio #1). - **Vertical integration**: Tight coupling with 𝕏 platform for data, distribution, and real‑time world knowledge. - **Speed of execution**: “Move quickly and fix things” culture; built Colossus in 122 days. - **Government contracts**: Early access to public‑sector AI markets (U.S. Department of War). - **Talent density**: Small, focused team of researchers and engineers; in‑person work accelerates iteration. ## Strategic Focus - **Continue advancing frontier reasoning** (Grok 4 Heavy, next‑gen models). - **Expand voice and multimodal APIs** for developers (Voice Agent Builder, Custom Voices Library). - **Deepen government / defense partnerships** (xAI for Government, Department of War). - **Scale infrastructure** post‑acquisition via SpaceX resources. - **Remain “open” where strategic** (open‑source release of Grok‑1 weights). ## Why Work Here - **Culture**: “Ambitious goals, fast execution” – rejects consensus in favor of first‑principles reasoning. In‑person work prioritized (Palo Alto, Seattle, Memphis, London). - **Benefits**: - Competitive cash + equity compensation. - Comprehensive medical, dental, vision, disability, life insurance. - 401(k) plan, fertility benefits, flexible vacation. - Visa sponsorship offered. - **Engineering culture**: No recruiters for assessments – applications reviewed directly by technical team. Interview process: application → screening → technical interviews → offer. - **Mission‑driven**: Opportunity to work on frontier AI models that aim to “advance humanity.” - **Perks**: In‑person collaboration, fast‑paced environment, ownership of meaningful projects. ## Sources 1. [x.ai/company](https://x.ai/company) 2. [x.ai/careers](https://x.ai/careers) 3. [job-boards.greenhouse.io/xai](https://job-boards.greenhouse.io/xai) 4. [x.ai/news](https://x.ai/news) 5. [x.ai/careers/open-roles](https://x.ai/careers/open-roles) ## Other roles at xAI - [Fraud Analyst (Sat-Weds)](https://feeny.ai/job/fraud-analyst-sat-weds-xai-palo-alto-ypm3merq7y8w) — Palo Alto, CA / New York, NY - [Legal Operations Specialist](https://feeny.ai/job/legal-operations-specialist-xai-palo-alto-yvgq5sxqmheq) — Palo Alto, CA - [Agency Development Manager](https://feeny.ai/job/agency-development-manager-xai-new-york-6zsw61w04xfw) — New York, NY - [AI Tutor - Design Specialist](https://feeny.ai/job/ai-tutor-design-specialist-xai-global-gyys8ysv19rj) — Global - [AI Tutor - Humanities](https://feeny.ai/job/ai-tutor-humanities-xai-global-f8b939y59f97) — Global - [Network Operations Center Specialist - Memphis](https://feeny.ai/job/network-operations-center-specialist-memphis-xai-southaven-5eaecc9mwq8n) — Southaven, MS / Memphis, TN - [Hardware Failure Analysis Engineer - Memphis](https://feeny.ai/job/hardware-failure-analysis-engineer-memphis-xai-southaven-111jdc07jvwy) — Southaven, MS / Memphis, TN - [Controls Engineer, Supercomputer Infrastructure - Memphis](https://feeny.ai/job/controls-engineer-supercomputer-infrastructure-memphis-xai-southaven-9tds2x1wpwjw) — Southaven, MS / Memphis, TN - [OT Systems Engineer (Supercomputer Infrastructure) - Memphis](https://feeny.ai/job/ot-systems-engineer-supercomputer-infrastructure-memphis-xai-southaven-qzd61f3q6z9j) — Southaven, MS / Memphis, TN - [Site Reliability Engineer - Memphis](https://feeny.ai/job/site-reliability-engineer-memphis-xai-southaven-cr8rfed4p2sp) — Southaven, MS / Memphis, TN