--- title: 'Director, Site Operations at xAI' canonical: 'https://feeny.ai/job/director-site-operations-xai-memphis-rwmzd6w0k8n7' type: 'job' last_seen: '2026-09-21' --- # Director, Site Operations at xAI - **Company:** xAI - **Location:** Memphis, TN - **Posted:** 2026-09-15 - **Last confirmed live:** 2026-09-21 - **Apply:** https://job-boards.greenhouse.io/xai/jobs/5238779007 ## Job description SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates. ## ABOUT THE ROLE: As the Director of Site Operations, you’ll own node and rack uptime for SpaceXAI's AI supercompute cluster—the most advanced of its kind. This role is the extreme owner of cluster health and customer Service Level Agreements across 5+ sites operating 24/7. You’ll lead a 250+ person organization of site managers, shift supervisors, and technicians, plus the site reliability engineering team that monitors cluster health and drives fault mitigation at scale. We’re looking for a hands-on operations leader who can build a culture of excellence and accountability, partner tightly across the company, and keep uptime exceptional as we grow. ## RESPONSIBILITIES: - Own Cluster Uptime: Serve as extreme owner of node, rack, and cluster health across 5+ sites running 24/7, accountable for customer Service Level Agreements and consistently exceptional uptime on SpaceXAI's supercompute cluster. - Lead a Large Operations Organization: Direct a 250+ person team spanning site managers, shift supervisors, and technicians across four 24/7 shifts, building a culture of excellence and accountability at every layer of the org. - Drive Node and Rack Remediation: Ensure systematic recovery of failed nodes and racks through command-line and physical intervention, driving mean time to repair to the feasible minimum. - Partner Across Functions: Coordinate with facilities operations to limit downtime from power and cooling faults and proactive maintenance; with network engineering on cluster upgrades; and with tenant representatives on node remediation and planned and unplanned downtime. - Own Vendor Execution: Direct vendors through hardware rework and field operations so repairs, replacements, and capacity work happen at the speed the cluster requires. - Lead Site Reliability Engineering: Own the SRE organization responsible for proactive cluster health monitoring, reactive fault mitigation at scale, root cause analyses for node, rack, and cluster issues, and site-wide reliability procedures and fault documentation. - Run Data-Driven Improvement: Lead continual improvement and efficiency initiatives, using operational data to balance team resources and raise uptime, repair time, and SLA performance across sites. - Command Incidents at Scale: Set the standard for incident response during cluster-impacting events, providing clear direction, fast recovery, and tight communication with internal and external partners. - Scale Operations: Standardize best practices across sites and grow the organization in step with cluster expansion, keeping operations consistent as SpaceXAI's footprint scales. ## BASIC QUALIFICATIONS: - Bachelor’s degree and 7+ years of experience working in a large scale operations with 5+ years leading people leaders of technical teams OR 10+ years of experience working in a large scale operations with 5+ years leading people leaders of technical teams. ## PREFERRED SKILLS AND EXPERIENCE: - Proven ability to lead large, multi-site, 24/7 operations organizations in fast-paced, high-responsibility settings. - Deep expertise in server hardware, cluster reliability, and data center technologies, from deployment through lifecycle management. - Experience supporting compute-heavy environments like AI, machine learning, or high-performance computing at scale. - A track record of owning uptime, Service Level Agreements, or reliability metrics for large compute clusters. - Experience leading site reliability engineering or equivalent reliability-focused teams, including root cause analysis and procedure ownership. - Strong analytical skills and the ability to explain technical concepts clearly to diverse audiences, from technicians to executive and tenant partners. - A history of partnering with vendors at scale, driving mean time to repair down, and scaling operations across multiple sites. - Familiarity with tooling and automation (e.g., Jira, Python, Bash) used to monitor cluster health and improve team efficiency. - Enthusiasm for SpaceXAI's mission to accelerate human discovery and unravel the universe. - Ability to thrive in a dynamic, mission-focused environment with on-call ownership of cluster-impacting events. ## ADDITIONAL REQUIREMENTS: - Willingness to travel frequently to data center locations to support operations across sites. - Physical capability to handle data center tasks, including lifting up to 50 lbs unassisted, standing for long periods, and occasional ladder use. - Must be willing to work extended hours and/or weekends as needed. SpaceXAI is an equal opportunity employer. For details on data processing, view our [Recruitment Privacy Notice](https://x.ai/legal/recruitment-privacy-notice). ## About xAI ## Company Overview - **One‑liner**: xAI builds artificial intelligence systems to accelerate human scientific discovery and “understand the universe,” with products spanning frontier reasoning models, real‑time voice, and generative media. - **Entity Type**: Private (last funding round: Series E, January 2026; acquired by SpaceX in February 2026) - **Headquarters**: Palo Alto, California, USA (additional offices in Seattle, WA; Memphis, TN; London, UK) - **Founded**: 2023 - **Founders**: Elon Musk ## Core Business - Primary industry: Artificial Intelligence (AI research, large language models, voice AI, generative media) - Target customers: B2B (developers via API), B2C (consumers via Grok on 𝕏, SuperGrok subscriptions), Enterprise & Government (xAI for Government) - Mission: “To understand the universe” – building AI that extends what humanity can know and do. ## Products & Services - **Grok (base models)** : A conversational AI assistant modeled after *The Hitchhiker’s Guide to the Galaxy*, available on the 𝕏 platform and via API. - **Grok 4 / Grok 4 Heavy**: The most intelligent reasoning models, featuring native tool use and real‑time search integration (SuperGrok and Premium+ tiers). - **Grok Voice Agent API / Voice Agent Builder**: Create custom voice agents in under two minutes (no‑code) – includes speech‑to‑text, text‑to‑speech, custom voice cloning, and multilingual support. - **Grok Imagine API**: State‑of‑the‑art video generation API (quality, cost, latency). - **Grok‑1.5 / Grok‑2 / Grok‑2 mini**: Earlier model releases with improved reasoning and 128K context length. - **PromptIDE**: An integrated development environment for prompt engineering and interpretability research. - **xAI for Government**: A suite of frontier AI products tailored for U.S. government customers. - **Colossus Supercomputer**: 200,000‑GPU cluster built in 122 days, used for training xAI models; now also leased to Anthropic (Colossus 1 agreement). All products are delivered as SaaS, API, and model weights (open‑source release of Grok‑1). ## Market Standing - **Valuation / Funding**: - Series B (May 2024): $6 B - Series C (Dec 2024): $6 B (led by A16Z, BlackRock, Fidelity, Kingdom Holdings, Lightspeed, MGX, Morgan Stanley, OIA, QIA, Sequoia, Valor, Vy Capital, others) - Series E (Jan 2026): $20 B - Acquisition by SpaceX announced Feb 2, 2026 (terms not disclosed). - **Key Metric**: Total disclosed funding ~$32 B; annual revenue not publicly available. - **Notable Investors/Partners**: A16Z, BlackRock, Fidelity, Sequoia Capital, Morgan Stanley, Kingdom Holdings, Lightspeed, MGX, Valor Equity Partners, Vy Capital, Anthropic (Colossus deal), U.S. Department of War (selected Dec 2025). - **Growth Signals**: - Rapid scaling from founding to 200K‑GPU cluster in 122 days. - Over 214 open roles listed on Greenhouse as of mid‑2026. - #1 on Big Bench Audio (Grok 4 audio performance). - Expanded into voice agents, video generation, and government verticals within 3 years. - Acquisition by SpaceX signals deep integration with space‑tech resources. ## Competitive Advantages - **Compute infrastructure**: Colossus supercomputer – one of the largest GPU clusters in existence, enabling faster training and larger models. - **Model performance**: Grok 4 claims “most intelligent model in the world” with top benchmarks (Big Bench Audio #1). - **Vertical integration**: Tight coupling with 𝕏 platform for data, distribution, and real‑time world knowledge. - **Speed of execution**: “Move quickly and fix things” culture; built Colossus in 122 days. - **Government contracts**: Early access to public‑sector AI markets (U.S. Department of War). - **Talent density**: Small, focused team of researchers and engineers; in‑person work accelerates iteration. ## Strategic Focus - **Continue advancing frontier reasoning** (Grok 4 Heavy, next‑gen models). - **Expand voice and multimodal APIs** for developers (Voice Agent Builder, Custom Voices Library). - **Deepen government / defense partnerships** (xAI for Government, Department of War). - **Scale infrastructure** post‑acquisition via SpaceX resources. - **Remain “open” where strategic** (open‑source release of Grok‑1 weights). ## Why Work Here - **Culture**: “Ambitious goals, fast execution” – rejects consensus in favor of first‑principles reasoning. In‑person work prioritized (Palo Alto, Seattle, Memphis, London). - **Benefits**: - Competitive cash + equity compensation. - Comprehensive medical, dental, vision, disability, life insurance. - 401(k) plan, fertility benefits, flexible vacation. - Visa sponsorship offered. - **Engineering culture**: No recruiters for assessments – applications reviewed directly by technical team. Interview process: application → screening → technical interviews → offer. - **Mission‑driven**: Opportunity to work on frontier AI models that aim to “advance humanity.” - **Perks**: In‑person collaboration, fast‑paced environment, ownership of meaningful projects. ## Sources 1. [x.ai/company](https://x.ai/company) 2. [x.ai/careers](https://x.ai/careers) 3. [job-boards.greenhouse.io/xai](https://job-boards.greenhouse.io/xai) 4. [x.ai/news](https://x.ai/news) 5. [x.ai/careers/open-roles](https://x.ai/careers/open-roles) ## Other roles at xAI - [Driver (CDL) - Memphis](https://feeny.ai/job/driver-cdl-memphis-xai-southaven-m6jdbnnekqmr) — Southaven, MS / Memphis, TN - [Material Handler - Memphis](https://feeny.ai/job/material-handler-memphis-xai-southaven-v05f9fzy4z8p) — Southaven, MS / Memphis, TN - [Construction Scheduler - Memphis](https://feeny.ai/job/construction-scheduler-memphis-xai-southaven-gzz5ftzytgjn) — Southaven, MS / Memphis, TN - [Iron Worker (Construction)](https://feeny.ai/job/iron-worker-construction-xai-southaven-6rd1x8sg2337) — Southaven, MS / Memphis, TN - [Manager, Spaceport Events](https://feeny.ai/job/manager-spaceport-events-xai-palo-alto-wfg7wnr03dt9) — Palo Alto, CA - [HVAC Technician - Memphis](https://feeny.ai/job/hvac-technician-memphis-xai-southaven-ds1b392y442w) — Southaven, MS / Memphis, TN - [Plumber - Memphis (Service and Repair)](https://feeny.ai/job/plumber-memphis-service-and-repair-xai-southaven-849jrykkeyh7) — Southaven, MS / Memphis, TN - [Director of Growth Marketing and Sales Enablement](https://feeny.ai/job/director-of-growth-marketing-and-sales-enablement-xai-new-york-nz4w06jyyrea) — New York, NY - [Supervisor, Production Coordination (Logistics) - Memphis](https://feeny.ai/job/supervisor-production-coordination-logistics-memphis-xai-southaven-352xcgwsqyfx) — Southaven, MS / Memphis, TN - [GCR Agency Development Manager](https://feeny.ai/job/gcr-agency-development-manager-xai-singapore-jz8jcnf12wyg) — Singapore