--- title: 'Member Of Technical Staff - Cloud Infrastructure at xAI' canonical: 'https://feeny.ai/job/member-of-technical-staff-cloud-infrastructure-xai-palo-alto-byk5vgw8mqky' type: 'job' last_seen: '2026-09-14' --- # Member Of Technical Staff - Cloud Infrastructure at xAI - **Company:** xAI - **Location:** Palo Alto, CA - **Posted:** 2026-05-08 - **Last confirmed live:** 2026-09-14 - **Apply:** https://job-boards.greenhouse.io/xai/jobs/5133071007 ## Job description SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates. ## ABOUT THE ROLE: We are seeking a highly skilled Senior Infrastructure Engineer to join our US Government Team, focused on designing, building, and operating secure, scalable infrastructure for critical government projects. In this role, you will develop and manage training and inference clusters, as well as highly reliable applications, across bare metal, classified cloud, and hybrid cloud architectures. You will leverage your expertise in Kubernetes and GPU hardware to deliver robust, secure systems that support large-scale AI workloads while meeting stringent federal compliance requirements. This role demands a passion for automation, observability, and ensuring system integrity in a fast-paced, high-security environment. ## RESPONSIBILITIES: - Develop and optimize software to provision and manage SpaceXAI's infrastructure across on-premise, virtual machine, and classified cloud environments, enabling efficient scaling for US government initiatives. - Enhance the reliability, performance, and cost-effectiveness of infrastructure to support large-scale AI and application workloads in secure, classified settings. - Collaborate with SpaceXAI engineers to understand workload requirements and design tailored solutions that meet government-specific needs and compliance standards. - Implement robust observability, monitoring, and security practices to ensure the integrity, availability, and confidentiality of critical systems, adhering to federal protocols. - Manage storage infrastructure using Infrastructure-as-Code (IaC) tools such as Pulumi, Terraform, or Ansible, with a focus on secure data handling. - Drive system reliability through incident management, postmortems, and the definition of clear SLAs and SLOs, while maintaining security and compliance. - This is an in-person role based in Palo Alto, CA or Washington, DC, with up to 50% travel required. ## BASIC QUALIFICATIONS: - Active Top Secret (TS) security clearance. - 5+ years of experience as an Infrastructure Engineer, Site Reliability Engineer, or similar role, with a focus on building and maintaining reliable, scalable systems, preferably in secure or government environments. - Proficiency in managing storage infrastructure with IaC tools such as Pulumi, Terraform, or Ansible. - Deep understanding of the Kubernetes stack, including CNI, CRI, CSI, and related components. - Demonstrated ability to improve system reliability through incident management, postmortems, and defining SLAs/SLOs. - Excellent communication and documentation skills, with the ability to handle sensitive information concisely and accurately. ## PREFERRED SKILLS AND EXPERIENCE: - Deep familiarity with installing and using GPU hardware, including setting up drivers, debugging issues, and ensuring reliability. - Experience with high-traffic web or mobile application workloads, including optimizing Kubernetes for large-scale deployments in classified or federal settings. - Familiarity with chaos engineering, capacity planning, or similar practices for ensuring system resilience in government projects. - Proficiency with tools such as Kyverno, ArgoCD, or Go programming for infrastructure automation. - Strong sense of ownership, curiosity, and enthusiasm for tackling complex technical challenges in secure environments. - Passion for problem-solving and a proactive drive to deliver impactful results while adhering to security protocols. - Certifications in security-related fields (e.g., CISSP) or experience in secure federal environments. ## COMPENSATION AND BENEFITS: $180,000 - $440,000 USD Base salary is just one part of our total rewards package at SpaceXAI, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks. SpaceXAI is an equal opportunity employer. For details on data processing, view our [Recruitment Privacy Notice](https://x.ai/legal/recruitment-privacy-notice). ## About xAI ## Company Overview - **One‑liner**: xAI builds artificial intelligence systems to accelerate human scientific discovery and “understand the universe,” with products spanning frontier reasoning models, real‑time voice, and generative media. - **Entity Type**: Private (last funding round: Series E, January 2026; acquired by SpaceX in February 2026) - **Headquarters**: Palo Alto, California, USA (additional offices in Seattle, WA; Memphis, TN; London, UK) - **Founded**: 2023 - **Founders**: Elon Musk ## Core Business - Primary industry: Artificial Intelligence (AI research, large language models, voice AI, generative media) - Target customers: B2B (developers via API), B2C (consumers via Grok on 𝕏, SuperGrok subscriptions), Enterprise & Government (xAI for Government) - Mission: “To understand the universe” – building AI that extends what humanity can know and do. ## Products & Services - **Grok (base models)** : A conversational AI assistant modeled after *The Hitchhiker’s Guide to the Galaxy*, available on the 𝕏 platform and via API. - **Grok 4 / Grok 4 Heavy**: The most intelligent reasoning models, featuring native tool use and real‑time search integration (SuperGrok and Premium+ tiers). - **Grok Voice Agent API / Voice Agent Builder**: Create custom voice agents in under two minutes (no‑code) – includes speech‑to‑text, text‑to‑speech, custom voice cloning, and multilingual support. - **Grok Imagine API**: State‑of‑the‑art video generation API (quality, cost, latency). - **Grok‑1.5 / Grok‑2 / Grok‑2 mini**: Earlier model releases with improved reasoning and 128K context length. - **PromptIDE**: An integrated development environment for prompt engineering and interpretability research. - **xAI for Government**: A suite of frontier AI products tailored for U.S. government customers. - **Colossus Supercomputer**: 200,000‑GPU cluster built in 122 days, used for training xAI models; now also leased to Anthropic (Colossus 1 agreement). All products are delivered as SaaS, API, and model weights (open‑source release of Grok‑1). ## Market Standing - **Valuation / Funding**: - Series B (May 2024): $6 B - Series C (Dec 2024): $6 B (led by A16Z, BlackRock, Fidelity, Kingdom Holdings, Lightspeed, MGX, Morgan Stanley, OIA, QIA, Sequoia, Valor, Vy Capital, others) - Series E (Jan 2026): $20 B - Acquisition by SpaceX announced Feb 2, 2026 (terms not disclosed). - **Key Metric**: Total disclosed funding ~$32 B; annual revenue not publicly available. - **Notable Investors/Partners**: A16Z, BlackRock, Fidelity, Sequoia Capital, Morgan Stanley, Kingdom Holdings, Lightspeed, MGX, Valor Equity Partners, Vy Capital, Anthropic (Colossus deal), U.S. Department of War (selected Dec 2025). - **Growth Signals**: - Rapid scaling from founding to 200K‑GPU cluster in 122 days. - Over 214 open roles listed on Greenhouse as of mid‑2026. - #1 on Big Bench Audio (Grok 4 audio performance). - Expanded into voice agents, video generation, and government verticals within 3 years. - Acquisition by SpaceX signals deep integration with space‑tech resources. ## Competitive Advantages - **Compute infrastructure**: Colossus supercomputer – one of the largest GPU clusters in existence, enabling faster training and larger models. - **Model performance**: Grok 4 claims “most intelligent model in the world” with top benchmarks (Big Bench Audio #1). - **Vertical integration**: Tight coupling with 𝕏 platform for data, distribution, and real‑time world knowledge. - **Speed of execution**: “Move quickly and fix things” culture; built Colossus in 122 days. - **Government contracts**: Early access to public‑sector AI markets (U.S. Department of War). - **Talent density**: Small, focused team of researchers and engineers; in‑person work accelerates iteration. ## Strategic Focus - **Continue advancing frontier reasoning** (Grok 4 Heavy, next‑gen models). - **Expand voice and multimodal APIs** for developers (Voice Agent Builder, Custom Voices Library). - **Deepen government / defense partnerships** (xAI for Government, Department of War). - **Scale infrastructure** post‑acquisition via SpaceX resources. - **Remain “open” where strategic** (open‑source release of Grok‑1 weights). ## Why Work Here - **Culture**: “Ambitious goals, fast execution” – rejects consensus in favor of first‑principles reasoning. In‑person work prioritized (Palo Alto, Seattle, Memphis, London). - **Benefits**: - Competitive cash + equity compensation. - Comprehensive medical, dental, vision, disability, life insurance. - 401(k) plan, fertility benefits, flexible vacation. - Visa sponsorship offered. - **Engineering culture**: No recruiters for assessments – applications reviewed directly by technical team. Interview process: application → screening → technical interviews → offer. - **Mission‑driven**: Opportunity to work on frontier AI models that aim to “advance humanity.” - **Perks**: In‑person collaboration, fast‑paced environment, ownership of meaningful projects. ## Sources 1. [x.ai/company](https://x.ai/company) 2. [x.ai/careers](https://x.ai/careers) 3. [job-boards.greenhouse.io/xai](https://job-boards.greenhouse.io/xai) 4. [x.ai/news](https://x.ai/news) 5. [x.ai/careers/open-roles](https://x.ai/careers/open-roles) ## Other roles at xAI - [AI Tutor - Video (Weekend)](https://feeny.ai/job/ai-tutor-video-weekend-xai-global-ah7fpybc60zs) — Global - [Expert Team Lead, Human Data Operations](https://feeny.ai/job/expert-team-lead-human-data-operations-xai-asia-europe-international-us-dd18bx8p1129) — Asia / Europe / International / US - [Team Lead, Human Data Operations - Post Training](https://feeny.ai/job/team-lead-human-data-operations-post-training-xai-asia-europe-international-us-7mqwd37c5cse) — Asia / Europe / International / US - [Controls Technician (Physical Infrastructure) - Memphis](https://feeny.ai/job/controls-technician-physical-infrastructure-memphis-xai-southaven-62rxn4p3hs02) — Southaven, MS / Memphis, TN - [Fraud Analyst (Sat-Weds)](https://feeny.ai/job/fraud-analyst-sat-weds-xai-palo-alto-ypm3merq7y8w) — Palo Alto, CA / New York, NY - [Agency Development Manager](https://feeny.ai/job/agency-development-manager-xai-austin-6zsw61w04xfw) — Austin, TX / New York, NY - [AI Tutor - Design](https://feeny.ai/job/ai-tutor-design-xai-global-gyys8ysv19rj) — Global - [AI Tutor - Humanities](https://feeny.ai/job/ai-tutor-humanities-xai-global-f8b939y59f97) — Global - [Network Operations Center Specialist - Memphis](https://feeny.ai/job/network-operations-center-specialist-memphis-xai-southaven-5eaecc9mwq8n) — Southaven, MS / Memphis, TN - [Hardware Failure Analysis Engineer - Memphis](https://feeny.ai/job/hardware-failure-analysis-engineer-memphis-xai-southaven-111jdc07jvwy) — Southaven, MS / Memphis, TN