Neura Robotics

GPU Cluster Engineer (human) at Neura Robotics (Metzingen, Germany / Riederich, Germany)

Neura Robotics· Metzingen, Germany / Riederich, Germany·

Role details

Work type
Onsite
Employment
Full-Time
Skills
AWS HyperPodSlurmKubernetesGPU Cluster OperationsDistributed TrainingInfrastructure as CodeCloud Cost ManagementSelf-service ToolingOperational DocumentationGerman

Neura Robotics at a glance

German cognitive-robotics company building humanoid and collaborative robots that see, hear, feel, and learn to work alongside people.

NEURA Robotics designs cognitive robots (humanoids, collaborative arms, and mobile bases) plus the AI and platform that run them, all built in-house. Its flagship 4NE1 humanoid targets industrial workflows and everyday assistance, and the Neuraverse platform links deployed robots into a shared learning network.

~$1.7B+ raised · latest: Series C · up to $1.4B · June 2026 · backed by Tether, Amazon, NVIDIA, Qualcomm

Summary

Design and operate NEURA's large-scale AWS HyperPod GPU cluster infrastructure, focusing on cluster stability, workload management, and self-service tooling for ML teams. The role involves optimizing GPU utilization, managing costs, and collaborating with AWS to influence platform platform at the hyperscaler level.

Job description

YOUR MISSION & CHALLENGES

  • You are the go-to expert for NEURA's GPU cluster infrastructure - a large-scale AWS HyperPod environment running cutting-edge GPU instances for foundation model training and customer fine-tuning workloads. You design the operational framework, build self-service tooling for ML teams, and work directly with AWS to influence the platform at the hyperscaler level.
  • Your focus is on cluster engineering and operations — not on ML research itself, but on making sure the people doing that research have rock-solid, efficient, and accessible infrastructure under them.
  • Setting up, configuring, and continuously evolving NEURA's HyperPod clusters, including HyperPod/Slurm and HyperPod/EKS orchestration models.
  • Designing and implementing strategies for cluster stability: node failure detection, automated job recovery, checkpoint coordination, and fault-tolerant multi-node training workflows.
  • Providing a workload priority management framework that allows multiple teams and use cases like foundation model pretraining, fine-tuning, customer workloads, to share cluster capacity efficiently and fairly.
  • Optimizing end-to-end GPU utilization: identifying and resolving bottlenecks across compute, GPU memory, EFA networking, and storage throughput.
  • Working directly and closely with the AWS HyperPod product and solutions engineering teams, escalating operational issues, sharing learnings from one of the platform's largest deployments, and placing concrete requirements on the roadmap.
  • Providing self-service tooling that allows ML researchers and engineers to launch, monitor, and manage training jobs independently, without requiring infrastructure intervention for routine operations.
  • Developing onboarding documentation, training materials, and internal workshops that enable users to operate efficiently, follow best practices, and understand cost implications of their workloads.
  • Infrastructure as Code is a given for you. Every cluster configuration, every operational change, every new environment is code first.
  • Owning the cost and capacity strategy: Spot instance management, Reserved Instance planning, Savings Plans, and ongoing commitment negotiations with AWS.

WHAT WE CAN LOOK FORWARD TO

  • 5+ years of experience in infrastructure or systems engineering, with a strong focus on GPU cluster or HPC operations.
  • Deep hands-on experience with AWS HyperPod and AWS instances; direct prior experience with HyperPod is a strong differentiator.
  • Solid understanding of both Slurm and Kubernetes as cluster orchestration layers, and the ability to evaluate their trade-offs for large-scale GPU workloads.
  • Practical knowledge of distributed training - you understand what affects throughput and how to debug it.
  • Experience building self-service tooling and operational documentation for technical end users.
  • You make complex infrastructure accessible, not just functional.
  • Strong understanding of cloud cost management at scale: Spot interruption handling, capacity reservations, cost attribution across teams and workloads.
  • Comfort working across organizational boundaries — your primary partners are ML researchers, but you'll also work closely with product, finance, and cloud vendor teams.
  • Strong English communication skills. German is a plus.

Why work at Neura Robotics

  • Culture: "Passion for Winning" — achievement-driven, self-responsible culture with flat hierarchies. Values include trust, honesty, speed, and "We Are Human" (people at the heart of everything).
  • Flexibility: Flexible working hours, 30 days of vacation, and a hybrid-friendly environment.
  • Team: Over 1,400 employees from more than 45 countries — highly skilled, international team of experts.
  • Growth & Development: Support for personal and professional growth via Udemy training platform and more.
  • Compensation: Competitive salary package with exclusive employee discounts.
  • Perks: Summer parties, company town hall meetings, and regular team celebrations.
  • Impact: Opportunity to work on cutting-edge cognitive robotics and AI that directly addresses labor shortages and makes a meaningful societal impact.
  • Diversity: Explicitly welcomes applications from all backgrounds, emphasizing building a team that reflects the world.
  • Application Process: Fast — initial feedback within a few days, interviews scheduled within two weeks.

Application questions