Watney

Staff Software Engineer - ML Infrastructure at Watney (San Francisco, CA)

Watney· San Francisco, CA·

Role details

Work type
Onsite
Employment
Full-Time

Job description

OUR MISSION

Expand human ambition in the physical world.

Critical infrastructure is constrained by labor shortages, hazardous working conditions, and operational complexity. Watney builds and deploys autonomous robotic systems that increase the speed and capacity of buildout, starting with data centers.

ABOUT THE ROLE

At Watney, ML Infrastructure engineers turn data collected from a live fleet of robots into better models. The fleet produces large volumes of video and telemetry data from real work in the field, and making that data trainable is one of the hardest systems problems at the company.

As we continue to scale, these systems will require larger training runs with more data, expanded clusters, and optimal GPU utilization.

WHAT YOU’LL DO

  • Own training and inference infrastructure
  • Build the data pipelines that these training runs depend on
  • Make experiments fast to launch and reproduce
  • Contribute to our core training code

YOU MAY BE A GOOD FIT IF YOU:

  • Have built ML infrastructure that carried real production training runs
  • Have scaled distributed training systems
  • Strong experience with Python, PyTorch or TensorFlow
  • Have experience identifying and troubleshooting GPU performance bottlenecks in large-scale training environments

We’re committed to building a diverse, inclusive team. At Watney Robotics, we welcome people of all backgrounds and identities, and we make hiring decisions based on skills, experience, and potential. If you’re passionate about robotics but don’t meet every requirement, we still encourage you to apply!

CURIOUS TO LEARN MORE?

Follow us here on X x.com/watneyrobotics and LinkedIn linkedin.com/about

Why work at Watney

  • In-Office culture: All roles are based in San Francisco (HQ); the company emphasizes in-person collaboration.
  • Ground-floor opportunity: With outsized ownership and visibility, employees shape the “rollout playbook” for a category-defining robotics company.
  • Technical challenge: Work on real systems that combine robotics, AI, hardware, and infrastructure at scale.
  • Notable peers: Colleagues include alumni from Scale AI, Amazon Web Services, Tools for Humanity, Warburg Pincus, and the US Army Corps of Engineers.
  • Growth trajectory: A very young company (founded 2024) with rapid hiring and clear product-market fit, ideal for early-stage career builders.

Application questions