--- title: 'ML Cloud Infrastructure Engineer at HavocAI' canonical: 'https://feeny.ai/job/ml-cloud-infrastructure-engineer-havocai-remote-kjajjwn4sss5' type: 'job' last_seen: '2026-09-17' --- # ML Cloud Infrastructure Engineer at HavocAI - **Company:** HavocAI - **Location:** Remote - **Compensation:** $150k–$175k - **Employment:** full-time - **Work type:** remote - **Posted:** 2026-09-14 - **Last confirmed live:** 2026-09-17 - **Apply:** https://jobs.ashbyhq.com/havocai/d1546df4-f0fd-4db6-8cc9-b75b93c9a951 ## Job description About Us: Havoc is a leader in all-domain collaborative autonomy. Its software-defined hardware approach powers military and commercial-grade autonomous systems across sea, air, and land to sense, decide, and act together in complex and contested environments. Havoc connects assets, enabling them to share information, adapt in real time, and continue operating even when communications are disrupted or denied. Havoc optimizes mission performance and minimizes human risk. Havoc was founded in 2024 and headquartered in Providence, Rhode Island. Learn more at [Havoc: All-Domain Collaborative Autonomy](http://havocai.com/) . ## About the Role As a Machine Learning Cloud Infrastructure Engineer, you will build and operate the infrastructure HavocAI teams use to train, evaluate, deploy, and monitor machine learning models safely and reliably. You will develop the pipelines, services, platforms, and integrations connecting data lakes, telemetry stores, simulation environments, training workloads, cloud compute, and deployed models. Your work will enable autonomy, data, and software engineers to move efficiently from raw field and simulation data to reproducible datasets, scalable training, rigorous evaluation, and production deployment. This role is ideal for a strong software, infrastructure, or data engineer with hands-on ML experience who enjoys building production systems at the intersection of data, models, compute, and autonomy. You should be comfortable operating in a fast-paced environment, solving ambiguous infrastructure problems, and building systems that must remain scalable, observable, secure, and reliable. ## What You’ll Do ML Infrastructure & Pipelines - Build pipelines that transform raw multi-modal data—including telemetry, imagery, video, sensor, and simulation data—into curated, versioned training datasets. - Develop reproducible training and evaluation workflows that scale across cloud compute and GPU resources. - Build and maintain model deployment infrastructure for packaging, serving, inference, versioning, and rollback. - Implement experiment tracking, dataset lineage, model versioning, and other capabilities required for reproducible ML development. - Own data schema versioning and migration across pipelines, data lakes, and services as datasets and models evolve. Cloud Platform & Infrastructure - Design, build, and operate scalable AWS infrastructure using Infrastructure as Code. - Build and maintain Kubernetes/EKS workloads and containerized environments for training, batch processing, evaluation, and model serving. - Develop self-service tooling and paved paths for compute scheduling, storage, data access, training, and deployment. - Improve utilization, scalability, and cost efficiency across cloud and accelerator infrastructure. - Build infrastructure that enables engineering teams to launch workloads safely without unnecessary operational overhead. Reliability, Evaluation & Observability - Build evaluation frameworks and regression testing for model quality, dataset integrity, and pipeline correctness. - Establish quality and reliability signals that help determine whether models are ready for production use. - Develop monitoring, logging, tracing, and observability across training jobs, data pipelines, and deployed models. - Diagnose and resolve performance, scaling, reliability, and infrastructure bottlenecks. - Maintain high standards for automation, testing, documentation, and operational readiness. Cross-Functional Engineering - Partner with Autonomy, Software, Data, Simulation, and Security teams to build ML infrastructure spanning edge data capture through cloud training and model deployment. - Contribute to CI/CD and release processes for models, datasets, and ML pipelines. - Translate engineering requirements into scalable platform capabilities that can support multiple teams and use cases. - Incorporate feedback from engineers and users to continuously improve ML development workflows. Security & Data Management - Implement secure infrastructure practices, including IAM least privilege, secrets management, access controls, and secure handling of sensitive and defense-related data. - Build data and ML workflows with reproducibility, traceability, and appropriate controls from the start. - Partner with security and infrastructure teams to ensure ML systems meet applicable operational and compliance requirements. ## What We’re Looking For - 3+ years of experience in software engineering, infrastructure engineering, data engineering, ML infrastructure, or a related field. - Strong programming experience in Python, with experience in Go, C++, or another systems-oriented language preferred. - Experience building and operating production services, APIs, data pipelines, developer platforms, or infrastructure. - Hands-on experience with ML workflows such as dataset preparation, model training, evaluation, or deployment. - Experience with cloud infrastructure, preferably AWS, and Infrastructure as Code. - Hands-on experience with Kubernetes and containerized environments. - Strong understanding of production engineering fundamentals, including reliability, observability, testing, automation, and maintainability. - Ability to work effectively across engineering disciplines and solve ambiguous technical problems with a high degree of ownership. - U.S. Citizenship and ability to obtain and maintain a U.S. Government security clearance. ## Nice to Have - Experience with MLOps and workflow platforms such as MLflow, Weights & Biases, Kubeflow, Ray, Airflow, or Dagster. - Experience with GPU/accelerator scheduling, distributed training, or large-scale ML workloads. - Experience working with multi-modal datasets including imagery, video, telemetry, sensor, or simulation data. - Experience supporting autonomy, robotics, simulation, or real-time systems. - Experience deploying ML models to edge or embedded environments. - Experience with AWS GovCloud, GCP Assured Workloads, or compliance-driven environments such as FedRAMP or IL4/IL5. What Success Looks Like Within your first 12 months, you will have: - Built reliable and reproducible pipelines that move data from raw capture through curated, versioned training datasets. - Enabled training and evaluation workloads to run at scale with clear quality and reliability signals before deployment. - Delivered self-service ML infrastructure that reduces friction and accelerates Autonomy, Data, and Software teams. - Improved the observability, reliability, security, and cost efficiency of HavocAI’s ML infrastructure. - Established scalable foundations that allow ML development and deployment to grow alongside HavocAI’s autonomy capabilities. Benefits: - 100% Employer paid Health, Dental and Vision Insurance for you and your families - Life Insurance (Employer Paid) - Ability to participate in the companies 401k program (Matching) - Unlimited PTO policy with an enforced 2 week minimum - Equity Package - Work / Home Office Stipend - Global Entry - 16 Week Paid Parental Leave - Monthly Health and Wellness Stipend Our Values: - Innovation: We are driven to break new ground. Every day presents an opportunity to challenge the status quo, think boldly, and deliver advanced solutions that transform the future of defense technology. - Integrity: We hold ourselves to the highest ethical standards, ensuring transparency, accountability, and trust in all our actions and partnerships. - Mission-Driven: We are focused on achieving impactful outcomes that align with our core mission—protecting lives through innovation. - Forward-Leaning: We continuously seek out new opportunities and remain at the forefront of technological advancements. We embrace change and anticipate the challenges of tomorrow with confidence and creativity. - Ownership of All Tasks: At HavocAI, no problem is too complex or too trivial. We believe that greatness comes from tackling the hardest challenges, but also in handling the smallest, sometimes thankless, tasks with the same level of commitment and care. - Servant Leadership: We lead by serving others, whether it’s supporting our employees, partners, or the broader community. Empowering those around us is key to achieving long-term success and making a lasting impact. HavocAI is an Equal Opportunity Employer and is committed to creating an inclusive and diverse workplace. We welcome applicants from all backgrounds and do not discriminate based on race, color, religion, gender, sexual orientation, age, national origin, disability, veteran status, or any other legally protected status. ## About HavocAI ## Company Overview - **One-liner**: HavocAI builds collaborative autonomy software that enables a single human operator to supervise thousands of autonomous assets across sea, air, and land in contested environments. - **Entity Type**: Private (Series A) - **Headquarters**: Providence, Rhode Island, United States - **Founded**: 2024 - **Founders**: Paul Lwin (CEO) and Joe Turner (COO) ## Core Business - **Primary industry/industries**: Defense and Space Manufacturing, Autonomous Systems, Maritime Security - **Target customers**: U.S. Navy, Army, and allied defense forces; commercial maritime security operators - **Mission or purpose statement**: "Minimize human risk in the world's most unstable environments, removing humans from the dull, dirty, and dangerous." ## Products & Services - **Havoc C2**: An intuitive command-and-control interface that lets operators define intent and direct large numbers of autonomous assets across distributed systems, scaling complex multi-asset missions without proportional increases in manpower. - **Havoc Connect**: A resilient peer-to-peer overlay network for data exchange across DDIL (Disrupted, Disconnected, Intermittent, Limited) networks, enabling systems to exchange commands, telemetry, and mission data even when connectivity is intermittent or degraded. - **Havoc OS**: An operating system for autonomous systems at the edge, enabling heterogeneous platforms to carry out collaborative missions, adapt to changing conditions, and continue operating in DDIL environments. - **Havoc Insights**: An intelligence layer that fuses AI and analytics, ingesting real-time data streams from vehicles, sensors, and system services to surface insights, trends, and anomalies for mission-level decision support. ## Market Standing - **Valuation/Market Cap**: Not publicly disclosed - **Key Metric**: Total Funding of $197.2M across three rounds (Seed, two Series A tranches) - **Notable Investors/Partners**: Scout Ventures, Trousdale Ventures, IQT (lead Series A), Azak (MOU for autonomous ground systems) - **Growth Signals**: Headcount grew 221.3% year-over-year to 98 employees; acquired Mavrik in March 2026; over 25,000 hours of autonomous operations; 11 active job postings with quarterly growth of +10.0% ## Competitive Advantages - **Software-defined autonomy**: Havoc's platform is platform-agnostic, integrating with existing vehicles, sensors, and C2 networks without major redesigns, making it interoperable across sea, air, and land domains. - **Human-on-the-loop architecture**: Enables a single operator to supervise thousands of autonomous assets, dramatically reducing manpower requirements for complex multi-asset missions. - **DDIL resilience**: Havoc Connect creates a resilient peer-to-peer overlay network that maintains operations even when communications are disrupted, intermittent, or denied — a critical capability for contested military environments. - **Proven track record**: Over 25,000 hours of autonomous operations, validated through real-world exercises with U.S. Navy and Army stakeholders. ## Strategic Focus - **All-domain collaborative autonomy**: Scaling from maritime to air and land domains, enabling heterogeneous platforms to work together seamlessly. - **Fully adaptive maritime autonomy**: Building systems that can anticipate, collaborate, and make independent decisions in dynamic environments. - **International expansion**: Head of Europe appointed (Theodora Dell), indicating a push into allied defense markets. - **Ground systems expansion**: MOU with Azak to produce high mobility autonomous ground systems, diversifying beyond maritime. ## Why Work Here - **Mission-driven work**: Employees work on cutting-edge autonomy technology that directly reduces human risk in dangerous environments — "removing humans from the dull, dirty, and dangerous." - **Strong growth trajectory**: 221.3% headcount growth year-over-year, with $197.2M in total funding, indicating a well-capitalized, scaling organization. - **Comprehensive benefits**: 100% employer-paid medical, dental, and vision for employees and families; stock options; 16 weeks fully paid parental leave; unlimited PTO with a two-week enforced minimum; monthly wellness stipend; home office stipend; 401(k) program; Global Entry reimbursement. - **Work model**: Remote and hybrid options available, with corporate HQ, test, and production facilities in Providence, RI. Open roles also listed in Honolulu, HI and San Diego, CA. - **Engineering culture**: Deep technical challenges in autonomy, robotics, edge computing, and defense-grade software. Team includes talent from US Navy, MIT Lincoln Laboratory, Anduril, Shield AI, and Lockheed Martin. ## Sources 1. [havocai.com](https://www.havocai.com/) 2. [havocai.com/careers](https://havocai.com/careers) 3. [havocai.com/about](https://havocai.com/about) 4. [linkedin.com/company/havocai](https://www.linkedin.com/company/havocai) 5. [jobs.ashbyhq.com/havocai](https://jobs.ashbyhq.com/havocai) ## Other roles at HavocAI - [Data and ML Infrastructure Engineer](https://feeny.ai/job/data-and-ml-infrastructure-engineer-havocai-remote-cvktwbd92gbt) - [Robotics Simulation Engineer](https://feeny.ai/job/robotics-simulation-engineer-havocai-remote-h61rw9a3gxhm) - [Engineering Manager, Perception and Machine Learning](https://feeny.ai/job/engineering-manager-perception-and-machine-learning-havocai-remote-8gamv127ddt4) - [Test Engineer](https://feeny.ai/job/test-engineer-havocai-north-kingstown-88kgx527zff0) — North Kingstown, RI - [Demo Operations Specialist](https://feeny.ai/job/demo-operations-specialist-havocai-north-kingstown-cp5k0x8ttcxk) — North Kingstown, RI - [Electrical Field Service Technician](https://feeny.ai/job/electrical-field-service-technician-havocai-north-kingstown-65kf1cmwvyrn) — North Kingstown, RI - [Rapid Prototyping Engineer – Drones & Unmanned Systems](https://feeny.ai/job/rapid-prototyping-engineer-drones-unmanned-systems-havocai-san-diego-cq5dzh1v2bvs) — San Diego, CA - [Mechanical Engineer – Drone Development, Design & Test](https://feeny.ai/job/mechanical-engineer-drone-development-design-test-havocai-san-diego-074mpqathxd2) — San Diego, CA - [Recruiter](https://feeny.ai/job/recruiter-havocai-san-diego-p0ty7dde3xaa) — San Diego, CA - [Senior Drone Engineer](https://feeny.ai/job/senior-drone-engineer-havocai-san-diego-5ban077zqqxg) — San Diego, CA