--- title: 'Staff DevOps Engineer at Nexxa.AI' canonical: 'https://feeny.ai/job/staff-devops-engineer-nexxa-ai-toronto-sk6nc60ttm7t' type: 'job' last_seen: '2026-09-08' --- # Staff DevOps Engineer at Nexxa.AI - **Company:** Nexxa.AI - **Location:** Toronto, Canada - **Employment:** full-time - **Work type:** remote - **Posted:** 2026-08-27 - **Last confirmed live:** 2026-09-08 - **Apply:** https://jobs.ashbyhq.com/nexxa/03c58f17-17b9-47f7-8558-d8544f764ac6 ## Job description Nexxa http://Nexxa.ai is building the best AI systems for heavy industries — enabling machines, systems, and operations to think, decide, and act autonomously across manufacturing, large-scale infrastructure, logistics, and legacy environments. Our mission is to translate deep technical breakthroughs into operational reality, solving some of the hardest systems-level problems in industry. ## ABOUT THE ROLE We're looking for a Senior/Staff DevOps Engineer who has spent the last several years building and operating the infrastructure that lets AI and industrial systems run reliably at scale. You understand what it takes to keep production ML and data workloads fast, observable, and resilient — from GPU-backed training and inference clusters to the pipelines that connect them to real-world industrial environments. This role is ideal for candidates who want deep infrastructure ownership at a company where uptime, latency, and reliability directly affect physical operations — not just software. You'll partner closely with AI, data, and product engineering teams to make sure the systems they build can actually run in production, safely and at scale. ## WHAT YOU'LL DO - Own and evolve Nexxa's core infrastructure — compute, networking, storage, and deployment systems — end-to-end - Design and operate CI/CD pipelines that support fast, safe iteration across AI, data, and product engineering teams - Build and maintain infrastructure-as-code (e.g., Terraform, Pulumi) for reproducible, auditable environments across cloud and on-prem/edge deployments - Architect and manage Kubernetes-based platforms for training, inference, and application workloads, including GPU scheduling and autoscaling - Partner with data and AI teams to support the infrastructure behind: - Data warehouses and lakehouse architectures (e.g., Snowflake, BigQuery, Redshift, Databricks) - Feature stores, embedding indices, and retrieval pipelines - Model training, evaluation, and serving infrastructure - Define and drive observability practices — metrics, logging, tracing, and alerting — across distributed systems - Establish and enforce reliability practices: SLOs/SLIs, incident response, postmortems, and on-call rotations - Design for security and compliance across cloud infrastructure, secrets management, and access control, particularly relevant to industrial and legacy-environment integrations - Make pragmatic tradeoffs across cost, latency, reliability, and developer velocity - Collaborate with engineering leadership to define infrastructure roadmap and platform strategy - Mentor engineers on infrastructure best practices and raise the bar for operational excellence across the org ## REQUIRED QUALIFICATIONS - 6+ years of experience in DevOps, Site Reliability Engineering, Platform Engineering, or infrastructure-focused software engineering roles - Deep hands-on experience with: - Cloud platforms (AWS, GCP, or Azure) at production scale - Kubernetes in production, including GPU workload scheduling - Infrastructure-as-code tooling (Terraform, Pulumi, or equivalent) - CI/CD systems (e.g., GitHub Actions, GitLab CI, CircleCI, Jenkins, ArgoCD) - Strong track record designing and operating observability stacks (e.g., Prometheus, Grafana, Datadog, OpenTelemetry) - Experience supporting ML/AI infrastructure — training clusters, model serving, data pipelines — a strong plus - Excellent scripting/programming skills (Python, Go, or Bash) for automation and tooling - Proven ability to independently scope and lead infrastructure projects from design through production rollout - Strong incident management instincts — you can lead through an outage calmly and drive toward root cause ## PREFERRED QUALIFICATIONS - Experience operating infrastructure that bridges cloud and edge/on-prem environments, especially in industrial or manufacturing contexts - Familiarity with data warehouse/lakehouse platforms (Snowflake, BigQuery, Redshift, Databricks) - Experience with service mesh, zero-trust networking, or compliance frameworks relevant to industrial/critical infrastructure (e.g., SOC 2, IEC 62443) - History of building internal developer platforms or self-service infrastructure tooling - Experience scaling infrastructure teams or setting technical direction at a Staff level ## WHAT SUCCESS LOOKS LIKE - You can own ambiguous, high-stakes infrastructure problems end-to-end - Systems you build stay reliable as usage and scale grow — you design for the next order of magnitude, not just today - You bring strong technical judgment on tradeoffs between reliability, cost, and speed - You raise the bar for operational rigor and engineering discipline across the team - You help define what's next for the platform, not just execute what's known WHY JOIN NEXXA.AI http://Nexxa.ai? - Innovative Environment: Play a critical role in transforming heavy industries through groundbreaking AI and automation technologies - Collaborative Culture: Be part of a team that values innovation, discipline, and continuous improvement - Professional Growth: Benefit from significant opportunities for career development and advancement - Competitive Compensation: Enjoy a comprehensive salary and equity package reflective of your expertise and contributions If you're passionate about building the infrastructure that powers advanced AI solutions in the real world, we'd love to connect. ## About Nexxa.AI ## Company Overview - **One-liner**: Nexxa.AI builds specialized multi-agent AI systems to automate industrial operations for heavy industries such as railroads, energy, construction, and manufacturing. - **Entity Type**: Private (Seed-stage, Privately Held) - **Headquarters**: Sunnyvale, California, United States - **Founded**: 2024 (some sources list 2023) - **Founders**: Philipp Wehn (CEO) and David Huang (CTO) ## Core Business - **Primary industries**: Artificial Intelligence, Industrial Automation, Heavy Industries (Railroad, Energy, Construction, Mining, Manufacturing) - **Target customers**: B2B, Enterprise – industrial engineering teams in project-heavy industries - **Mission**: Build an AI system that works alongside industrial engineers to accelerate engineering productivity without pre-training, ultimately achieving industrial autonomy. ## Products & Services - **Agentic AI Platform**: A multi-agent system that autonomously sets goals, defines success criteria, and executes operations directly into customers’ systems of record. Deployable on cloud or on-premise. Uses a combination of computer-use AI and computer vision. - **Nitro (Intelligence/Own)**: Captures decision logic, institutional knowledge, and operational judgment from fragmented, legacy software stacks, enabling AI agents to act on unstructured data. - **Full Self Computing**: A platform that processes unstructured data using generative AI for tasks like document analysis and decision-making. ## Market Standing - **Valuation**: Not publicly disclosed - **Total Funding**: USD 14.4 million - Pre-Seed (July 2025): $4.4M led by Andreessen Horowitz (a16z Speedrun) - Seed (January 2026): $9.0M led by Construct Capital - Non-Equity Assistance (November 2025): $1.0M from Amazon Web Services - **Key Metric**: Annual Recurring Revenue (ARR) grew 4× in production within the automotive vertical - **Growth Signals**: 400% headcount growth year-over-year (from ~2 to 26 employees); 8 active job postings; expanded presence to Canada - **Notable Investors**: Andreessen Horowitz (a16z), Construct Capital, Amazon Web Services ## Competitive Advantages - Deep specialization for heavy industries (railroad, energy, mining, manufacturing) rather than a horizontal AI platform - AI agents are entirely owned by the customer – no shared models or data leakage - Agents understand technical domain knowledge without any pre-training required - Combines computer-use AI with computer vision to interact directly with legacy industrial systems ## Strategic Focus - Expand across additional heavy industries (railroad, energy, construction, mining, manufacturing) - Achieve “industrial autonomy” – AI that writes directly into systems of record, moving from recommendations to autonomous execution - Scale from prototype to production in mission-critical environments, with deployment flexibility (cloud or on-premise) ## Why Work Here - **High-growth startup**: 400% headcount growth YoY, backed by top-tier investors (a16z, Construct Capital, AWS) - **Impactful mission**: Transform how industrial engineering teams work – days of repetitive work become minutes - **Team culture**: Emphasis on “work together and grow together”, resilience, and open collaboration - **Engineering-centric**: ~50% of team in technical roles (Applied AI Engineer, Senior AI Architect, etc.); active hiring across AI, product, and customer-facing roles - **Locations**: Sunnyvale, CA (HQ) and Canada; remote/hybrid policy not explicitly stated but presence in two countries suggests flexibility - **Perks**: Recent non-equity assistance from AWS, strong investor network, early-stage equity opportunity ## Sources 1. [nexxa.ai](https://nexxa.ai) 2. [LinkedIn - Nexxa.ai](https://www.linkedin.com/company/nexxa-ai) 3. [Crunchbase - Nexxa.ai](https://www.crunchbase.com/organization/nexxa-ai) 4. [Bloomberg - Nexxa AI Inc](https://www.bloomberg.com/profile/company/2584046D:US) 5. [PR Newswire - Funding Announcement](https://www.prnewswire.com) (implied from search results) ## Other roles at Nexxa.AI - [Security & Infrastructure Engineer](https://feeny.ai/job/security-infrastructure-engineer-nexxa-ai-toronto-6kvymg7sm3c7) — Toronto, Canada - [Security & Infrastructure Engineer](https://feeny.ai/job/security-infrastructure-engineer-nexxa-ai-san-francisco-5fvtj0rsjcyk) — San Francisco, CA - [Staff DevOps Engineer](https://feeny.ai/job/staff-devops-engineer-nexxa-ai-san-francisco-61tncfk1x35m) — San Francisco, CA - [QA Engineer (AI Systems)](https://feeny.ai/job/qa-engineer-ai-systems-nexxa-ai-toronto-33wekf9anshf) — Toronto, Canada - [QA Engineer (AI Systems)](https://feeny.ai/job/qa-engineer-ai-systems-nexxa-ai-san-francisco-1rwfk7sapb46) — San Francisco, CA - [Backend AI Engineer](https://feeny.ai/job/backend-ai-engineer-nexxa-ai-toronto-ctswqxgx1cq1) — Toronto, Canada - [AI Vision Engineer](https://feeny.ai/job/ai-vision-engineer-nexxa-ai-toronto-61gs4v0bx84p) — Toronto, Canada - [Backend AI Engineer](https://feeny.ai/job/backend-ai-engineer-nexxa-ai-san-francisco-mk2v68wtjfcd) — San Francisco, CA - [AI Vision Engineer](https://feeny.ai/job/ai-vision-engineer-nexxa-ai-san-francisco-abermqzmp3wj) — San Francisco, CA - [Applied AI Engineer](https://feeny.ai/job/applied-ai-engineer-nexxa-ai-munich-7g7m9w14c4f3) — Munich, Germany