--- title: 'AI Engineer at NetBrain' canonical: 'https://feeny.ai/job/ai-engineer-netbrain-toronto-xnfaf69qtea8' type: 'job' last_seen: '2026-09-20' --- # AI Engineer at NetBrain - **Company:** NetBrain - **Location:** Toronto, Canada - **Work type:** hybrid - **Posted:** 2026-09-18 - **Last confirmed live:** 2026-09-20 - **Apply:** https://job-boards.greenhouse.io/netbrain/jobs/5241572007 ## Job description Founded in 2004, NetBrain is the leader in no-code network automation. Its ground-breaking Next-Gen platform provides IT operations teams with the ability to scale their hybrid multi-cloud connected networks by automating the processes associated with Diagnostic Troubleshooting, Outage Prevention and Protected Change Management.  Today, over 2,500 of the world’s largest enterprises and managed services providers leverage NetBrain’s platform. ## What We Need We’re looking for a Senior AI Engineer to design and build production-grade agent and RAG systems that power intelligent, reliable automation across our platform. This role combines hands-on engineering with system-level thinking—owning everything from architecture and evaluation to scalability, observability, and reliability in production. The ideal candidate thrives in ambiguity, moves quickly from prototype to production, and brings a strong focus on quality, safety, and real-world impact. ## What You'll Do Agent Platform Architecture - Design and implement core capabilities for an enterprise-grade Agent platform, including orchestration patterns such as ReAct, Plan-and- - Execute, and Supervisor, as well as tool execution, context and memory management, and safety guardrails. - Design enterprise-grade Agent execution and governance mechanisms, including Human-in-the-Loop approval workflows, multi-tenant - permission isolation, policy enforcement, and secure execution controls. - Build reusable Agent Skills, standardized tool interfaces, and a scalable tool ecosystem deeply integrated with NetBrain platform - capabilities and business workflows. LLM and Model Optimization - Design and implement LLM post-training strategies, including domain-specific Supervised Fine-Tuning (SFT), DPO/RLHF-based - preference alignment, and parameter-efficient fine-tuning techniques such as LoRA, to continuously improve model performance in the - network operations domain. - Build an Agent self-learning feedback loop that converts production execution traces, user feedback, and evaluation results into high- - quality datasets for continuously improving prompts, skills, models, and retrieval strategies. - Analyze and optimize LLM behavior across areas such as instruction following, tool calling, structured output generation, contextual - understanding, reasoning stability, and hallucination mitigation. Evaluation, Reliability, and Observability - Build production-grade LLM and Agent evaluation frameworks and automated regression pipelines, including benchmark datasets, - deterministic checks, LLM-as-a-Judge, tool-call validation, retrieval-quality evaluation, and end-to-end task success metrics. - Establish release quality gates and hallucination-detection mechanisms for AI features to prevent significant accuracy, reliability, and - performance regressions from reaching production. - Build comprehensive AI system observability capabilities, including distributed tracing, structured logging, metrics, dashboards, and - alerting. - Rapidly diagnose and resolve production AI failures, including hallucinations, incorrect tool selection, invalid tool parameters, Agent - loops, retrieval-quality degradation, structured-output failures, latency regressions, and unexpected model behavior changes. Production Engineering and Technical Execution - Design and implement highly reliable backend services for production AI and Agent workloads, including asynchronous and concurrent - processing, retries, timeouts, caching, rate limiting, and fault isolation. - Continuously optimize latency, throughput, token consumption, and infrastructure cost to meet platform SLA requirements and support - large-scale production workloads. - Independently diagnose and resolve complex AI system issues spanning prompts, models, RAG, tools, Agent workflows, backend - services, and infrastructure. - Lead technical design for critical modules and system-level capabilities, ensuring solutions align with platform architecture, security - requirements, engineering standards, and product requirements. - Drive technical improvements based on production data, evaluation results, and benchmarks, and collaborate closely with Engineering, - Product, QA, and other teams to deliver solutions into production. Applied Research and Technical Strategy - Prototype, benchmark, and productionize emerging technologies such as GraphRAG, Knowledge Graphs, MCP, LLM Post-Training, and - Agent Self-Learning to improve grounding, multi-hop reasoning, and domain expertise. - Continuously evaluate Agent frameworks and supporting infrastructure, including LangChain, LangGraph, AutoGen, and LlamaIndex, - and provide technical recommendations for platform architecture evolution and product technology strategy. - Stay current with developments in LLM and Agent technologies and rapidly translate promising technologies into measurable, testable, - and production-ready engineering capabilities. ## What You Bring - Bachelor's degree or higher in Computer Science, Artificial Intelligence, Electrical Engineering, or a related technical field. Master's or Ph.D. - preferred; equivalent practical experience will also be considered. - 3+ years of experience in software engineering, machine learning, or applied AI, including 2+ years building, deploying, and operating production- - grade LLM or Agent applications. Must have delivered at least one LLM-powered feature end-to-end and owned its ongoing operation and - improvement after production launch. - Deep understanding of Agent architectures and LLM behavioral characteristics, including instruction following, tool-calling behavior, and context - sensitivity, with hands-on experience building multi-step workflows involving reasoning, tool execution, state management, structured outputs, - validation, and error recovery. - Proven ability to diagnose and resolve production LLM/Agent failures, including hallucinations, incorrect tool calls, retrieval-quality degradation, - Agent loops, structured-output failures, latency regressions, and regressions introduced by prompt or model changes. - Strong Python and distributed backend engineering skills, including API and service development, asynchronous and concurrent programming, - retries, timeouts, caching, rate limiting, testing, logging, and cross-service performance debugging. - Hands-on experience designing evaluation systems for LLM applications, including dataset construction, metric definition, regression testing, and - release quality gates. - Strong understanding of security risks associated with LLM and Agent applications, including prompt injection, data leakage, unsafe tool - execution, permission boundaries, and uncontrolled Agent autonomy. - Ability to independently design, implement, debug, deploy, and operate complex production systems. ## Preferred Qualifications - Strong experience with RAG and advanced retrieval systems, including embeddings, vector and hybrid search, reranking, chunking strategies, - grounding and citation mechanisms, as well as multi-hop retrieval, Knowledge Graphs, or GraphRAG. - Familiarity with Agent frameworks such as LangGraph, LangChain, AutoGen, and LlamaIndex, as well as standardized protocols such as MCP - (Model Context Protocol) for connecting Agents with external tools and systems. A strong understanding of the underlying architecture is more - important than expertise in any specific framework.Experience designing Agent runtime mechanisms, including Human-in-the-Loop workflows such as risk classification, approval, interruption, pause - /resume, as well as context management for long-running Agent workflows, including state persistence, history compression, and memory - systems. - Experience with LangSmith or similar LLM observability and evaluation platforms. - Experience with LLM fine-tuning, including LoRA or other parameter-efficient fine-tuning techniques. - Experience applying LLM technologies to networking, infrastructure, cybersecurity, observability, or other complex technical domains. - Fluent in both English and Chinese, with strong cross-regional communication and collaboration skills. ## What We Offer Our comprehensive compensation package is vital in how we recognize our people for the impact they make on us reaching our goals as a company. For this role, the estimated base is CAD $130,000 - CAD $165,000 + Bonus. The actual salary may vary based on a range of factors, including market and individual qualifications objectively assessed during the interview process. The range listed above is a guideline and may be modified. People Experience offers a comprehensive benefits package in addition to cash compensation that includes but is not limited to RRSP and medical/dental coverage. Speak with your Recruiter for more details on our Total Rewards philosophy. ## #LI-BW1 NetBrain invites all interested and qualified candidates to apply for employment opportunities. Qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, disability, protected veteran status, or other characteristics protected by law. If you have a disability that prevents or limits your ability to use or access the site, or if you require any other accommodation in the application process due to a disability, you may request a reasonable accommodation. To make a request, please contact our People Team at: people@netbraintech.com and we will be happy to assist you. In compliance with applicable laws, NetBrain conducts holistic, individual background reviews in support of all hiring decisions. It is unlawful in Massachusetts to require or administer a lie detector test as a condition of employment or continued employment. An employer who violates this law shall be subject to criminal penalties and civil liability. ## About NetBrain ## Company Overview - **One-liner**: NetBrain provides an intent-based network automation and visibility platform that uses AI agents to diagnose, decide, and act on network issues, helping enterprises prevent outages and reduce manual operations. - **Entity Type**: Private (Acquired by Blackstone Group in July 2025) - **Headquarters**: Burlington, Massachusetts, United States - **Founded**: 2004 - **Founders**: Not publicly listed on primary sources ## Core Business - **Primary industry**: Network Automation, IT Operations, Cybersecurity - **Target customers**: Large enterprises, including approximately one-third of the Fortune 100 and Fortune 500; B2B, Enterprise - **Mission or purpose**: To deliver autonomous network operations through AI agents that diagnose, decide, and act with full network context, ensuring seamless network reliability for critical services worldwide. ## Products & Services - **NetBrain Next-Gen Platform**: An intent-based network automation and visibility platform that automates more than 95% of network tickets and prevents over 50% of network problems before they impact production. It integrates with major ITSM systems and provides hybrid-cloud visibility, map-based diagnostics, and proactive compliance assessment. - **Agentic NetOps**: AI agents that autonomously diagnose, decide, and act on network issues with full context, reducing manual operations and escalations. - **Hybrid-Cloud Visibility and Mapping**: Instantly creates documentation, hybrid-cloud maps, A-B paths, and inventory reports. - **Compliance & Security**: Proactively assesses networks 24/7 to prevent outages and ensure compliance, reducing security risk. ## Market Standing - **Valuation/Market Cap**: Acquired by Blackstone Group in July 2025 for $750 million [linkedin.com](https://www.linkedin.com/company/netbraintech) - **Key Metric**: Annual Revenue of $95.0M (as of most recent data) [linkedin.com](https://www.linkedin.com/company/netbraintech) - **Total Funding**: $40.0M in a single venture round led by Summit Partners in 2014 [linkedin.com](https://www.linkedin.com/company/netbraintech) - **Notable Investors/Partners**: Summit Partners (prior investor), Blackstone Group (acquirer); integrates with major ITSM systems; talent sources include Cisco, Oracle, Dell Technologies, HPE, Juniper Networks [linkedin.com](https://www.linkedin.com/company/netbraintech) - **Growth Signals**: Serves approximately one-third of the Fortune 100 and Fortune 500; acquired by Blackstone for $750M in July 2025; operates in 10 countries with 9 offices globally; 331 employees with +1.4% monthly headcount growth; 28 active job postings with a +1300% quarterly increase in postings [linkedin.com](https://www.linkedin.com/company/netbraintech) [netbrain.com](https://www.netbrain.com/about/) ## Competitive Advantages - **Intent-based automation**: Manages networks from the top down, focusing on results rather than individual device details, which automates over 95% of network tickets and prevents over 50% of problems. - **AI-driven Agentic NetOps**: Uses AI agents that operate with full network context, enabling autonomous diagnosis and remediation. - **Deep enterprise integration**: Tightly integrates with major ITSM systems, making it a core part of existing IT workflows. - **Proven scale**: Trusted by one-third of the Fortune 100 and Fortune 500, demonstrating reliability in the most complex enterprise environments. ## Strategic Focus - **Agentic NetOps leadership**: Pioneering autonomous network operations through AI agents that diagnose, decide, and act. - **Hybrid-cloud expansion**: Continuing to support edge-to-cloud hybrid networks as enterprises adopt multi-cloud strategies. - **Proactive prevention**: Shifting from reactive ticket resolution to intent-based enforcement that prevents problems before they impact production. - **Post-acquisition growth**: Under Blackstone ownership, likely focused on scaling the platform and expanding market reach. ## Why Work Here - **Culture**: Emphasizes innovation, collaboration, and customer centricity. Employees describe a culture of passionate innovators with a desire to grow and learn [netbrain.com](https://www.netbrain.com/about/careers/). - **Work-life balance**: Prioritized as a core value within the company culture [netbrain.com](https://www.netbrain.com/about/careers/). - **Hybrid/Remote policy**: Most roles are hybrid (e.g., Burlington, MA; Toronto, ON; Hyderabad, India) with some fully remote positions (e.g., VP of Product, Strategic Technology Alliances Director, Sales roles) [greenhouse.io](http://job-boards.greenhouse.io/netbrain). - **Global presence**: Offices in Boston, London, Hyderabad, Beijing, and Toronto, offering opportunities to work across diverse geographies [netbrain.com](https://www.netbrain.com/about/). - **Benefits & Perks**: Competitive benefits, fun perks, support, and recognition aimed at supporting teammates and their families [netbrain.com](https://www.netbrain.com/about/careers/). - **Employee rating**: 4.3/5.0 on LinkedIn (242 reviews), with strong scores for work-life balance (4.1), compensation (4.1), culture (4.0), and career opportunities (4.0) [linkedin.com](https://www.linkedin.com/company/netbraintech). - **Team size**: Over 400 employees globally, with a diverse team spanning technical, sales, marketing, and customer success roles [netbrain.com](https://www.netbrain.com/about/careers/). - **Engineering culture**: Emphasis on innovation, collaboration, and customer-centricity; teams are encouraged to constantly innovate across all functions [netbrain.com](https://www.netbrain.com/about/). ## Sources 1. [netbrain.com](https://www.netbrain.com/about/) 2. [netbrain.com](https://www.netbrain.com/) 3. [greenhouse.io](http://job-boards.greenhouse.io/netbrain) 4. [netbrain.com](https://www.netbrain.com/about/careers/) 5. [linkedin.com](https://www.linkedin.com/company/netbraintech) ## Other roles at NetBrain - [AI Engineer](https://feeny.ai/job/ai-engineer-netbrain-burlington-erw57tvv8ftr) — Burlington, MA - [Strategic Account Executive (New York Metro)](https://feeny.ai/job/strategic-account-executive-new-york-metro-netbrain-new-york-0yqks5wkwt28) — New York, NY - [Senior Software Engineer, Customer Engineering](https://feeny.ai/job/senior-software-engineer-customer-engineering-netbrain-toronto-z40kfejjfgqg) — Toronto, Canada - [Marketing Operations Manager](https://feeny.ai/job/marketing-operations-manager-netbrain-united-states-pfdze2h0pagf) — United States - [Data Analytics Engineer](https://feeny.ai/job/data-analytics-engineer-netbrain-united-states-bvrrwhkxma2g) — United States - [Senior Security Compliance Analyst](https://feeny.ai/job/senior-security-compliance-analyst-netbrain-burlington-jcg52k0f3wez) — Burlington, MA - [Strategic Account Executive (Twin Cities)](https://feeny.ai/job/strategic-account-executive-twin-cities-netbrain-minneapolis-kj231qtybxec) — Minneapolis, MN - [Controller](https://feeny.ai/job/controller-netbrain-burlington-vzbb394yc3sh) — Burlington, MA - [Strategic Account Executive (Boston)](https://feeny.ai/job/strategic-account-executive-boston-netbrain-new-wr9thmagb3s4) — New, United Kingdom - [Associate Product Manager](https://feeny.ai/job/associate-product-manager-netbrain-toronto-0y1s5vhv164p) — Toronto, Canada