Boson AI

Datacenter Technician at Boson AI (Barrie, Canada)

Boson AI· Barrie, Canada· $50k–$100k·

Role details

Salary
$50k–$100k
Work type
Onsite
Employment
Full-Time

Job description

Boson AI is an early-stage startup building large language tools for everyone to use. Our founders (Alex Smola, Mu Li), and a team of Deep Learning, Optimization, NLP, AutoML and Statistics scientists and engineers are working on high quality generative AI models for language and beyond.

You keep the machines running

We are looking for a Datacenter Hardware Technician to keep the physical infrastructure behind our AI research running in Barrie, ON. You will work hands-on with the latest NVIDIA GPUs, thousands of disks, terabit networking and hundreds of Supermicro servers — racking them, repairing them, keeping firmware current, and making sure every machine that should be online is online.

Our researchers train models around the clock, so the difference between a good day and a bad one is often a technician who noticed a marginal cable, logged a serial number correctly, or caught a failing drive before it took a training job down with it. This is careful, methodical work, and we treat it that way.

This role is for our Barrie datacenter. You must live within 50 km of the site, such that you can be onsite on short notice. Workload can be bursty, i.e. periods of smooth sailing mixed with periods of intense work during hardware failures, upgrade and maintenance cycles.

A day in the life

  • Install, rack, cable and commission new Supermicro servers, storage and network equipment.
  • Diagnose and repair hardware failures — replace DIMMs, drives, power supplies, fans, GPUs, cables and mainboards — and drive RMA cases with vendors through to resolution.
  • Install and update drivers, BIOS and firmware across servers, NICs, HBAs and switches, keeping fleet versions consistent and documented.
  • Test and install network connections, including structured cabling, optics and link validation, and troubleshoot physical-layer faults.
  • Perform preventive maintenance: inspections, cable management, airflow and filter checks, and spare-parts inventory.
  • Keep accurate records of every asset, serial number, part swap and rack location.
  • Respond to hardware failure alerts, and escalate to the SRE team when a fault is not purely physical.
  • Follow runbooks and ESD and safety procedures precisely — and improve them where they are unclear.

If you take pride in a tidy rack, a clean cable run, and a fleet where every machine is accounted for, we'd love to hear from you.

Why work at Boson AI

  • Culture: Described as “a diverse group of researchers, engineers, and industry specialists united by a passion for innovation.” Emphasis on building scalable AI that serves millions.
  • Work policy: Most roles are on-site at Santa Clara HQ or Toronto office. A Site Reliability Engineer role was listed as remote (Toronto). Likely hybrid/on-site for core engineering.
  • Engineering environment: Deep technical stack (PyTorch, TensorFlow, Kubernetes, NVIDIA, Supermicro, etc.). Opportunity to work on frontier AI models and real-time systems.
  • Growth stage: Small team (~32 people) with strong research roots; employees have previously worked at AWS, Google, Robinhood, and alumni go to OpenAI, Anthropic, xAI – indicating high-caliber talent and career mobility.
  • Hiring process: Uses AI tools to assist with resume review and analysis, but final decisions made by humans.

Application questions