--- title: 'Senior ML Data Processing Developer at LawZero' canonical: 'https://feeny.ai/job/senior-ml-data-processing-developer-lawzero-montreal-5wp0w4r3spk7' type: 'job' last_seen: '2026-09-10' --- # Senior ML Data Processing Developer at LawZero - **Company:** LawZero - **Location:** Montréal, Canada - **Posted:** 2026-07-02 - **Last confirmed live:** 2026-09-10 - **Apply:** https://job-boards.greenhouse.io/lawzero/jobs/4305208009 ## Job description We are seeking a Senior ML Data Processing Developer to participate in the development, curation, and scaling of our core data asset pipeline. Sitting at the intersection of data engineering, data curation, and machine learning, you will own the end-to-end pipeline that transforms raw web-scale data into high-signal datasets used to train the Scientist AI. In this role, you will not just manage data; you will engineer its quality. You will design algorithmic filtering, build model-based scoring mechanics, and ensure rigorous benchmark integrity to power the next generation of AI. And as our models push beyond established paradigms, you will design and implement novel data transformations that don't yet have playbooks, working at the frontier of what training data can be. We are hiring multiple people for this role, and responsibilities may be distributed across the team based on individual experience, skills, and interests. ## Key Responsibilities - Partner with the Research team to define, build, automate, scale, and manage data pipelines that transform raw web-scale data into training datasets for the Scientist AI. - Build and maintain data processing pipelines, including deduplication, model-based quality scoring, heuristic filtering, toxicity removal, PII scrubbing, metadata extraction, and proprietary data transformations, with full dataset versioning and provenance tracking, optimizing for throughput and cost at scale. - Ensure all ingested data meets compliance requirements, internal Data Governance policies, and legal obligations. - Develop and refine the scoring and filtering toolchain: heuristics, LLM-as-a-judge evaluators, ML classifiers, metadata extraction modules, and human-in-the-loop review workflows required for data processing and quality assurance. - Instrument data processing pipelines with data-quality monitoring, guardrails, and alerting to catch regressions before they propagate downstream. - Collaborate with the Research team and other teams to understand evolving data requirements, then identify and acquire large-scale text corpora that meet those requirements. This includes conducting systematic coverage analyses to identify gaps in the corpus and develop targeted acquisition strategies to address them, and working with the Legal & Governance Team to license new data sources. - Design and maintain strict leakage detection mechanisms to guard against evaluation contamination across all stages of the data processing pipeline. - Build internal tooling and interfaces that let researchers explore, query, and understand available datasets with minimal friction. ## Skills and Qualifications - Degree in computer science, software engineering, or a related field. - Proven track record of handling massive unstructured text datasets (trillion-token scale), with 5+ years of experience in data processing, machine learning engineering or Natural Language Processing (NLP). - Hands-on experience with distributed processing frameworks (e.g., Spark, Ray, Flink), designing and optimizing high-throughput pipelines. - Experience with data privacy implementation (PII scrubbing), content-safety filtering (toxicity, bias), and evaluation-contamination prevention. - Demonstrated ability to work across Research, Engineering, and/or Legal/Governance teams, translating varied requirements into concrete pipeline work. - Strong Python proficiency, including experience writing production-grade data-processing code. - Experience with pipeline orchestration frameworks (e.g., Airflow, Prefect, Dagster). ## Nice to have - Experience training, fine-tuning, or deploying ML models for data-quality tasks (classifiers, LLM-based evaluators) and familiarity with LLM inference optimization (e.g. vLLM, SGLang). - Familiarity with containerized deployment (Docker, Kubernetes) and infrastructure-as-code practices. - Familiarity with ML experiment tracking tools (e.g. Weights and Biases). - Experience with data licensing workflows or web-scale data acquisition. - Contributions to open-source data processing or NLP tooling. ## What we offer - The chance to contribute meaningfully to a globally critical initiative - Comprehensive health benefits (including mental health and wellness management account) - 20 days of vacation per year upon start - Employer contribution of 4% to your retirement savings, with no required employee match - Additional compensation totaling 8% of your salary to apply towards additional retirement savings or bonuses (independent of group and individual performance) - A team of passionate world-class experts in their field - A collaborative and inclusive work environment in our vibrant office space in the heart of Little Italy, in the trendy Mile-Ex district, close to public transportation ## About LawZero LawZero is a non-profit organization committed to advancing research and creating technical solutions that enable safe-by-design AI systems. Its scientific direction is based on new research and methods proposed by Professor Yoshua Bengio, the most cited AI researcher in the world. Based in Montreal, LawZero’s research aims to build non-agentic AI that learns primarily to understand the world rather than to act in it, giving truthful answers to questions based on transparent and externalized probabilistic reasoning. Such AI systems could be used to accelerate scientific discovery, to provide oversight for agentic AI systems, and to advance the understanding of AI risks and how to avoid them. LawZero believes that AI should be cultivated as a global public good—developed and used safely towards human flourishing. For more information, visit [www.lawzero.org](https://www.lawzero.org/) You belong here At LawZero, diversity is important to us. We value a work environment that is fair, open and respectful of differences. We welcome applications from highly qualified individuals interested in working towards our mission in a respectful, inclusive and collaborative setting. Your personal information will be collected and processed by LawZero to evaluate your application for employment in compliance with our [Privacy Policy](https://lawzero.org/en/website-privacy-notice). Under privacy laws in force in your country of residence, you may have several privacy rights, such as to request access to your personal information or to request that your personal information be rectified or erased. Details on how you can exercise your rights can be found in our Privacy Policy. ## About LawZero ## Company Overview - **One-liner**: LawZero is a Canadian nonprofit research organization developing "Scientist AI," a non-agentic, safe-by-design artificial intelligence system. - **Entity Type**: Private (Nonprofit research organization; funded through philanthropic grants) - **Headquarters**: Montréal, Quebec, Canada - **Founded**: June 3, 2025 - **Founders**: Yoshua Bengio ## Core Business - **Primary industry**: AI Safety Research - **Target customers**: The organization is not a commercial entity; its research outputs are intended for the global public good, with potential downstream users including governments, AI companies, and scientific researchers. - **Mission or purpose**: To develop AI systems that are "safe by design" through technical research insulated from commercial and government pressures, cultivated as a global public good. ## Products & Services - **Scientist AI**: A novel, non-agentic AI system designed to understand the world through probabilistic reasoning, produce transparent and auditable predictions, and have no hidden goals or preferences. The system is built on three mechanisms: contextualization (separating facts from opinions), consequence invariance (preventing feedback from downstream outcomes), and a generator–estimator architecture (a creative generator held accountable by a neutral estimator). Intended uses include accelerating scientific discovery, providing guardrails for agentic AI systems, and advancing understanding of AI risks. **Type**: Research / Technical Framework ## Market Standing - **Valuation/Market Cap**: Not applicable (nonprofit) - **Key Metric**: Total Philanthropic Funding: > US$35 million (as of August 2025, including a grant from the Gates Foundation). An additional > CAD$100 million in financial backing from the Government of Canada was under discussion as of February 2026. - **Notable Investors/Partners**: Funders include Jaan Tallinn, Schmidt Sciences, Coefficient Giving, the Future of Life Institute, and the Gates Foundation. Operating partner is Mila – Quebec AI Institute. Affiliated with MATS (AI-alignment training program). - **Growth Signals**: Launched in June 2025 with ~15 researchers; grew to ~30 employees by February 2026, with stated plans to expand to more than 100 within a year. Secured significant government interest (letter of intent from the Government of Canada). Inaugural board and global advisory council announced in January 2026, including high-profile members such as historian Yuval Noah Harari and former Prime Minister of Sweden Stefan Löfven. ## Competitive Advantages - **Founder Reputation**: Founded by Yoshua Bengio, the world’s most cited living AI researcher and a Turing Prize winner, lending unparalleled scientific credibility. - **Mission-Driven Position**: Operates as a nonprofit, explicitly insulated from market and government pressures, which differentiates its research from major for-profit AI labs (e.g., OpenAI, Google DeepMind, Anthropic). - **Unique Technical Approach**: Its core concept, "Scientist AI," is a fundamentally different path to advanced AI—non-agentic and designed for safety from the ground up—rather than a modification of existing agentic systems. - **Compute Advantage**: A potential partnership with the Government of Canada is structured to provide substantial compute resources (the majority of expenditure is for compute), a critical and expensive resource for frontier AI research. ## Strategic Focus - **Research & Development**: The entire organization is focused on advancing the "Scientist AI" framework, with primary activities being theoretical research (as seen in arXiv preprints and white papers), algorithm development, and building the technical architecture. - **Talent Acquisition**: Aggressively hiring to scale the research team from ~30 to over 100 people within the next year. - **Global Governance**: Building a global advisory council and board to influence AI safety policy and standards, in addition to technical research. - **Scientific Application**: One concrete application area being explored is the use of Scientist AI to accelerate drug discovery and scientific research, supported by the Gates Foundation grant. ## Why Work Here - **Highly Mission-Driven**: A rare opportunity to work on fundamental AI safety research at a nonprofit, free from the commercial pressures of major AI labs. The explicit goal is to build AI as a global public good. - **World-Class Leadership**: Work directly under and alongside Yoshua Bengio and a leadership team with deep expertise in AI. - **Growth Trajectory**: The organization is in a rapid scaling phase (growing from 15 to 100+ people), offering early employees significant influence and rapid career growth. - **Prime Location**: Based in Montréal, a major AI research hub with access to talent from Mila and top universities. - **Policy Impact**: Involvement with a high-profile global advisory council that includes former heads of state, offering a chance to shape AI governance. - **Work Environment**: The culture emphasizes safety, mission integrity, and rigorous scientific research. The organization is structured as a research lab, not a commercial entity. ## Sources 1. [LawZero Official Website](https://lawzero.org/en) 2. [LawZero Careers Page (Greenhouse)](http://job-boards.greenhouse.io/lawzero) 3. [Wikipedia - LawZero](https://en.wikipedia.org/wiki/LawZero) 4. [LawZero Team Page](https://lawzero.org/en/team) 5. [LawZero Research Page](https://lawzero.org/en/research) ## Other roles at LawZero - [Senior Data Platform Engineer](https://feeny.ai/job/senior-data-platform-engineer-lawzero-montreal-x9384bvj4059) — Montréal, Canada - [Platform Engineer](https://feeny.ai/job/platform-engineer-lawzero-montreal-4bef2adt2wpn) — Montréal, Canada - [Senior Manager, Talent Acquisition](https://feeny.ai/job/senior-manager-talent-acquisition-lawzero-montreal-j5eqjx54dfe3) — Montréal, Canada - [Operations Admin Lead – Executive Office](https://feeny.ai/job/operations-admin-lead-executive-office-lawzero-montreal-fvvna9c8aper) — Montréal, Canada - [Senior Researcher Communications Specialist](https://feeny.ai/job/senior-researcher-communications-specialist-lawzero-montreal-70swcv9djdbd) — Montréal, Canada - [Machine Learning Manager](https://feeny.ai/job/machine-learning-manager-lawzero-montreal-ddxs7d0jnk2v) — Montréal, Canada - [Join our Talent Community (future opportunities)](https://feeny.ai/job/join-our-talent-community-future-opportunities-lawzero-multiple-a0gzf9abwev7) — Multiple - [Responsible AI & Data Governance Lead](https://feeny.ai/job/responsible-ai-data-governance-lead-lawzero-montreal-fj7e83h98pqa) — Montréal, Canada - [Mathematical Scientist for AI Safety Research](https://feeny.ai/job/mathematical-scientist-for-ai-safety-research-lawzero-montreal-kv2e1etp7wk0) — Montréal, Canada - [Director, Evaluations](https://feeny.ai/job/director-evaluations-lawzero-montreal-nhj2b35b8jy2) — Montréal, Canada