--- title: 'ML Lead, AI Data Labeling at NewtonX' canonical: 'https://feeny.ai/job/ml-lead-ai-data-labeling-newtonx-remote-2xy06953cq4m' type: 'job' last_seen: '2026-09-07' --- # ML Lead, AI Data Labeling at NewtonX - **Company:** NewtonX - **Location:** Remote - **Employment:** full-time - **Work type:** remote - **Posted:** 2026-05-18 - **Last confirmed live:** 2026-09-07 - **Apply:** https://jobs.ashbyhq.com/newtonx/1d43a564-38af-4020-857a-fce1f053b46f ## Job description ## About NewtonX NewtonX is a B2B insights company trusted by the world's most innovative companies to make high-stakes decisions with confidence. We combine a verified network of business professionals with AI-powered research tools to deliver research intelligence faster, more precise, and more defensible than traditional methods. Our clients include Google, Microsoft, TikTok, DoorDash, Stripe, and Coinbase. Our research has been cited by Fortune, Forbes, TechCrunch, Adweek, and the Wall Street Journal. NewtonX has raised $47M from investors including Two Sigma Ventures, Third Prime, XFund, and Citi Ventures. ## About the Role NewtonX is rapidly expanding into the AI data annotation and RLHF (Reinforcement Learning from Human Feedback) space. We are leveraging our core superpower—recruiting the world’s leading domain experts—to provide high-quality, expert-led data labeling for AI labs and enterprises. Beyond our unmatched B2B recruiting, we utilize powerful, automated project management processes that allow us to scale projects rapidly, adapt to shifting requirements, and manage our subject matter experts with professional, industry-best practices and fair compensation. This role is the technical owner of data quality. You will partner directly with the Program Lead and act as technical lead in communicating with clients, interpreting their core AI model testing goals and assisting the Program Lead in creating concrete technical specs that will accomplish these goals. The core question you own: will the data we produce do its job once the customer trains on it or evaluates with it? Clean data that passes every operational gate can still fail — a training set that yields a weak or misleading signal, or an eval that technically runs but doesn't surface the weaknesses that matter. You are the person who can look at an approved submission and say "this is technically correct and still won't do the job, and here's why." You are hands-on and close to the work. This is a foundational role — what it covers will grow as the business does. In this role you will: Training Signal Integrity (core) - Own the judgment of whether task designs and rubrics produce a useful training signal for the consuming method (SFT, RLHF, RLVR, agentic RL, CoT, evals). Catch mismatches between what the data rewards and what the customer is actually training for. - Design the task, environment, and rubric structure for agentic workflows. - Own the ML-level validation of the dataset before delivery — not just whether individual submissions meet spec. Technical Feedback Loop with Operations - Partner with the Program Lead to convert customer requirements into concrete technical specs: expert profiles, screener trees, task interfaces, task templates, QC rubrics, statistical thresholds. - Partner with the Program Lead to define and defend quality metrics: inter-annotator agreement targets, gold-standard injection rates, statistical power thresholds — providing the statistical and methodological grounding. - Calibrate the ops team on what "good" looks like per engagement; run alignment sessions when standards shift. Customer Technical Credibility - Serve as the technical counterpart to the customer's ML, applied science, and product teams. Hold your ground on the technical questions that decide data quality across training and evaluation — task and reward design, data quality for SFT and preference methods, RL and RLVR, contamination, statistical rigor, and agentic workflows. - Help diagnose data-quality questions when a customer reports the data underperformed — reason through whether the issue is the data, the quantity, or something in their training setup, and make the case defensibly. ## Who you are - Deep applied ML experience centered on post-training human data — you've owned a human-data workstream as an applied scientist or ML engineer. - At least 2 years working hands-on with RL, including how tool-use trajectories are rewarded and evaluated. If you're not fluent in RL, this isn't the role — it's the foundation of the core judgment you'll be making. - Working fluency across modern LLM post-training and evaluation: SFT, RLHF/preference data quality, RLVR, chain-of-thought, eval harness construction, contamination handling, statistical significance, and agentic/tool-use evaluation. - Genuine understanding of how training data becomes model behavior — you can reason about what a model will learn from a given dataset, not just whether the data meets spec. - Strong programming foundation: read and reason about an eval harness, write Python comfortably, work with model APIs, prototype scoring pipelines. Not a production engineer, but not hands-off. - Statistical fluency: you know when an effect is real vs. noise and can defend a sample size or significance threshold. - Client-facing presence: you've defended technical design choices in real time to skeptical audiences and adjusted scope without losing rigor. Range matters — you can talk to a Series B CTO and a Fortune 100 AI lead in the same week. - Strong written communication: methodology sections, technical reports, and specs that hold up to expert review. If the profile above describes you and your passions, we'd love to hear from you! ## What we offer - Massive Impact: Opportunity to have an astounding impact, build a brand new business unit from the ground up, and have direct C-level influence at an extremely fast-growing late-stage startup. - Fast-track career growth: This foundational role will enable you to progress quickly within NewtonX towards commercial and operational leadership. - Comprehensive Benefits: Excellent medical, dental, and vision insurance. - Retirement: 401k match with immediate vesting. - Perks: Health savings/flexible savings account, and pre-tax commuter benefits. - Work-Life Balance: Paid time off: vacation, holidays, sick, and parental leave. - Great Culture: A diverse, collaborative, and positive culture where we invest in and celebrate each other's success (happy hours, team projects, and retreats). - Visa sponsorship is not available for this role. NewtonX is proud to be an equal opportunity workplace. We do not discriminate based upon race, religion, color, national origin, sex, sexual orientation, gender identity/expression, age, status as a protected veteran, status as an individual with a disability, or any other applicable legally protected characteristics. ## About NewtonX ## Company Overview - **One-liner**: NewtonX is an end-to-end B2B market research platform that combines AI‑enabled collection and analysis with a proprietary community of verified professionals to power confident, mission‑critical decisions. - **Entity Type**: Private (Series B) - **Headquarters**: New York, New York, United States - **Founded**: 2017 - **Founders**: Sascha Eder (CEO & Co‑founder) ## Core Business - **Primary industry**: Market Research / B2B Insights - **Target customers**: Fortune 500 enterprises, MBB consulting firms (McKinsey, BCG), top market research agencies, and private equity / venture capital firms. - **Mission statement**: “We’re on a mission to uncover world‑class knowledge.” ## Products & Services - **NewtonX Platform**: AI‑enabled end‑to‑end market research platform that automates expert recruitment, survey design, data collection, and analysis. Includes automated analysis and visualization for instant insights. - **Hub Researcher**: An AI‑powered research analyst tool that provides continuous, context‑aware insights for high‑stakes B2B research (new product). - **Custom Research Services**: Bespoke qualitative and quantitative research programs, including product research, market opportunity assessment, customer research, brand & comms research, and international expansion studies. ## Market Standing - **Valuation / Market Cap**: Not publicly disclosed - **Key Metric**: Annual revenue of **$30M** (LinkedIn estimate); total funding of **$47M** - **Notable Investors / Partners**: - **Series B (2021)**: $32M led by Marbruck - **Series A (2019)**: $12M led by Two Sigma Ventures - **Seed (2017)**: $3M led by Third Prime - **Non‑equity assistance**: Grand Central Tech - **Acquisition**: Merlin Guides (2019) - **Growth Signals**: - Headcount grew **24.2% YoY** to 266 employees (LinkedIn) - LinkedIn follower growth of **39.6% YoY** - Operates in **30 countries** with a global, multilingual team (12+ languages spoken) - 97% of expert requests are feasible; every expert is 100% verified ## Competitive Advantages - **Precision expert targeting**: Proprietary community of the world’s most elusive and verified B2B professionals. - **AI‑enabled speed**: Automated collection and analysis deliver insights in days, not weeks. - **High data quality**: SOC 2, ESOMAR, Insights Association, MRS, and BBB DPF certified. - **Track record**: Powers insights for Fortune 500, MBB firms, and top market research companies. ## Strategic Focus - **AI & automation**: Scaling AI‑driven research tools (Hub Researcher) to replace slow, manual processes. - **International expansion**: Growing presence in EMEA, LATAM, and APAC. - **Product innovation**: Building a research platform that combines human expertise with machine intelligence to reduce time‑to‑insight. ## Why Work Here - **Culture**: Emphasizes learning, regular feedback, career path discussions, and development programs. Values include innovation, integrity, ownership, and community. - **Remote / hybrid**: The company is based in New York with a global remote‑friendly culture; many roles list “Remote” (Indeed). Over half the team was born or has lived internationally. - **Benefits & perks**: - Medical, dental, vision insurance - 401K with 3% match, immediate vesting - Paid vacation, public holidays, sick days - Pre‑tax commuter benefits; HSA / FSA - Paid parental / family leave - Office snacks, lunch & learns, monthly outings, bimonthly happy hours - Annual company retreat, volunteering, virtual social activities - **Engineering culture**: Strong technical leadership (CTO Jeremy Wood formerly co‑founded OpenStore, was engineer at Google). Active hiring for Staff Software Engineers, ML Lead, Senior Data Engineer/Analyst. - **Employee rating**: 4.0/5 on LinkedIn (167 reviews); Culture 4.1, Career 4.0, Compensation 3.9, Work‑Life 3.8. ## Sources 1. [newtonx.com](https://www.newtonx.com/) 2. [newtonx.com/careers](https://www.newtonx.com/careers/) 3. [newtonx.com/team](https://www.newtonx.com/team/) 4. [linkedin.com/company/newtonx](https://www.linkedin.com/company/newtonx) 5. [indeed.com/cmp/Newtonx-1](https://www.indeed.com/cmp/Newtonx-1) 6. [jobs.ashbyhq.com/newtonx](https://jobs.ashbyhq.com/newtonx) ## Other roles at NewtonX - [Project Lead, AI Model Training](https://feeny.ai/job/project-lead-ai-model-training-newtonx-remote-8z500skpa6p7) - [Accelerated Growth Lead](https://feeny.ai/job/accelerated-growth-lead-newtonx-remote-camat19yke4d) - [Account Executive (New Logo)](https://feeny.ai/job/account-executive-new-logo-newtonx-remote-7yrgzn5y9055) - [Senior Frontend React Engineer](https://feeny.ai/job/senior-frontend-react-engineer-newtonx-medellin-q123g2p70g5f) — Medellín, Colombia - [Senior Backend Python Engineer](https://feeny.ai/job/senior-backend-python-engineer-newtonx-medellin-m1vv76s85gyz) — Medellín, Colombia - [Senior Sales Manager, EMEA](https://feeny.ai/job/senior-sales-manager-emea-newtonx-london-g0xkbbrrs8vq) — London, United Kingdom - [Senior Sales Manager, US Consulting Practice](https://feeny.ai/job/senior-sales-manager-us-consulting-practice-newtonx-new-york-ts0zr781hb0b) — New York, NY - [Client Delivery Associate](https://feeny.ai/job/client-delivery-associate-newtonx-new-york-tcmqq10ssm1a) — New York, NY - [Software Engineer - LLM Systems](https://feeny.ai/job/software-engineer-llm-systems-newtonx-remote-f24qbeceeke5) - [Staff Software Engineer](https://feeny.ai/job/staff-software-engineer-newtonx-new-york-xphx5rztr8em) — New York, NY