--- title: 'Data Engineer at Matterworks, Inc.' canonical: 'https://feeny.ai/job/data-engineer-matterworks-inc-somerville-4x8pr0kjgk7n' type: 'job' last_seen: '2026-09-13' --- # Data Engineer at Matterworks, Inc. - **Company:** Matterworks, Inc. - **Location:** Somerville, MA - **Employment:** full-time - **Work type:** hybrid - **Posted:** 2026-08-05 - **Last confirmed live:** 2026-09-13 - **Apply:** https://jobs.ashbyhq.com/matterworks/02f5d512-49f7-43c7-b989-b9ddd1b43923 ## Job description ## About Us Most of the molecules driving human biology are invisible to us. Mass spectrometers already detect metabolites, lipids, and peptides, but the vast majority of those signals never get identified. A typical experiment names a small fraction of its features and discards the rest. We call this biology's dark matter. It’s signal-rich and mechanism-defining, yet almost entirely opaque. Matterworks is building the foundation models that make that dark matter legible. Our Large Spectral Models do for biochemical biology what AlphaFold and ESM did for proteins: turning a library-bound discipline into something predictable and generative, and embedding it at every stage of R&D. Come build the future of biological discovery with us. Position Overview Matterworks is seeking a Data Engineer to build and run the pipelines behind our models. Our platform acquires mass spectrometry and molecular data at scale and turns it into the datasets our AI team trains on and our product reasons over. You will own real pieces of that path end to end. Our data serves two different customers. The AI team needs training corpora that are complete, correctly split, and reproducible. The product and agentic layer needs values a scientist can explain and we can safely show a customer. You will build inside the contracts and checks that keep both honest, and you will help extend them. This is a hands-on role with a clear growth path. You will start by owning well-scoped pipelines and datasets and grow toward owning larger parts of the platform. You will report to the Head of Engineering and work daily with our machine learning researchers, scientists, and product team. ## Key Responsibilities - Build and Operate Pipelines: Own well-scoped pipelines end to end: designed, tested, instrumented, documented, and running on a schedule. - Labels and Enrichment: Turn raw data into datasets people can actually use, with consistent schemas, trustworthy metadata, and documented definitions. - Data Quality Checks: Design, build and extend our quality checks so that each build gets compared against the last one before it publishes, and a failure stops the pipeline instead of shipping. - Ingest and Acquisition: Bring new public and partner datasets into the platform: fetching, converting, validating, and reconciling them against what we already hold. Expect messy scientific and vendor formats and file that require continuous improvements to our systems to handle at scale. ## About You - 2+ years of professional experience building data pipelines in production. - Proficient in Python and SQL. - Working knowledge of cloud data infrastructure. We run Argo Workflows and Metaflow on EKS, Glue and Athena over Apache Iceberg and Parquet, DuckDB, and Terraform. Depth in any comparable stack transfers fine. - Demonstrated experience owning a pipeline or dataset end to end, including the tests, the monitoring, and the failures. - Comfort with messy data and messy formats, and the patience to track down why two sources disagree. - Daily use of AI coding tools, paired with healthy skepticism about their output on questions of production data correctness. - Clear written communication, particularly when explaining what broke and what you changed. - Curiosity about the science. Experience in life sciences, biotechnology, or biochemistry is a plus but not a requirement, and you will work alongside strong in-house chemistry every day. - A passion for contributing to an early-stage startup where autonomy, eagerness to learn, and enthusiasm for solving novel scientific challenges prevail over rigid processes and egos. Working at Matterworks Given the cross-disciplinary and innovative nature of our work, effective collaboration and communication are critical to our progress. We operate in a flexible hybrid model that accommodates both fully remote team members and those who work full-time from our Somerville, MA office. While some positions may require regular in-person presence for hands-on work or local collaboration, many roles can be performed remotely with team members distributed across various locations. ## Compensation and Benefits Matterworks offers full-time employees a competitive base salary, stock options, and benefits (health & dental, vision, long- and short-term disability, life insurance, 401k with company match). Employees enjoy a flexible work & unlimited time away policy, commuter benefits and parking, regular team meals and outings, and company support for continued education/coursework and conference participation. Matterworks, Inc. is an equal opportunity employer. All candidates for employment at Matterworks are considered without regard to race, color, religion, national origin, age, sex, marital status, ancestry, physical or mental disability, veteran status, gender identity, sexual orientation, or any other category protected by law. ## About Matterworks, Inc. ## Company Overview - **One-liner**: Matterworks builds AI foundation models for biochemical omics, enabling predictive biology across the life sciences R&D pipeline. - **Entity Type**: Private (Series A) - **Headquarters**: Somerville, Massachusetts, United States - **Founded**: 2019 - **Founders**: Jack (J.M.) Geremia, PhD (CEO/Co-Founder) ## Core Business - **Primary industry**: Biotechnology, AI-driven life sciences research tools - **Target customers**: B2B – pharmaceutical and biotechnology companies, CROs, academic research labs - **Mission or purpose statement**: “Making biology a science of prediction” (LinkedIn company description) ## Products & Services - **Large Spectral Model (LSM)**: A proprietary frontier AI model trained on billions of small molecule, lipid, peptide, and protein spectra. It enables untargeted biochemical annotation, absolute quantitation, and direct phenotype prediction from raw mass spectrometry data. - **Pyxis Platform**: Cloud-based AI platform that guides users through analysis with a predictive omics assistant. Supports self-service onboarding and enterprise-grade integration. - **Application Modules**: Targeted solutions for Asset & Target Discovery, Toxicology & Safety, Clinical Development, and CMC & Biomanufacturing (including bioprocess optimization and cell line development). ## Market Standing - **Valuation/Market Cap**: Not publicly disclosed - **Total Funding**: USD $13.75 million (across Pre-Seed, Seed, Convertible Note, and Series A rounds) - **Notable Investors/Partners**: OMX Ventures (Series A lead), Lewis & Clark Capital, Germin8 Ventures, Intermountain Ventures, Pillar, and others. Scientific advisory board includes former Genentech SVP Andrew C. Chan and venture partner Steve Hitchcock of 5AM Ventures. - **Growth Signals**: 45 employees (+29.4% YoY), LinkedIn follower growth +75.3% yearly, recent hire of Paul Stone as CFO/Head of Corporate Development (Feb 2026), 4 active job postings (Product Designer, Product Management Intern), 6 patents filed. ## Competitive Advantages - Proprietary Large Spectral Model covers billions of spectra across millions of biological contexts – a breadth unmatched by conventional analysis tools. - Ability to predict complex biological outcomes (efficacy, toxicity, clinical phenotypes) directly from raw mass spec data, including absolute quantitation for untargeted biomolecules. - Multiple validated use cases across the full R&D lifecycle, from target discovery to bioprocess manufacturing. - Broad access offered via self-service free tier, enabling rapid adoption and feedback loops. ## Strategic Focus - Scale commercial and partnership development with pharma and biotech companies. - Deepen the Pyxis platform with additional predictive models and application modules (e.g., generative discovery, bioprocess optimization). - Grow the team in product design, engineering, and management to support expanding enterprise customer base. ## Why Work Here - **Culture**: Collaborative, cross-functional environment with all-hands lunches, happy hours, and team offsite adventures. Emphasis on camaraderie and shared scientific purpose. - **Benefits**: Health/vision/dental insurance, stock options, unlimited PTO, 401(k) matching, parental leave, HSA, commuter benefits, learning & conference budget. - **Work Policy**: Not explicitly stated; likely hybrid/flexible given office presence in Somerville, MA, plus satellite offices in Canada and Singapore. Open roles appear to be based in the United States. - **Mission**: Opportunity to work at the cutting edge of AI and biology, building tools that directly accelerate life science innovation. ## Sources 1. [matterworks.ai](https://www.matterworks.ai/) 2. [matterworks.ai/company](https://www.matterworks.ai/company) 3. [matterworks.ai/careers](https://www.matterworks.ai/careers) 4. [linkedin.com/company/matterworksbio](https://www.linkedin.com/company/matterworksbio) 5. [cbinsights.com/company/matterworks](https://www.cbinsights.com/company/matterworks) ## Other roles at Matterworks, Inc. - [Senior Software Engineer, Data Platform](https://feeny.ai/job/senior-software-engineer-data-platform-matterworks-inc-somerville-3rxsw2gw3jgw) — Somerville, MA - [Future Opportunities](https://feeny.ai/job/future-opportunities-matterworks-inc-somerville-vpm9ysh7637k) — Somerville, MA - [Data Engineer](https://feeny.ai/job/data-engineer-siteminder-pune-bmcc0yj3ttby) — Pune, India - [Data Engineer](https://feeny.ai/job/data-engineer-future-processing-gliwice-f8cp0tctfjmd) — Gliwice, Poland - [Data Engineer](https://feeny.ai/job/data-engineer-xcimer-energy-denver-nf6k4sr1dzcn) — Denver, CO - [Data Engineer](https://feeny.ai/job/data-engineer-scott-logic-bristol-2x8fz0r00jjn) — Bristol, United Kingdom - [Data Engineer](https://feeny.ai/job/data-engineer-oxylabs-vilnius-zaj01zcyf0kq) — Vilnius, Lithuania - [Data Engineer](https://feeny.ai/job/data-engineer-moniepoint-poland-s1fvb5j64pw7) — Poland - [Data Engineer](https://feeny.ai/job/data-engineer-wpp-warsaw-dcxnw1rg6eeh) — Warsaw, Poland - [Data Engineer](https://feeny.ai/job/data-engineer-centerfield-los-angeles-dea9e54mc8jk) — Los Angeles, CA