--- title: 'Lead Data Scientist at Rossum' canonical: 'https://feeny.ai/job/lead-data-scientist-rossum-prague-p0khcgp6zqcs' type: 'job' last_seen: '2026-09-06' --- # Lead Data Scientist at Rossum - **Company:** Rossum - **Location:** Prague, Czech Republic - **Employment:** full-time - **Posted:** 2026-09-01 - **Last confirmed live:** 2026-09-06 - **Apply:** https://jobs.ashbyhq.com/rossum.ai/43eccd98-3cff-4e87-bc57-d5c0efd74396 ## Job description ## ABOUT COUPA Coupa is the platform companies run their spending on - sourcing, procurement, invoices, payments, suppliers, contracts. It is the system of record for how large organisations decide what to buy, from whom, and at what price. That makes it one of the most unusual datasets in enterprise software: a single network of 10M+ buyers and suppliers, and $10 trillion of transacted spend to date - quotes, bids, awards, orders, invoices, contracts, and the documents behind every one of them. Multimodal, longitudinal, and tied to outcomes measured in real money. ## ABOUT THE TEAM Rossum joined Coupa earlier this year. We brought the document understanding layer - our proprietary T-LLM (transactional LLM) architectures, which we design and train from scratch, and which read the world's messiest business documents in production, millions of them every week. Now we are pointing the same in-house research capability at a much bigger problem: not just reading the documents, but acting on them. ## ABOUT THE ROLE SOURCING IS WHERE THE MONEY IS ACTUALLY DECIDED. We are expanding our Data Capture Research team in Prague with a Senior Data Scientist to work on Sourcing. Which suppliers get invited. How the event is structured. How bids that differ in price, lead time, quality, risk and carbon get compared at all. What a fair price even is. When to award, to whom, and how to split the award across suppliers. For a researcher this is unusually open ground - not one model family, but several, on the same data: - Recommendation and retrieval. Supplier discovery: matching demand to the right suppliers across a 10M-node network. - Forecasting and should-cost modelling. What this category, in this region, at this volume, should cost right now. - Game theory and mechanism design. Auction formats, bidding behaviour, incentives, and competitive dynamics between real counterparties. - Combinatorial optimisation. Award allocation under volume, capacity and multi-sourcing constraints. - Multimodal document understanding. RFPs, specs, quotes and contracts carry the actual requirements - and our T-LLM foundations already give us a head start there. Part of the work is choosing the right instrument for each - neural networks, gradient boosting, optimisation, bandits, mechanism design - instead of forcing one. Very little of this has been built with modern ML yet. That is the point of the role: real greenfield problems, a dataset nobody else has, and a product that ships to companies whose margins depend on getting these decisions right. You will work in a small, senior team of researchers and engineers - the group that built Rossum's production models from scratch - with direct access to Product and the AI Platform team. Ideas that work do not sit on paper; they roll into systems used at scale. ## WHAT YOU'LL DO - Own sourcing research initiatives end to end - from framing the problem on real spend data, through experiments, to a model running in production. - Build what does not exist yet. Most of these problems have no baseline to beat and no off-the-shelf answer - you decide what the first version looks like. - Turn the network into a training set. Define the datasets, labels and benchmarks that make sourcing problems learnable at all: deriving supervision from historical events and their outcomes, and building evaluation the team can trust. - Design and run bold experiments with a hacker mindset - fast prototypes, honest baselines and offline evaluation that actually predicts online behaviour. - Build on our T-LLM foundations wherever documents carry the signal - RFPs, specs, quotes, contracts - reusing in-house architectures we train ourselves. - Ship with deployment in mind: inference cost, latency, robustness, and what happens when a recommendation is wrong in front of a buyer. - Work across the company - Coupa's Sourcing product and data teams and our AI Platform team. Document decisions and make the people around you better. ## WHO YOU ARE Sourcing needs several kinds of modelling, so we are deliberately open about which one you bring. - 5+ years in applied ML, data science, ML engineering or quantitative research, with models or decision systems you took into production. - Strong Python and real comfort with messy, large-scale data - including SQL and the unglamorous work of making a dataset trustworthy. - Depth in at least one modelling discipline, curiosity about the rest. Deep learning, recommendation and ranking, forecasting and time series, optimisation and operations research, causal inference and econometrics, RL and bandits, market and mechanism design, or LLM-based systems. Sourcing touches most of these; nobody arrives holding all of them. - Scientific rigour. Strong experiment design, healthy scepticism about your own metrics, and real care about leakage, baselines, and evaluation that survives contact with production. - Ownership and curiosity. You are comfortable in a greenfield, ambiguous problem space, and you will talk to product people and procurement experts to find where the value actually is. - Interest in procurement, supply chains or market design is welcome but not required - we will teach the domain. ## WHY JOIN US - Greenfield problems in a mature product: Modern ML has barely been applied to sourcing, inside a platform that already has the users, the workflows and the data. - Data nobody else has: $10 trillion of transacted spend, 10M+ buyers and suppliers, and the documents behind all of it. - We train our own models: Proprietary T-LLM architectures, designed and trained in-house - not a wrapper around someone else's API. - Real ownership, short path to customers: You frame the problem, choose the method, and see it working in front of buyers - no research-to-product handoff. - Global impact: Technology used every day by companies around the world. - Experiment-driven culture: Pragmatic delivery, and quarterly recognition for standout research contributions. - Compute and tools: Frontier LLMs on tap for your own work, and our high-end GPU and large-memory clusters to train on. - Conferences: A budget to attend the ones that matter in your field. - 33 days off: PTO, personal days, your birthday and two company wellness days. Parental leave on top. - Prague, Karlín: Inspiring workspace and full tech setup, including a 200 m² terrace with views of Prague Castle. ## About Rossum ## Company Overview - **One-liner**: Rossum provides an AI-powered intelligent document processing platform that automates transactional document workflows, such as invoices and purchase orders. - **Entity Type**: Private (Series A; $105.5M total funding) - **Headquarters**: London, United Kingdom - **Founded**: 2017 - **Founders**: Tomáš Gogár, Petr Baudiš, Tomáš Tunys ## Core Business - **Primary industry/industries**: Intelligent Document Processing (IDP), Enterprise Automation, Artificial Intelligence - **Target customers**: Enterprise businesses across finance, accounting, and supply chain departments (B2B) - **Mission or purpose statement**: "To enable one person to effortlessly process one million transactions from start to finish, in a year." ## Products & Services - **Rossum Aurora (AI Platform)**: The core cloud-native platform powered by a proprietary Transactional Large Language Model (LLM). It ingests documents via email, scanners, PEPPOL, and shared drives; extracts data with high accuracy; validates and augments data against ERPs and business rules; and triggers automated communications and approvals. - **AI Agents**: Autonomous agents that read documents, capture, validate, transform data, send emails, ask for approval, and write data to ERPs, all in line with standard operating procedures. - **Integrations**: Offers turnkey integrations with SAP, NetSuite, Coupa, Workday, and other major enterprise systems. ## Market Standing - **Valuation/Market Cap**: Not publicly available - **Key Metric**: Total Funding of $105.5M (6 funding rounds); Annual Revenue estimated at $8.0M - **Notable Investors/Partners**: Not disclosed in search results, but company has 200+ partners and integrations. - **Growth Signals**: 200+ employees across 5 global offices (Prague, US, UK, Luxembourg); 450+ customers; recognized as a Leader in the Everest Group Intelligent Document Processing PEAK Assessment 2026; named in the Forrester Wave™: Document Mining And Analytics Platforms, Q2 2024; included in Sifted B2B SaaS Rising 100 (2024); employee headcount grew 12.5% YoY. ## Competitive Advantages - **Proprietary Transactional LLM**: A specialized AI model designed for transactional documents, supporting 276 languages and handwriting, with zero hallucinations and continuous learning from user feedback. - **End-to-End Automation**: Goes beyond simple data capture to include validation, augmentation, workflow approval, automated communications, and direct ERP integration. - **Enterprise-Grade Security**: ISO 27001 certified, HIPAA compliant, with 99.9% uptime SLA and 24/7 support. - **Industry Recognition**: Named a Leader by Everest Group and a Strong Performer by Forrester and IDC MarketScape. ## Strategic Focus - **AI-First Innovation**: Continuously raising the performance bar by capitalizing on the latest advancements in LLMs and Gen AI. - **Global Expansion**: Growing presence in the US, UK, and Luxembourg from its Czech roots. - **Customer Value**: Focus on reducing document chaos and unlocking real-time strategic insights from transactional data for enterprise clients. ## Why Work Here - **Culture**: Values-driven, with a focus on "Strong opinions, weakly held" and data-driven decision-making. Emphasizes autonomy, transparency, and learning from failure. - **Work Model**: Flexible hybrid model (majority of employees prefer this). Headquarters in Prague (Karlín), with offices in the US and remote workers globally. - **Perks**: Stock options, generous holidays, high-end gear, and a commitment to employee well-being. - **Engineering Culture**: Fast-paced, inclusive, and innovative. Employees are empowered to experiment and share learnings across teams. - **Hiring Process**: Transparent and efficient (typically 2-3 weeks), with skills assessments, deep interviews, and a focus on cultural fit. The internal language is English. - **Career Growth**: Strong emphasis on internal mobility and progression; managers actively support career development. ## Sources 1. [rossum.ai/company/](https://rossum.ai/company/) 2. [rossum.ai/](https://rossum.ai/) 3. [rossum.ai/careers/this-is-rossum/](https://rossum.ai/careers/this-is-rossum/) 4. [rossum.ai/careers/how-we-hire/](https://rossum.ai/careers/how-we-hire/) 5. [uk.linkedin.com/company/rossum](https://uk.linkedin.com/company/rossum) ## Other roles at Rossum - [Senior Data Scientist](https://feeny.ai/job/senior-data-scientist-rossum-prague-8xcjw7sexjkr) — Prague, Czech Republic - [Senior AI Engineer](https://feeny.ai/job/senior-ai-engineer-rossum-prague-3xakzcpvd505) — Prague, Czech Republic - [Lead AI Engineer](https://feeny.ai/job/lead-ai-engineer-rossum-prague-nd9gggfa0znc) — Prague, Czech Republic - [Senior Solution Architect](https://feeny.ai/job/senior-solution-architect-rossum-prague-xekwmmmrvs13) — Prague, Czech Republic - [Principal Solution Architect](https://feeny.ai/job/principal-solution-architect-rossum-prague-pab2her259h6) — Prague, Czech Republic - [Lead Data Scientist](https://feeny.ai/job/lead-data-scientist-faculty-london-3w05crhh4hvh) — London, United Kingdom - [Lead Data Scientist](https://feeny.ai/job/lead-data-scientist-netspend-austin-21vct1zspfsm) — Austin, TX - [Lead Data Scientist](https://feeny.ai/job/lead-data-scientist-gymshark-solihull-england-j9mkrhcem96g) — Solihull England, United Kingdom - [Lead Data Scientist](https://feeny.ai/job/lead-data-scientist-onos-health-san-francisco-yyt0hr2bx33g) — San Francisco, CA - [Lead Data Scientist](https://feeny.ai/job/lead-data-scientist-bondora-finland-latvia-spain-tallinn-estonia-emj8m763hhqn) — Finland / Latvia / Spain / Tallinn, Estonia