--- title: 'Research Scientist – Large Tabular Models (LTMs) at Granica' canonical: 'https://feeny.ai/job/research-scientist-large-tabular-models-ltms-granica-bay-area-2qa1k66qnecd' type: 'job' last_seen: '2026-09-08' --- # Research Scientist – Large Tabular Models (LTMs) at Granica - **Company:** Granica - **Location:** Bay Area - **Employment:** full-time - **Work type:** onsite - **Posted:** 2026-07-02 - **Last confirmed live:** 2026-09-08 - **Apply:** https://jobs.ashbyhq.com/granica/53b843e5-0d00-4595-8283-f6590fadab61 ## Job description Location: Mountain View, CA (On-site) ## OVERVIEW Most of today's generative AI is built for text, images, and video. Enterprise data isn't. The world's most valuable data lives in tables: customer records, transactions, financial systems, telemetry, operational data, and business workflows. Today's generative AI stack wasn't designed to learn efficiently from this kind of information. At Granica, we're building Large Tabular Models (LTMs)—foundation models that learn natively from structured and relational enterprise data. Our research is led by Prof. Andrea Montanari (Stanford) and focuses on one central question: How can we build generative AI that learns efficiently from tabular data? That requires solving problems well beyond model architecture, including intelligent data selection, dataset augmentation, representation learning, and information-preserving compression. If you're excited about inventing the algorithms that make Large Tabular Models possible, we'd love to talk. ## WHAT YOU'LL WORK ON - Develop new machine learning algorithms for Large Tabular Models. - Research methods for selecting, augmenting, and compressing training data without losing information. - Build representation learning techniques for structured and relational datasets. - Prototype and evaluate new approaches for generative modeling over enterprise data. - Design rigorous experiments and benchmarks to measure progress. - Collaborate closely with Prof. Andrea Montanari and Granica's research team to translate research into production systems. ## WHAT WE'RE LOOKING FOR - PhD in Machine Learning, Computer Science, Statistics, Applied Mathematics, or a related field. - Strong research record in machine learning. - Experience developing new models or learning algorithms. - Hands-on experience with PyTorch or JAX. - Strong programming skills in Python. - Ability to turn research ideas into working systems. - Experience in structured learning, representation learning, generative modeling, probabilistic modeling, statistical learning, or scalable ML systems is particularly relevant. ## BONUS - Research on tabular, relational, or graph data. - Experience with diffusion or other generative modeling approaches. - Publications at NeurIPS, ICML, ICLR, COLT, KDD, or related venues. - Open-source or production ML systems experience. ## COMPENSATION & BENEFITS - Competitive salary, meaningful equity, and performance bonus for top performers - 401(k) with company match, comprehensive health coverage, and unlimited PTO - Daily catered meals in our Mountain View office - Support for research, publication, and conference participation At Granica, you'll help build the next generation of enterprise AI—from exabyte-scale data infrastructure, Large Tabular Models (LTMs), and stateful AI agents. Together, we're creating the infrastructure that enables enterprises to own their data, own the intelligence built on it, and scale both efficiently. ## About Granica ## Company Overview - **One-liner**: Granica builds self-optimizing data infrastructure that compresses enterprise tabular data and enables structured intelligence for AI workloads. - **Entity Type**: Private (Series A) - **Headquarters**: Mountain View, California, United States - **Founded**: 2023 - **Founders**: Not publicly available ## Core Business - Primary industry/industries: AI Infrastructure, Data Compression, Research Services - Target customers: Enterprise B2B — SaaS, consumer-internet, healthcare, and transportation companies with petabyte-scale data estates - Mission or purpose statement: "Turning entropy to intelligence" — building a new class of data infrastructure that makes data estates efficient, reliable, and steerable for AI ## Products & Services - **[Crunch]**: A self-optimizing, lossless compression layer for structured data (Iceberg, Delta, Trino, Spark, Snowflake, BigQuery, Databricks). Reduces storage by 45–80% and cuts cloud query spend by 15–35%. Deploys inside a customer's VPC with zero code changes and zero downtime. Continuously adapts to query patterns and data drift. - **[EΣL (Extract, Signify, Load)]**: A reimagining of ETL. During "Signify," the system learns distributions, keys, and temporal drift while storing data, enabling real-time inference over a latent space without scanning cold blocks. - **[Large Tabular Models]**: In-development systems that learn cross-column and relational structure to deliver trustworthy answers and automation with provenance and governance. ## Market Standing - **Valuation/Market Cap**: Not disclosed - **Key Metric**: Total Funding — $45.0M (Series A, June 2023) - **Lead Investor**: New Enterprise Associates (led the $45M Series A, with 6 total investors) - **Notable Investors/Partners**: New Enterprise Associates, plus 5 other undisclosed institutional investors - **Growth Signals**: 44.4% headcount growth year-over-year (35 employees), LinkedIn followers up 238.9% yearly, active 9 open job postings, deployments ranging from 1 PB to 100+ PB across dozens of enterprise customers ## Competitive Advantages - **Entropy-aware compression**: Delivers state-of-the-art compression ratios (45–80% byte reduction) that are continuously adaptive to query patterns and data drift, unlike static compression schemes. - **Zero disruption deployment**: Operates inside the customer's VPC with no code changes, no downtime, and day-zero activation — dashboards show savings before "coffee cools." - **Research moat**: Foundational research published at NeurIPS 2024 (weighted empirical risk minimization with surrogate data) and ongoing work on statistical theory of data selection under weak supervision. Chief Scientist Andrea Montanari (Stanford) leads the research agenda. - **Dual value proposition**: Simultaneously reduces storage costs (pennies per GB) and accelerates query latency (petabytes queried like terabytes), while also optimizing LLM token utilization by up to 50%. ## Strategic Focus - **Near-term**: Scale Crunch adoption across enterprise data lakes, with a focus on Snowflake, Databricks, and BigQuery ecosystems. - **Medium-term**: Expand from compression into advanced subsampling and safe synthetic data generation, turning any lake into a "self-optimizing data factory." - **Long-term**: Build Large Tabular Models that enable real-time reasoning over exabyte-scale data without scanning cold blocks — replacing traditional warehouse scans with inferred answers. ## Why Work Here - **High-impact engineering culture**: 51% of the team is in technical roles, 18% in research — the company is deeply engineering-first and research-driven. Engineers work on foundational data systems for AI at petabyte scale. - **Cutting-edge ML research**: Opportunity to work alongside a Chief Scientist from Stanford and publish at top venues (NeurIPS 2024). The company is advancing the state-of-the-art in data compression, subsampling, and synthetic data. - **Remote/hybrid/office policy**: Headquarters in Mountain View, CA (287 Castro Street). Job postings indicate Mountain View is onsite. Also has offices in India (9 employees) and Austria (1 employee). - **Notable perks**: "Pays for itself" ROI philosophy — the product delivers measurable cost savings to customers. The company is well-funded ($45M Series A) with strong investor backing. - **Team composition**: Small, high-leverage team (35 people) with alumni from Meta, Salesforce, Dremio, StackRox, UiPath, and Stanford. Alums go on to LangChain, Google, Rippling, Uber, and Temporal Technologies. - **Active hiring**: 9 open positions including Senior Software Engineer (Foundational Data Systems), Engineering Manager, Research Scientist (Tabular & Structured ML), Staff Software Engineer, and Research Product Manager. ## Sources 1. [granica.ai](https://www.granica.ai/) 2. [granica.ai/about](https://www.granica.ai/about) 3. [LinkedIn](https://www.linkedin.com/company/granica-ai) 4. [PitchBook](https://pitchbook.com/profiles/company/528930-82) 5. [jobs.ashbyhq.com](https://jobs.ashbyhq.com/granica) ## Other roles at Granica - [Senior Software Engineer — Distributed Compute / Spark Systems](https://feeny.ai/job/senior-software-engineer-distributed-compute-spark-systems-granica-bay-area-7k9r7jezn5t1) — Bay Area - [Senior Software Engineer — Lakehouse Systems](https://feeny.ai/job/senior-software-engineer-lakehouse-systems-granica-bay-area-4apk25mqqrb0) — Bay Area - [Research Product Manager – AI Systems](https://feeny.ai/job/research-product-manager-ai-systems-granica-bay-area-dmfh4stjx4qb) — Bay Area - [Forward Deployed Engineer](https://feeny.ai/job/forward-deployed-engineer-granica-bay-area-3mqk9yj6gyez) — Bay Area - [Enterprise Account Executive - Mountain View, onsite](https://feeny.ai/job/enterprise-account-executive-mountain-view-onsite-granica-bay-area-v1mcbe10jx24) — Bay Area - [Enterprise Account Executive — New York Metro, remote](https://feeny.ai/job/enterprise-account-executive-new-york-metro-remote-granica-new-york-9qt45q7e6n7g) — New York, NY - [Research Scientist – Diffusion Models](https://feeny.ai/job/research-scientist-diffusion-models-granica-bay-area-g1gcawhdzsns) — Bay Area - [Head of Finance — Strategic Finance & Corporate Development](https://feeny.ai/job/head-of-finance-strategic-finance-corporate-development-granica-bay-area-77nfgrymjsbe) — Bay Area - [People Operations Manager](https://feeny.ai/job/people-operations-manager-granica-bay-area-pt66p3r3ty66) — Bay Area