--- title: 'Research Scientist – Diffusion Models at Granica' canonical: 'https://feeny.ai/job/research-scientist-diffusion-models-granica-bay-area-g1gcawhdzsns' type: 'job' last_seen: '2026-09-08' --- # Research Scientist – Diffusion Models at Granica - **Company:** Granica - **Location:** Bay Area - **Compensation:** $160k–$240k - **Employment:** full-time - **Work type:** onsite - **Posted:** 2026-07-02 - **Last confirmed live:** 2026-09-08 - **Apply:** https://jobs.ashbyhq.com/granica/a0383e4c-d3b6-4cf7-9b43-e57a176a8c69 ## Job description Location: Mountain View, CA (On-site) ## OVERVIEW Diffusion models have transformed image, video, and multimodal AI. We're applying those ideas to one of the next frontiers in machine learning. At Granica, we're building Large Tabular Models (LTMs)—foundation models designed to learn natively from enterprise data. Realizing that vision requires new generative modeling techniques capable of learning from structured information at scale. Our research is led by Prof. Andrea Montanari (Stanford) and explores a fundamental question: How can diffusion models enable the next generation of AI for enterprise data? If you're excited about inventing new generative learning algorithms and applying them to entirely new domains, we'd love to talk. ## WHAT YOU'LL WORK ON - Develop novel diffusion models and generative learning algorithms. - Research new representation learning techniques for Large Tabular Models. - Design efficient training methods for large-scale generative models. - Prototype and evaluate new generative modeling approaches. - Design rigorous experiments and benchmarks to measure model quality and efficiency. - Collaborate closely with Prof. Andrea Montanari and Granica's research team to translate research into production systems. ## WHAT WE'RE LOOKING FOR - PhD in Machine Learning, Computer Science, Statistics, Applied Mathematics, or a related field. - Strong research record in generative machine learning. - Experience developing new generative models or learning algorithms. - Hands-on experience with PyTorch or JAX. - Strong programming skills in Python. - Ability to turn research ideas into working systems. - Experience with diffusion models, score-based generative modeling, representation learning, probabilistic modeling, or scalable ML systems is particularly relevant. ## BONUS - Research applying diffusion models beyond traditional vision tasks. - Publications at NeurIPS, ICML, ICLR, CVPR, ICCV, ECCV, or related venues. - Open-source or production ML systems experience. ## COMPENSATION & BENEFITS - Competitive salary, meaningful equity, and performance bonus for top performers - 401(k) with company match, comprehensive health coverage, and unlimited PTO - Daily catered meals in our Mountain View office - Support for research, publication, and conference participation At Granica, you'll help build the next generation of enterprise AI—from exabyte-scale data infrastructure, Large Tabular Models (LTMs), and stateful AI agents. Together, we're creating the infrastructure that enables enterprises to own their data, own the intelligence built on it, and scale both efficiently. ## About Granica ## Company Overview - **One-liner**: Granica builds self-optimizing data infrastructure that compresses enterprise tabular data and enables structured intelligence for AI workloads. - **Entity Type**: Private (Series A) - **Headquarters**: Mountain View, California, United States - **Founded**: 2023 - **Founders**: Not publicly available ## Core Business - Primary industry/industries: AI Infrastructure, Data Compression, Research Services - Target customers: Enterprise B2B — SaaS, consumer-internet, healthcare, and transportation companies with petabyte-scale data estates - Mission or purpose statement: "Turning entropy to intelligence" — building a new class of data infrastructure that makes data estates efficient, reliable, and steerable for AI ## Products & Services - **[Crunch]**: A self-optimizing, lossless compression layer for structured data (Iceberg, Delta, Trino, Spark, Snowflake, BigQuery, Databricks). Reduces storage by 45–80% and cuts cloud query spend by 15–35%. Deploys inside a customer's VPC with zero code changes and zero downtime. Continuously adapts to query patterns and data drift. - **[EΣL (Extract, Signify, Load)]**: A reimagining of ETL. During "Signify," the system learns distributions, keys, and temporal drift while storing data, enabling real-time inference over a latent space without scanning cold blocks. - **[Large Tabular Models]**: In-development systems that learn cross-column and relational structure to deliver trustworthy answers and automation with provenance and governance. ## Market Standing - **Valuation/Market Cap**: Not disclosed - **Key Metric**: Total Funding — $45.0M (Series A, June 2023) - **Lead Investor**: New Enterprise Associates (led the $45M Series A, with 6 total investors) - **Notable Investors/Partners**: New Enterprise Associates, plus 5 other undisclosed institutional investors - **Growth Signals**: 44.4% headcount growth year-over-year (35 employees), LinkedIn followers up 238.9% yearly, active 9 open job postings, deployments ranging from 1 PB to 100+ PB across dozens of enterprise customers ## Competitive Advantages - **Entropy-aware compression**: Delivers state-of-the-art compression ratios (45–80% byte reduction) that are continuously adaptive to query patterns and data drift, unlike static compression schemes. - **Zero disruption deployment**: Operates inside the customer's VPC with no code changes, no downtime, and day-zero activation — dashboards show savings before "coffee cools." - **Research moat**: Foundational research published at NeurIPS 2024 (weighted empirical risk minimization with surrogate data) and ongoing work on statistical theory of data selection under weak supervision. Chief Scientist Andrea Montanari (Stanford) leads the research agenda. - **Dual value proposition**: Simultaneously reduces storage costs (pennies per GB) and accelerates query latency (petabytes queried like terabytes), while also optimizing LLM token utilization by up to 50%. ## Strategic Focus - **Near-term**: Scale Crunch adoption across enterprise data lakes, with a focus on Snowflake, Databricks, and BigQuery ecosystems. - **Medium-term**: Expand from compression into advanced subsampling and safe synthetic data generation, turning any lake into a "self-optimizing data factory." - **Long-term**: Build Large Tabular Models that enable real-time reasoning over exabyte-scale data without scanning cold blocks — replacing traditional warehouse scans with inferred answers. ## Why Work Here - **High-impact engineering culture**: 51% of the team is in technical roles, 18% in research — the company is deeply engineering-first and research-driven. Engineers work on foundational data systems for AI at petabyte scale. - **Cutting-edge ML research**: Opportunity to work alongside a Chief Scientist from Stanford and publish at top venues (NeurIPS 2024). The company is advancing the state-of-the-art in data compression, subsampling, and synthetic data. - **Remote/hybrid/office policy**: Headquarters in Mountain View, CA (287 Castro Street). Job postings indicate Mountain View is onsite. Also has offices in India (9 employees) and Austria (1 employee). - **Notable perks**: "Pays for itself" ROI philosophy — the product delivers measurable cost savings to customers. The company is well-funded ($45M Series A) with strong investor backing. - **Team composition**: Small, high-leverage team (35 people) with alumni from Meta, Salesforce, Dremio, StackRox, UiPath, and Stanford. Alums go on to LangChain, Google, Rippling, Uber, and Temporal Technologies. - **Active hiring**: 9 open positions including Senior Software Engineer (Foundational Data Systems), Engineering Manager, Research Scientist (Tabular & Structured ML), Staff Software Engineer, and Research Product Manager. ## Sources 1. [granica.ai](https://www.granica.ai/) 2. [granica.ai/about](https://www.granica.ai/about) 3. [LinkedIn](https://www.linkedin.com/company/granica-ai) 4. [PitchBook](https://pitchbook.com/profiles/company/528930-82) 5. [jobs.ashbyhq.com](https://jobs.ashbyhq.com/granica) ## Other roles at Granica - [Senior Software Engineer — Distributed Compute / Spark Systems](https://feeny.ai/job/senior-software-engineer-distributed-compute-spark-systems-granica-bay-area-7k9r7jezn5t1) — Bay Area - [Senior Software Engineer — Lakehouse Systems](https://feeny.ai/job/senior-software-engineer-lakehouse-systems-granica-bay-area-4apk25mqqrb0) — Bay Area - [Research Product Manager – AI Systems](https://feeny.ai/job/research-product-manager-ai-systems-granica-bay-area-dmfh4stjx4qb) — Bay Area - [Forward Deployed Engineer](https://feeny.ai/job/forward-deployed-engineer-granica-bay-area-3mqk9yj6gyez) — Bay Area - [Enterprise Account Executive - Mountain View, onsite](https://feeny.ai/job/enterprise-account-executive-mountain-view-onsite-granica-bay-area-v1mcbe10jx24) — Bay Area - [Enterprise Account Executive — New York Metro, remote](https://feeny.ai/job/enterprise-account-executive-new-york-metro-remote-granica-new-york-9qt45q7e6n7g) — New York, NY - [Research Scientist – Large Tabular Models (LTMs)](https://feeny.ai/job/research-scientist-large-tabular-models-ltms-granica-bay-area-2qa1k66qnecd) — Bay Area - [Head of Finance — Strategic Finance & Corporate Development](https://feeny.ai/job/head-of-finance-strategic-finance-corporate-development-granica-bay-area-77nfgrymjsbe) — Bay Area - [People Operations Manager](https://feeny.ai/job/people-operations-manager-granica-bay-area-pt66p3r3ty66) — Bay Area