--- title: 'Senior Software Engineer — Lakehouse Systems at Granica' canonical: 'https://feeny.ai/job/senior-software-engineer-lakehouse-systems-granica-bay-area-4apk25mqqrb0' type: 'job' last_seen: '2026-09-08' --- # Senior Software Engineer — Lakehouse Systems at Granica - **Company:** Granica - **Location:** Bay Area - **Employment:** full-time - **Work type:** onsite - **Posted:** 2026-08-27 - **Last confirmed live:** 2026-09-08 - **Apply:** https://jobs.ashbyhq.com/granica/dd24067b-7bb0-43f8-bb8f-ad679095078f ## Job description ## SENIOR SOFTWARE ENGINEER — LAKEHOUSE SYSTEMS Location: Mountain View, CA — On-site ## ABOUT GRANICA Granica builds AI infrastructure for enterprises operating massive data environments. Our platform helps data and engineering teams reduce storage and compute costs, improve performance and reliability, and prepare large datasets for analytics and AI. Granica’s products include: - Crunch — continuous optimization for enterprise lakehouse data - Myelin — stateful infrastructure for long-running AI agents - Large Tabular Models — foundation models designed for enterprise tables Together, we are building the infrastructure that enables enterprises to own their data, own the intelligence built on it, and scale both efficiently. Granica has demonstrated approximately $200K in annualized value per petabyte and verified customer value within weeks. ## ABOUT THE ROLE Granica is hiring a Senior Software Engineer to build foundational lakehouse systems for AI. You will work on the core infrastructure behind Crunch, Granica’s continuous optimization product for enterprise lakehouse data. This includes systems for metadata management, transaction semantics, table maintenance, object-store-backed storage layouts, file-level optimization, and lakehouse cost/performance across petabyte- and exabyte-scale environments. You will own core systems that directly affect customer infrastructure cost, query performance, table reliability, and the operational health of large lakehouse environments. This is a hands-on engineering role for someone who has deep systems experience and wants to build at the intersection of data lakes, table formats, metadata systems, storage layout, query performance, and AI infrastructure. You will work on lakehouse systems involving Apache Iceberg, Delta Lake, Apache Hudi, Parquet, ORC, cloud object stores, and query engines such as Spark, Trino, Presto, Flink, Databricks, and Snowflake-adjacent environments. ## WHAT YOU’LL DO - Build metadata and transaction systems for large-scale tabular datasets - Design systems that support time travel, schema evolution, partition evolution, snapshot isolation, and atomic consistency - Develop table-maintenance infrastructure for lakehouse formats such as Apache Iceberg, Delta Lake, and Apache Hudi - Build systems for manifests, snapshots, transaction logs, metadata pruning, snapshot expiration, table garbage collection, and catalog consistency - Optimize file layout, clustering, compaction, file sizing, data skipping, indexing, and read-path performance - Improve performance and cost efficiency across object-store-backed lakehouse environments such as S3, GCS, and ADLS - Work with columnar formats such as Parquet and ORC, including encoding, compression, layout, pruning, and read-path optimization - Build systems that make lakehouse tables faster, cheaper, and more reliable across engines and platforms such as Spark, Flink, Trino, Presto, Databricks, and Snowflake-adjacent environments - Debug performance bottlenecks across storage, metadata, table maintenance, query execution, network, and compute layers - Develop workload-aware table optimization systems that learn from access patterns and reorganize data automatically - Implement algorithms in compression, representation, layout optimization, and data efficiency - Contribute to open-source or publish research when appropriate ## WHAT WE’RE LOOKING FOR - Strong engineering depth in distributed systems, storage systems, databases, or data infrastructure - Production experience with modern data lake or lakehouse technologies such as Iceberg, Delta Lake, Hudi, Spark, Trino, Presto, Flink, Hive Metastore, Unity Catalog, or similar systems - Hands-on experience with columnar formats such as Parquet or ORC - Understanding of metadata-driven architectures, table formats, transaction semantics, query planning, and physical data layout - Experience with table maintenance, compaction, clustering, file sizing, metadata pruning, snapshot expiration, or garbage collection - Familiarity with cloud object storage systems such as S3, GCS, or ADLS and the performance tradeoffs of building lakehouse systems on top of them - Strong programming skills in Java, Scala, Go, Rust, C++, or similar systems-oriented languages - Curiosity about compression, entropy, information theory, and how data representation affects AI efficiency - A pragmatic builder’s mindset: rigorous, hands-on, and comfortable owning complex systems end to end ## BONUS - Experience contributing to Apache Iceberg, Delta Lake, Apache Hudi, Spark, Flink, Trino, Presto, Velox, DuckDB, Polars, Parquet, ORC, or related systems - Experience with manifests, snapshots, metadata catalogs, schema evolution, partition evolution, delete handling, transaction logs, or table garbage collection - Experience solving the small-file problem, optimizing object-store access patterns, or improving table health at scale - Background in storage engines, query engines, indexing, caching, encoding, compression, or adaptive query optimization - Research or open-source contributions in distributed systems, databases, storage, compression, indexing, or data processing - Interest in how physical data representation affects model training, inference, retrieval, and reasoning efficiency ## WHY JOIN GRANICA - Build foundational infrastructure for enterprise data and AI - Work on deep systems problems across lakehouse metadata, transaction semantics, table maintenance, storage layout, object-store behavior, query performance, and AI efficiency - Partner directly with Product, Engineering, and company leadership - Help shape Crunch, Granica’s production data optimization platform for enterprise-scale lakehouse environments - Work with a small, high-caliber team solving high-value infrastructure problems at massive scale - Have direct influence on architecture, product direction, customer outcomes, and company growth ## COMPENSATION & BENEFITS - Competitive salary, meaningful equity, and performance bonus for top performers - 401(k) with company match, comprehensive health coverage, and unlimited PTO - Daily catered meals in our Mountain View office - Support for research, publication, and conference participation At Granica, you'll help build the next generation of enterprise AI—from exabyte-scale data infrastructure, Large Tabular Models (LTMs), and stateful AI agents. Together, we're creating the infrastructure that enables enterprises to own their data, own the intelligence built on it, and scale both efficiently. ## About Granica ## Company Overview - **One-liner**: Granica builds self-optimizing data infrastructure that compresses enterprise tabular data and enables structured intelligence for AI workloads. - **Entity Type**: Private (Series A) - **Headquarters**: Mountain View, California, United States - **Founded**: 2023 - **Founders**: Not publicly available ## Core Business - Primary industry/industries: AI Infrastructure, Data Compression, Research Services - Target customers: Enterprise B2B — SaaS, consumer-internet, healthcare, and transportation companies with petabyte-scale data estates - Mission or purpose statement: "Turning entropy to intelligence" — building a new class of data infrastructure that makes data estates efficient, reliable, and steerable for AI ## Products & Services - **[Crunch]**: A self-optimizing, lossless compression layer for structured data (Iceberg, Delta, Trino, Spark, Snowflake, BigQuery, Databricks). Reduces storage by 45–80% and cuts cloud query spend by 15–35%. Deploys inside a customer's VPC with zero code changes and zero downtime. Continuously adapts to query patterns and data drift. - **[EΣL (Extract, Signify, Load)]**: A reimagining of ETL. During "Signify," the system learns distributions, keys, and temporal drift while storing data, enabling real-time inference over a latent space without scanning cold blocks. - **[Large Tabular Models]**: In-development systems that learn cross-column and relational structure to deliver trustworthy answers and automation with provenance and governance. ## Market Standing - **Valuation/Market Cap**: Not disclosed - **Key Metric**: Total Funding — $45.0M (Series A, June 2023) - **Lead Investor**: New Enterprise Associates (led the $45M Series A, with 6 total investors) - **Notable Investors/Partners**: New Enterprise Associates, plus 5 other undisclosed institutional investors - **Growth Signals**: 44.4% headcount growth year-over-year (35 employees), LinkedIn followers up 238.9% yearly, active 9 open job postings, deployments ranging from 1 PB to 100+ PB across dozens of enterprise customers ## Competitive Advantages - **Entropy-aware compression**: Delivers state-of-the-art compression ratios (45–80% byte reduction) that are continuously adaptive to query patterns and data drift, unlike static compression schemes. - **Zero disruption deployment**: Operates inside the customer's VPC with no code changes, no downtime, and day-zero activation — dashboards show savings before "coffee cools." - **Research moat**: Foundational research published at NeurIPS 2024 (weighted empirical risk minimization with surrogate data) and ongoing work on statistical theory of data selection under weak supervision. Chief Scientist Andrea Montanari (Stanford) leads the research agenda. - **Dual value proposition**: Simultaneously reduces storage costs (pennies per GB) and accelerates query latency (petabytes queried like terabytes), while also optimizing LLM token utilization by up to 50%. ## Strategic Focus - **Near-term**: Scale Crunch adoption across enterprise data lakes, with a focus on Snowflake, Databricks, and BigQuery ecosystems. - **Medium-term**: Expand from compression into advanced subsampling and safe synthetic data generation, turning any lake into a "self-optimizing data factory." - **Long-term**: Build Large Tabular Models that enable real-time reasoning over exabyte-scale data without scanning cold blocks — replacing traditional warehouse scans with inferred answers. ## Why Work Here - **High-impact engineering culture**: 51% of the team is in technical roles, 18% in research — the company is deeply engineering-first and research-driven. Engineers work on foundational data systems for AI at petabyte scale. - **Cutting-edge ML research**: Opportunity to work alongside a Chief Scientist from Stanford and publish at top venues (NeurIPS 2024). The company is advancing the state-of-the-art in data compression, subsampling, and synthetic data. - **Remote/hybrid/office policy**: Headquarters in Mountain View, CA (287 Castro Street). Job postings indicate Mountain View is onsite. Also has offices in India (9 employees) and Austria (1 employee). - **Notable perks**: "Pays for itself" ROI philosophy — the product delivers measurable cost savings to customers. The company is well-funded ($45M Series A) with strong investor backing. - **Team composition**: Small, high-leverage team (35 people) with alumni from Meta, Salesforce, Dremio, StackRox, UiPath, and Stanford. Alums go on to LangChain, Google, Rippling, Uber, and Temporal Technologies. - **Active hiring**: 9 open positions including Senior Software Engineer (Foundational Data Systems), Engineering Manager, Research Scientist (Tabular & Structured ML), Staff Software Engineer, and Research Product Manager. ## Sources 1. [granica.ai](https://www.granica.ai/) 2. [granica.ai/about](https://www.granica.ai/about) 3. [LinkedIn](https://www.linkedin.com/company/granica-ai) 4. [PitchBook](https://pitchbook.com/profiles/company/528930-82) 5. [jobs.ashbyhq.com](https://jobs.ashbyhq.com/granica) ## Other roles at Granica - [Senior Software Engineer — Distributed Compute / Spark Systems](https://feeny.ai/job/senior-software-engineer-distributed-compute-spark-systems-granica-bay-area-7k9r7jezn5t1) — Bay Area - [Research Product Manager – AI Systems](https://feeny.ai/job/research-product-manager-ai-systems-granica-bay-area-dmfh4stjx4qb) — Bay Area - [Forward Deployed Engineer](https://feeny.ai/job/forward-deployed-engineer-granica-bay-area-3mqk9yj6gyez) — Bay Area - [Enterprise Account Executive - Mountain View, onsite](https://feeny.ai/job/enterprise-account-executive-mountain-view-onsite-granica-bay-area-v1mcbe10jx24) — Bay Area - [Enterprise Account Executive — New York Metro, remote](https://feeny.ai/job/enterprise-account-executive-new-york-metro-remote-granica-new-york-9qt45q7e6n7g) — New York, NY - [Research Scientist – Diffusion Models](https://feeny.ai/job/research-scientist-diffusion-models-granica-bay-area-g1gcawhdzsns) — Bay Area - [Research Scientist – Large Tabular Models (LTMs)](https://feeny.ai/job/research-scientist-large-tabular-models-ltms-granica-bay-area-2qa1k66qnecd) — Bay Area - [Head of Finance — Strategic Finance & Corporate Development](https://feeny.ai/job/head-of-finance-strategic-finance-corporate-development-granica-bay-area-77nfgrymjsbe) — Bay Area - [People Operations Manager](https://feeny.ai/job/people-operations-manager-granica-bay-area-pt66p3r3ty66) — Bay Area