--- title: 'Senior Software Engineer — Distributed Compute / Spark Systems at Granica' canonical: 'https://feeny.ai/job/senior-software-engineer-distributed-compute-spark-systems-granica-bay-area-7k9r7jezn5t1' type: 'job' last_seen: '2026-09-08' --- # Senior Software Engineer — Distributed Compute / Spark Systems at Granica - **Company:** Granica - **Location:** Bay Area - **Employment:** full-time - **Work type:** hybrid - **Posted:** 2026-08-27 - **Last confirmed live:** 2026-09-08 - **Apply:** https://jobs.ashbyhq.com/granica/68e0b1db-0f03-4ad3-85e9-586aeb46f6c2 ## Job description SENIOR SOFTWARE ENGINEER — DISTRIBUTED COMPUTE / SPARK SYSTEMS Location: Mountain View, CA — On-site ## ABOUT GRANICA Granica builds AI infrastructure for enterprises operating massive data environments. Our platform helps data and engineering teams reduce storage and compute costs, improve performance and reliability, and prepare large datasets for analytics and AI. Granica’s products include: - Crunch — continuous optimization for enterprise lakehouse data - Myelin — stateful infrastructure for long-running AI agents - Large Tabular Models — foundation models designed for enterprise tables Together, we are building the infrastructure that enables enterprises to own their data, own the intelligence built on it, and scale both efficiently. Granica has demonstrated approximately $200K in annualized value per petabyte and verified customer value within weeks. ## ABOUT THE ROLE Granica is hiring a Senior Software Engineer to build distributed compute systems for enterprise-scale data and AI workloads. You will work on the core infrastructure behind Crunch, Granica’s continuous optimization product for enterprise lakehouse data. This includes systems for distributed execution, workload optimization, query performance, scheduling, resource management, and compute cost reduction across petabyte- and exabyte-scale environments. You will own core systems that directly affect customer compute spend, query latency, workload reliability, cluster efficiency, and the performance of large-scale analytical data processing. This is a hands-on engineering role for someone who has deep systems experience and wants to build at the intersection of Spark, distributed query execution, lakehouse compute, workload scheduling, storage-aware optimization, and AI infrastructure. You will work on distributed compute systems involving Apache Spark, Spark SQL, Trino, Presto, Flink, Databricks, Snowflake-adjacent environments, cloud object stores, and lakehouse formats such as Apache Iceberg, Delta Lake, and Apache Hudi. ## WHAT YOU’LL DO - Build distributed compute systems for large-scale analytical and AI workloads - Improve performance and cost efficiency across Spark, Trino, Presto, Flink, Databricks, and Snowflake-adjacent environments - Design workload-aware systems for query execution, resource allocation, scheduling, and compute optimization - Optimize execution performance across joins, aggregations, scans, shuffles, spills, caching, partitioning, and task scheduling - Build systems that learn from workload patterns and automatically improve execution plans, cluster usage, and compute efficiency - Develop infrastructure for adaptive workload routing, execution planning, and data-processing reliability across large customer environments - Debug performance bottlenecks across query execution, metadata, storage, network, memory, CPU, and distributed compute layers - Work with lakehouse tables and columnar formats such as Iceberg, Delta Lake, Hudi, Parquet, and ORC to improve end-to-end workload performance - Build systems that reduce compute waste caused by inefficient scans, poor partitioning, small files, skew, unnecessary shuffles, and suboptimal workload placement - Improve reliability and failure recovery for large distributed data-processing jobs - Implement algorithms in workload optimization, execution efficiency, cost modeling, and data-processing performance - Contribute to open-source or publish research when appropriate ## WHAT WE’RE LOOKING FOR - Strong engineering depth in distributed systems, data processing systems, query engines, databases, or cloud infrastructure - Production experience with distributed compute or query systems such as Apache Spark, Spark SQL, Trino, Presto, Flink, Databricks, EMR, Glue, Hive, or similar systems - Hands-on experience improving performance, reliability, or cost efficiency for large-scale data-processing workloads - Understanding of distributed execution, query planning, scheduling, resource management, fault tolerance, and workload isolation - Experience with Spark internals, Spark SQL, Catalyst, Adaptive Query Execution, shuffle, joins, aggregation, spill, memory management, or task scheduling - Familiarity with lakehouse formats and columnar data such as Iceberg, Delta Lake, Hudi, Parquet, or ORC - Familiarity with cloud object storage systems such as S3, GCS, or ADLS and the performance tradeoffs of running distributed compute on top of them - Strong programming skills in Scala, Java, Go, Rust, C++, or similar systems-oriented languages - Curiosity about workload optimization, cost modeling, adaptive execution, and how compute efficiency affects AI and analytics at scale - A pragmatic builder’s mindset: rigorous, hands-on, and comfortable owning complex systems end to end ## BONUS - Experience contributing to Apache Spark, Spark SQL, Trino, Presto, Flink, Velox, DuckDB, DataFusion, Iceberg, Delta Lake, Hudi, Parquet, ORC, or related systems - Experience with Catalyst, Adaptive Query Execution, cost-based optimization, query planning, vectorized execution, or distributed runtime systems - Experience optimizing joins, aggregations, shuffles, scans, spills, caching, partitioning, skew handling, or task scheduling - Experience building workload schedulers, execution control planes, resource managers, or multi-engine compute platforms - Experience reducing compute cost or improving workload efficiency in large-scale production data environments - Background in query engines, distributed runtimes, storage-aware execution, indexing, caching, encoding, compression, or adaptive query optimization - Research or open-source contributions in distributed systems, databases, query processing, data processing, or cloud infrastructure ## WHY JOIN GRANICA - Build foundational infrastructure for enterprise data and AI - Work on deep systems problems across distributed compute, query execution, workload optimization, scheduling, resource management, and compute efficiency - Partner directly with Product, Engineering, and company leadership - Help shape Crunch, Granica’s production data optimization platform for enterprise-scale lakehouse environments - Work with a small, high-caliber team solving high-value infrastructure problems at massive scale - Have direct influence on architecture, product direction, customer outcomes, and company growth ## COMPENSATION & BENEFITS - Competitive salary, meaningful equity, and performance bonus for top performers - 401(k) with company match, comprehensive health coverage, and unlimited PTO - Daily catered meals in our Mountain View office - Support for research, publication, and conference participation At Granica, you'll help build the next generation of enterprise AI—from exabyte-scale data infrastructure, Large Tabular Models (LTMs), and stateful AI agents. Together, we're creating the infrastructure that enables enterprises to own their data, own the intelligence built on it, and scale both efficiently. ## About Granica ## Company Overview - **One-liner**: Granica builds self-optimizing data infrastructure that compresses enterprise tabular data and enables structured intelligence for AI workloads. - **Entity Type**: Private (Series A) - **Headquarters**: Mountain View, California, United States - **Founded**: 2023 - **Founders**: Not publicly available ## Core Business - Primary industry/industries: AI Infrastructure, Data Compression, Research Services - Target customers: Enterprise B2B — SaaS, consumer-internet, healthcare, and transportation companies with petabyte-scale data estates - Mission or purpose statement: "Turning entropy to intelligence" — building a new class of data infrastructure that makes data estates efficient, reliable, and steerable for AI ## Products & Services - **[Crunch]**: A self-optimizing, lossless compression layer for structured data (Iceberg, Delta, Trino, Spark, Snowflake, BigQuery, Databricks). Reduces storage by 45–80% and cuts cloud query spend by 15–35%. Deploys inside a customer's VPC with zero code changes and zero downtime. Continuously adapts to query patterns and data drift. - **[EΣL (Extract, Signify, Load)]**: A reimagining of ETL. During "Signify," the system learns distributions, keys, and temporal drift while storing data, enabling real-time inference over a latent space without scanning cold blocks. - **[Large Tabular Models]**: In-development systems that learn cross-column and relational structure to deliver trustworthy answers and automation with provenance and governance. ## Market Standing - **Valuation/Market Cap**: Not disclosed - **Key Metric**: Total Funding — $45.0M (Series A, June 2023) - **Lead Investor**: New Enterprise Associates (led the $45M Series A, with 6 total investors) - **Notable Investors/Partners**: New Enterprise Associates, plus 5 other undisclosed institutional investors - **Growth Signals**: 44.4% headcount growth year-over-year (35 employees), LinkedIn followers up 238.9% yearly, active 9 open job postings, deployments ranging from 1 PB to 100+ PB across dozens of enterprise customers ## Competitive Advantages - **Entropy-aware compression**: Delivers state-of-the-art compression ratios (45–80% byte reduction) that are continuously adaptive to query patterns and data drift, unlike static compression schemes. - **Zero disruption deployment**: Operates inside the customer's VPC with no code changes, no downtime, and day-zero activation — dashboards show savings before "coffee cools." - **Research moat**: Foundational research published at NeurIPS 2024 (weighted empirical risk minimization with surrogate data) and ongoing work on statistical theory of data selection under weak supervision. Chief Scientist Andrea Montanari (Stanford) leads the research agenda. - **Dual value proposition**: Simultaneously reduces storage costs (pennies per GB) and accelerates query latency (petabytes queried like terabytes), while also optimizing LLM token utilization by up to 50%. ## Strategic Focus - **Near-term**: Scale Crunch adoption across enterprise data lakes, with a focus on Snowflake, Databricks, and BigQuery ecosystems. - **Medium-term**: Expand from compression into advanced subsampling and safe synthetic data generation, turning any lake into a "self-optimizing data factory." - **Long-term**: Build Large Tabular Models that enable real-time reasoning over exabyte-scale data without scanning cold blocks — replacing traditional warehouse scans with inferred answers. ## Why Work Here - **High-impact engineering culture**: 51% of the team is in technical roles, 18% in research — the company is deeply engineering-first and research-driven. Engineers work on foundational data systems for AI at petabyte scale. - **Cutting-edge ML research**: Opportunity to work alongside a Chief Scientist from Stanford and publish at top venues (NeurIPS 2024). The company is advancing the state-of-the-art in data compression, subsampling, and synthetic data. - **Remote/hybrid/office policy**: Headquarters in Mountain View, CA (287 Castro Street). Job postings indicate Mountain View is onsite. Also has offices in India (9 employees) and Austria (1 employee). - **Notable perks**: "Pays for itself" ROI philosophy — the product delivers measurable cost savings to customers. The company is well-funded ($45M Series A) with strong investor backing. - **Team composition**: Small, high-leverage team (35 people) with alumni from Meta, Salesforce, Dremio, StackRox, UiPath, and Stanford. Alums go on to LangChain, Google, Rippling, Uber, and Temporal Technologies. - **Active hiring**: 9 open positions including Senior Software Engineer (Foundational Data Systems), Engineering Manager, Research Scientist (Tabular & Structured ML), Staff Software Engineer, and Research Product Manager. ## Sources 1. [granica.ai](https://www.granica.ai/) 2. [granica.ai/about](https://www.granica.ai/about) 3. [LinkedIn](https://www.linkedin.com/company/granica-ai) 4. [PitchBook](https://pitchbook.com/profiles/company/528930-82) 5. [jobs.ashbyhq.com](https://jobs.ashbyhq.com/granica) ## Other roles at Granica - [Senior Software Engineer — Lakehouse Systems](https://feeny.ai/job/senior-software-engineer-lakehouse-systems-granica-bay-area-4apk25mqqrb0) — Bay Area - [Research Product Manager – AI Systems](https://feeny.ai/job/research-product-manager-ai-systems-granica-bay-area-dmfh4stjx4qb) — Bay Area - [Forward Deployed Engineer](https://feeny.ai/job/forward-deployed-engineer-granica-bay-area-3mqk9yj6gyez) — Bay Area - [Enterprise Account Executive - Mountain View, onsite](https://feeny.ai/job/enterprise-account-executive-mountain-view-onsite-granica-bay-area-v1mcbe10jx24) — Bay Area - [Enterprise Account Executive — New York Metro, remote](https://feeny.ai/job/enterprise-account-executive-new-york-metro-remote-granica-new-york-9qt45q7e6n7g) — New York, NY - [Research Scientist – Diffusion Models](https://feeny.ai/job/research-scientist-diffusion-models-granica-bay-area-g1gcawhdzsns) — Bay Area - [Research Scientist – Large Tabular Models (LTMs)](https://feeny.ai/job/research-scientist-large-tabular-models-ltms-granica-bay-area-2qa1k66qnecd) — Bay Area - [Head of Finance — Strategic Finance & Corporate Development](https://feeny.ai/job/head-of-finance-strategic-finance-corporate-development-granica-bay-area-77nfgrymjsbe) — Bay Area - [People Operations Manager](https://feeny.ai/job/people-operations-manager-granica-bay-area-pt66p3r3ty66) — Bay Area