--- title: 'Compiler Optimization Engineer at Lemurian Labs' canonical: 'https://feeny.ai/job/compiler-optimization-engineer-lemurian-labs-santa-clara-1gzxed062b0j' type: 'job' last_seen: '2026-09-08' --- # Compiler Optimization Engineer at Lemurian Labs - **Company:** Lemurian Labs - **Location:** Santa Clara, CA / Toronto, Canada - **Work type:** remote - **Posted:** 2026-02-13 - **Last confirmed live:** 2026-09-08 - **Apply:** https://job-boards.greenhouse.io/lemurianlabs/jobs/4118628009 ## Job description ## About Us At Lemurian Labs, we're reimagining the foundations of computing to make AI accessible to everyone. Our mission is to remove the limits of scale, hardware, and cost that hold back innovation, so the people solving humanity's hardest problems can move faster. We're building a new kind of software stack: a hardware-agnostic platform that makes every system — from a laptop to a supercomputer — feel like one seamless engine. Developers can write once, run anywhere, and get state-of-the-art performance across any chip, any cloud, at any scale. It's a complete rethink of how software and hardware interact — designed for the era beyond Moore's Law. We're not looking for the comfortable or the conventional; we're looking for the bold. The engineers who crave frontier problems, who want to bend the limits of what's possible, who see infrastructure not as a constraint but as a canvas. If you want to build the foundation for the next era of AI and change what humanity can achieve in the process, join us. ## About the Role We're looking for a Graph Optimization Compiler Engineer to own the middle tier of our AI compiler stack — the layer where high-level model graphs are transformed, simplified, and made ready for efficient code generation. You'll design and implement the optimization passes that make the difference between a model that runs and a model that flies. This role sits between our compiler front end and code generation backend. You'll work on graph-level transformations — fusion, layout optimization, dead code elimination, constant folding, and more — with a direct line of sight to the performance outcomes your work produces. If you think in data flow graphs and optimization passes, and you want that thinking to power the next generation of AI infrastructure, we'd love to talk. ## What You'll Do - Design, develop, and maintain the graph optimization layer of our heterogeneous AI compiler - Implement and extend graph-level transformation passes including operator fusion, layout propagation, dead code elimination, constant folding, and algebraic simplification - Define and evolve our intermediate representation (IR) to support new optimization opportunities as ML model architectures advance - Analyze performance data to identify optimization gaps and drive measurable improvements in throughput and latency - Collaborate with front end and code generation teams to ensure clean IR interfaces and well-structured optimization pipelines - Propose and prototype new optimization strategies in response to advances in model design and hardware capabilities - Contribute to testing and validation infrastructure to ensure optimization correctness across model types and hardware targets ## Requirements Essential Skills and Experience - BS degree in Computer Science, Computer Engineering, or equivalent practical experience - 4+ years of experience working with compilers, with a focus on intermediate representation design or optimization passes - Deep knowledge of graph-level compiler optimization techniques — fusion, tiling, layout transformations, and related methods - 4+ years of experience with C/C++ - Strong written and verbal communication skills; ability to write clear and concise technical documentation ## Preferred Skills and Experience - Master's or PhD in Computer Science, Computer Engineering, or equivalent - Experience with polyhedral models or affine analysis for loop and tensor optimization - Familiarity with hardware memory hierarchies and how layout decisions impact performance on GPUs or accelerators - Experience working with MLIR, XLA, or similar graph-level IR frameworks - Experience with ML framework internals — PyTorch eager/compile mode, JAX/XLA, or TensorRT - Strong understanding of ML model architectures and their computational patterns (attention, convolution, normalization, etc.) - Knowledge of quantization, sparsity, or other model-level optimization techniques - Contributions to open-source compiler or ML infrastructure projects ## Why Join Lemurian Labs - Own a critical layer of our compiler stack where optimization decisions have direct, measurable impact on model performance - Work on the hardest graph-level problems in AI infrastructure — across diverse hardware targets and model architectures - Collaborate with a team that treats infrastructure as a canvas and optimization as a craft - Competitive compensation including equity, medical/dental/vision, retirement savings, and wellness benefits Lemurian Labs is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees, regardless of gender identity, race, ethnicity, sexual orientation, disability status, age, or background. Compensation depends on experience and geographic location and will be narrowed during the interview process. Additional benefits include equity, company bonus opportunities, medical, dental, and vision coverage, a retirement savings plan, and supplemental wellness benefits. ## About Lemurian Labs ## Company Overview - **One-liner**: Lemurian Labs builds a hardware-agnostic AI software stack (Tachyon) that enables organizations to write AI workloads once and deploy them across any hardware—CPUs, GPUs, accelerators—with performance matching hand-tuned kernels, eliminating vendor lock-in and kernel rewrites. - **Entity Type**: Private (Startup) - **Headquarters**: Santa Clara, California, United States (with offices in Canada and India) - **Founded**: 2021 - **Founders**: Jay Dawani (CEO & Co-Founder), Dr. Vassil Dimitrov (Chief Scientist & Co-Founder) ## Core Business - **Primary industry**: AI Infrastructure / Software - **Target customers**: B2B, Enterprise (AI teams at cloud providers, hardware vendors, and large-scale AI model developers) - **Mission or purpose statement**: "We're building the hardware-agnostic software infrastructure that makes AI accessible to everyone—fast, affordable, and scalable. Not just better tools. The foundation for what comes next." ## Products & Services - **Tachyon**: A complete software stack comprising five components: - **Tachyon DSL**: A Pythonic tensor language for expressing end-to-end AI pipelines. - **Hardware-Independent Concurrent IR**: Represents computation using parallel programming primitives that explicitly capture concurrency and data flow. - **Hardware-Parameterized Graph Optimizer**: Restructures computation to maximize data reuse, minimize data movement, and eliminate synchronization barriers based on detailed hardware characteristics. - **Portable Optimizing Code Generator**: Generates high-performance object code for any target hardware (NVIDIA, AMD, Intel, etc.) using hardware characteristics as parameters. - **Hierarchical Dynamic Runtime**: Explicitly manages data movement and orchestrates execution across clusters, adapting dynamically to changing runtime conditions. - **Type**: Software (SaaS/Platform) ## Market Standing - **Valuation/Market Cap**: Not publicly disclosed - **Key Metric**: Total Funding of **$15.1M** across 6 rounds; Annual Revenue estimated at **$2.3M** (per LinkedIn data) - **Notable Investors/Partners**: Oval Park Capital (led Seed Round), Silicon Catalyst (Non-Equity Assistance), ventureLAB (Non-Equity Assistance); team alumni from Sun Microsystems, Intel, NVIDIA, IBM, Huawei, Qualcomm, AMD, Uber, Google - **Growth Signals**: - Headcount: 34 employees (+11.1% YoY) - LinkedIn followers: 5,181 (+113.2% yearly growth) - Active job postings: 8 (+166.7% yearly growth) - Beta testing for Tachyon opens Summer 2027 - Claims 60-80% reduction in serving costs from higher utilization ## Competitive Advantages - **Hardware-Agnostic Performance**: Tachyon matches or exceeds hand-tuned kernel performance on any hardware without vendor lock-in—a direct solution to the fragmentation and cost problems plaguing AI infrastructure. - **Accelerated Software Paradigm**: The company positions itself as leading a shift from "accelerated hardware" to "accelerated software," targeting the post-Moore's Law era. - **Deep Technical Expertise**: Founders and team have decades of experience at top chip and systems companies (NVIDIA, Intel, AMD, Google, Uber, etc.), giving them unique insight into the weaknesses of current approaches. - **Portability Without Compromise**: Claims to enable moving between clouds and hardware in hours, not months, and to support new accelerators in <90 days via a plug-in model. ## Strategic Focus - **Current Priorities**: Completing Tachyon development and launching beta testing in Summer 2027; scaling the engineering team (especially compiler engineers, ML performance engineers, and runtime engineers); building developer community and product-market fit. - **Direction for Growth**: Becoming the standard software layer for AI deployment across all hardware, enabling companies to focus on AI innovation rather than infrastructure battles. ## Why Work Here - **Culture Highlights**: Values include "Relentless execution," "Fail fast, learn faster," "Extreme ownership," "Radical candor," and "Be a force multiplier." The team emphasizes moving quickly, testing aggressively, and iterating relentlessly. - **Remote/Hybrid/Office Policy**: Has 4 offices across 2 countries (Santa Clara HQ, Menlo Park, Santa Clara, and Oakville, Canada); LinkedIn data suggests presence in US (55%), Canada (28%), and India (2%). Likely hybrid or flexible given distributed team. - **Notable Perks/Engineering Culture**: Employer rating of **4.7/5.0** (3 reviews) with Work-Life: 5.0, Compensation: 4.0, Culture: 4.7, Career: 4.7. The company is building a groundbreaking software stack from the ground up—ideal for engineers passionate about compilers, systems programming, and AI infrastructure. Team composition is 50% technical roles. Notable talent sources include Intel, Huawei Canada, Luminous Computing, and Untether AI. - **Mission-Driven**: Explicitly focused on enabling AI to solve humanity's hardest problems (climate change, healthcare, education, scientific breakthroughs) by removing infrastructure barriers. ## Sources 1. [Lemurian Labs Website](https://www.lemurianlabs.com/) 2. [Lemurian Labs About Page](https://www.lemurianlabs.com/about) 3. [Lemurian Labs Technology Page](https://www.lemurianlabs.com/technology) 4. [Lemurian Labs Careers Page](https://job-boards.greenhouse.io/lemurianlabs) 5. [Lemurian Labs LinkedIn](https://www.linkedin.com/company/lemurianlabs) ## Other roles at Lemurian Labs - [Senior DSL Engineer](https://feeny.ai/job/senior-dsl-engineer-lemurian-labs-santa-clara-california-united-states-toronto-aa87rvqkddce) — Santa Clara California United States Toronto Ontario Canada, United States - [Senior Developer Tools Engineer](https://feeny.ai/job/senior-developer-tools-engineer-lemurian-labs-santa-clara-ww9gbhb1h1c9) — Santa Clara, CA / Toronto, Canada - [Runtime Engineer](https://feeny.ai/job/runtime-engineer-lemurian-labs-santa-clara-rhjmgrt5gf44) — Santa Clara, CA / Toronto, Canada - [Front End Compiler](https://feeny.ai/job/front-end-compiler-lemurian-labs-santa-clara-6b4bkrr13m3f) — Santa Clara, CA / Toronto, Canada - [Compiler Code Gen Engineer](https://feeny.ai/job/compiler-code-gen-engineer-lemurian-labs-santa-clara-st5wp0zdt0mk) — Santa Clara, CA / Toronto, Canada