--- title: 'Developer Advocate, MAX Inference & Serving at Modular' canonical: 'https://feeny.ai/job/developer-advocate-max-inference-serving-modular-united-states-canada-s1r9bnna72jz' type: 'job' last_seen: '2026-09-10' --- # Developer Advocate, MAX Inference & Serving at Modular - **Company:** Modular - **Location:** United States / Canada - **Employment:** full-time - **Work type:** remote - **Posted:** 2026-08-07 - **Last confirmed live:** 2026-09-10 - **Apply:** https://jobs.gem.com/modular/am9icG9zdDr8mLYVA244cbADr5tGwP5M ## Job description About the role: We are looking for a Developer Advocate to evangelize the MAX Platform's inference and serving capabilities with our user base and developer community. This involves creating technical content such as user guides and blog posts as well as giving talks at conferences, leading workshops, all with the goal of enabling our community of builders deploying models in production. Join our world-leading product team and be part of redefining how AI infrastructure is built and deployed. LOCATION:  Candidates based in the US or Canada are welcome to apply. To support growth and collaboration, those in earlier career stages work in a hybrid capacity at our Los Altos, CA. More senior staff can work out of our office in Los Altos, CA or remotely from home. Onboarding for new hires is conducted in-person in our Los Altos, CA office. Additionally, this role requires travel to conferences and developer events, which may be as often as once per month, as well as travel for team and company events (typically 2-4 times per year). What you will do: - Build and publish reproducible benchmarks comparing MAX against vLLM, SGLang, TensorRT-LLM, and Triton Inference Server, including methodology, harness code, and hardware configurations so others can verify the numbers. - Deploy and profile real inference workloads on MAX across CPU and GPU targets, and investigate performance gaps in latency, throughput, and cost per token. - Write and maintain technical content grounded in that work: performance deep dives, serving architecture explainers, and posts that show how MAX handles batching, KV cache management, quantization, and multi-GPU serving. - Author runnable tutorials, examples, and video walkthroughs covering model deployment on MAX, from pip install modular to a served endpoint under load. - Provide technical support to engineers evaluating MAX for inference and serving, including teams at some of the world's largest companies, and debug their deployment and performance issues directly. - Answer technical questions across GitHub, Discord, X, and LinkedIn, reproducing reported issues and filing them with enough detail for engineering to act on. - Translate community and customer findings into prioritized inference and serving feedback for engineering and product, shaping the MAX roadmap. - Present technical talks, benchmark results, and live demos at conferences, summits, and meetups. - Help set the technical direction of Modular's developer relations content, including what gets benchmarked, documented, and demoed next. What you bring to the table: - 3-5 years of professional engineering experience, with at least some of it spent deploying or operating ML systems in production. You have run inference workloads yourself, not just written about them. - Deep familiarity with the ML inference stack. You can explain continuous batching, KV cache management, quantization trade-offs, and tensor and pipeline parallelism, and you have hands-on experience with at least one of vLLM, Triton Inference Server, TensorRT-LLM, or SGLang. - Strong Python and systems programming experience in C++, Rust, Mojo, or CUDA is a significant advantage, especially if you have profiled and optimized GPU code. - Comfort with the deployment surface around serving: containers, Kubernetes, GPU drivers and runtimes, and the usual ways a cluster refuses to cooperate. - You benchmark rigorously. You know why one number is not a result, how to control for warmup and batch size, and when a comparison is not apples to apples. - You write well about technical work, and you have a portfolio to show it: blog posts, tutorials, videos, docs, or courses. You explain complex ideas without losing precision, and you cut the filler. - You learn new tools fast and turn them into accurate content quickly. A feature ships Tuesday, your tutorial goes out Thursday and the code runs. - A growth and leadership mindset, with a collaborative attitude that seeks to learn more from our customers, team members, and the broader market Minimum Qualifications: - Bachelor's degree in Engineering, Information Systems, Computer Science, or technical related field. - 2+ years of Product Management or related work experience. *Completed advanced degrees in a relevant field may be substituted for up to two years of work experience. Helpful, but not required: - Experience programming GPUs using CUDA or ROCm. - Familiarity with the Mojo 🔥 programming language and MAX AI framework. - Familiarity with open source software development practices and communities. - Experience producing high-quality videos covering technical topics. What Modular brings to the table: - Amazing Team. We are a progressive and agile team with some of the industry’s best engineering and product leaders. - World-class Benefits. In order to attract the best, we need to offer the best. Your benefits package may include comprehensive healthcare coverage, retirement and savings programs, employee stock purchase opportunities, paid time off, wellbeing resources, family support programs, and learning and development opportunities. Please note that specific benefit packages may vary based on your location, you can read more about [benefits offered by Qualcomm here](https://www.qualcomm.com/company/careers/benefits). - Competitive Compensation. We offer very strong compensation packages, including RSU grants. We want people to be focused on their best work and believe in tailoring compensation plans to meet the needs of our workforce. - Team Building Events. We organize regular team onsites and local meetups in Los Altos, CA as well as different cities. Traveling 2-4 times a year is expected for all roles. Working at Modular will enable you to grow quickly as you work alongside incredibly motivated and talented people who have high standards, possess a growth mindset, and a purpose to truly change the world. The estimated base salary range for this role to be performed in the US is $150,200.00 - $225,400.00 USD. The estimated base salary range for this role to be performed in Canada is $111,500.00 - $167,300.00 CAD. The salary for the successful applicant will depend on a variety of permissible, non-discriminatory job-related factors, which include but are not limited to education, training, work experience, business needs, or market demands. This range may be modified in the future. The total compensation for a candidate will also include annual target bonus, equity, and benefits, with equity making up a significant portion of your total compensation. For candidates who fall outside of the listed requirements, we nevertheless encourage you to apply as we may have upcoming openings that are lower/higher level than the ones advertised. ## About Modular ## Company Overview - **One-liner**: Modular builds a unified, high-performance AI inference platform that enables developers to run AI workloads efficiently across any hardware. - **Entity Type**: Private (Acquired by Qualcomm in July 2026) - **Headquarters**: Los Altos, California, United States - **Founded**: 2022 - **Founders**: Chris Lattner and Tim Davis ## Core Business - **Primary industry/industries**: AI Infrastructure, Deep Learning, Software Development - **Target customers**: AI/ML engineers, data scientists, and enterprises deploying large-scale AI models (B2B, Enterprise) - **Mission statement**: "Make AI’s compute layer unified, efficient, and accessible to all." ## Products & Services - **Modular Platform (MAX)**: A unified AI inference platform offering text, audio, and image inference. It provides state-of-the-art performance with shared or dedicated endpoints, deployment in Modular's cloud or the customer's VPC, and support for custom models. It includes a high-performance, hardware-agnostic serving framework that automatically optimizes kernels across accelerators. - **Mojo**: A high-performance systems language designed for writing composable GPU kernels, enabling maximum performance across different hardware. ## Market Standing - **Valuation/Market Cap**: Not publicly available (acquired by Qualcomm). Total funding raised was $380M. - **Key Metric**: Total Funding of $380M, with an annual revenue of $1M as of the latest data. - **Notable Investors/Partners**: General Catalyst (lead, Series B), Google Ventures (lead, Seed), US Innovative Technology Fund (lead, Series C). - **Growth Signals**: Headcount grew by 39.9% year-over-year to 157 employees. The company was acquired by Qualcomm in July 2026, signaling a major strategic validation and exit. It maintains a strong technical workforce (110 out of 157 employees are in technical roles). ## Competitive Advantages - **Hardware Agnosticism**: The platform runs seamlessly across NVIDIA, AMD, Trainium, TPU, Qualcomm, Intel, ARM, and Apple silicon, providing true portability and preventing vendor lock-in. - **Full-Stack Optimization**: Optimizes from low-level GPU kernels to API endpoints, delivering significant performance gains (e.g., 2x improvement over vLLM on diverse hardware) and up to 50% cost savings. - **Founding Team**: Led by Chris Lattner (creator of LLVM, Clang, and Swift) and Tim Davis, with decades of experience building AI infrastructure at Google and other big tech companies. ## Strategic Focus - **Unifying AI Compute**: The company’s core mission is to solve the fragmentation in AI infrastructure by creating a modular and composable platform. - **Open Source & Ecosystem**: Modular is open-sourcing the Mojo language and the MAX engine to drive adoption and build a community. - **Post-Acquisition Integration**: Following the acquisition by Qualcomm, the strategic focus will likely shift towards integrating its technology into Qualcomm’s hardware ecosystem and scaling its deployment to edge and mobile devices. ## Why Work Here - **Cutting-Edge Technology**: Employees work on building deep learning infrastructure from silicon to system, offering full-stack mastery of the AI software/hardware stack. - **Strong Compensation & Benefits**: Offers leading medical, dental, and vision insurance, strong compensation and equity packages, a 401k plan with up to 5% match, generous parental leave, unlimited paid time off, and a $1,500 work-from-home stipend. - **Culture**: Described as intellectually curious, humble, and collaborative. The company values work/life balance and offers flexible work hours and hybrid work options. - **High Employee Ratings**: On LinkedIn, the company has a 4.3/5.0 employer rating, with perfect 5.0 scores for Culture and Career, and 4.6 for Work-Life and Compensation. - **Onboarding & Process**: Onboarding occurs onsite at the Los Altos office. The interview process is straightforward, generally taking about 4 weeks, and includes a culture interview. ## Sources 1. [Modular Careers Page](https://www.modular.com/company/careers) 2. [Modular About Us](https://www.modular.com/company/about) 3. [Modular Website](https://www.modular.com/) 4. [Modular LinkedIn](https://www.linkedin.com/company/modular-ai) 5. [Modular Careers on Gem](https://jobs.gem.com/modular) ## Other roles at Modular - [Senior Open Source Community Engineer](https://feeny.ai/job/senior-open-source-community-engineer-modular-united-states-canada-4mtpdvn8xgw9) — United States / Canada - [Developer Advocate, Mojo](https://feeny.ai/job/developer-advocate-mojo-modular-united-states-canada-hgq4cjnkdxh8) — United States / Canada - [AI Inference Tools Engineer](https://feeny.ai/job/ai-inference-tools-engineer-modular-united-states-8sk0kdjq8g93) — United States - [Inference Optimization Manager](https://feeny.ai/job/inference-optimization-manager-modular-united-states-canada-9910s7qpd0fy) — United States / Canada - [Inference Optimization Engineer](https://feeny.ai/job/inference-optimization-engineer-modular-united-states-canada-h70h6szsvg8e) — United States / Canada - [Product Manager, Cloud](https://feeny.ai/job/product-manager-cloud-modular-united-states-canada-9ks8p0xdqtk4) — United States / Canada - [Senior AI Kernel Engineer](https://feeny.ai/job/senior-ai-kernel-engineer-modular-united-states-canada-crsvzs19n2q8) — United States / Canada - [Cloud Inference Engineer](https://feeny.ai/job/cloud-inference-engineer-modular-united-states-canada-nnas25s7s2g3) — United States / Canada