--- title: 'Senior Software Engineer - Model Performance at Inference' canonical: 'https://feeny.ai/job/senior-software-engineer-model-performance-inference-san-francisco-ey1dbf876pf7' type: 'job' last_seen: '2026-09-11' --- # Senior Software Engineer - Model Performance at Inference - **Company:** Inference - **Location:** San Francisco, CA - **Employment:** full-time - **Work type:** onsite - **Posted:** 2026-01-21 - **Last confirmed live:** 2026-09-11 - **Apply:** https://jobs.ashbyhq.com/inference/7a2963de-1b33-4dfc-b711-990faa93a6a5 ## Job description Help us make inference blazingly fast. If you love squeezing every last drop of performance out of GPUs, diving deep into CUDA kernels, and turning optimization techniques into production systems, we'd love to meet you. About [Inference.net](http://Inference.net) [Inference.net](http://Inference.net) trains and hosts specialized language models for companies that need frontier-quality AI at a fraction of the cost. The models we train match GPT-5 accuracy but are smaller, faster, and up to 90% cheaper. Our platform handles everything end-to-end: distillation, training, evaluation, and planet-scale hosting. We are a well-funded ten-person team of engineers who work in-person in downtown San Francisco on difficult, high-impact engineering problems. Everyone on the team has been writing code for over 10 years, and has founded and run their own software companies. We are high-agency, adaptable, and collaborative. We value creativity alongside technical prowess and humility. We work hard, and deeply enjoy the work that we do. Most of us are in the office 4 days a week in SF; hybrid works for Bay Area candidates. ## About the Role You will be responsible for making our inference stack as fast and efficient as possible. Your work spans from implementing known optimization techniques to experimenting with novel approaches, always with the goal of serving models faster and cheaper at scale. Your north star is inference performance: latency, throughput, cost efficiency, and how quickly we can bring new model architectures into production. You'll work across the full inference stack—from CUDA kernels to serving frameworks—to find and eliminate bottlenecks. This role reports directly to the founding team. You'll have the autonomy, a large compute budget, and technical support to push the limits of what's possible in model serving. ## Key Responsibilities - Implement and productionize optimization techniques including quantization, speculative decoding, KV cache optimization, continuous batching, and LoRA serving - Deep dive into inference frameworks (vLLM, SGLang, TensorRT-LLM) and underlying libraries to debug and improve performance - Profile and optimize CUDA kernels and GPU utilization across our serving infrastructure - Add support for new model architectures, ensuring they meet our performance standards before going to production - Experiment with novel inference techniques and bring successful approaches into production - Build tooling and benchmarks to measure and track inference performance across our fleet - Collaborate with applied ML engineers to ensure trained models can be served efficiently ## Requirements - 2+ years of experience in ML systems, inference optimization, or GPU programming - Strong proficiency in Python and familiarity with C++ - Hands-on experience with LLM inference frameworks (vLLM, SGLang, TensorRT-LLM, or similar) - Deep understanding of GPU architecture and experience profiling GPU workloads - Familiarity with LLM optimization techniques (quantization, speculative decoding, continuous batching, KV cache management) - Experience with PyTorch and understanding of how models execute on hardware - Track record of measurably improving system performance Nice-to-Have - Experience with CUDA programming - Familiarity with serving non-LLM models (TTS, vision, embeddings) - Experience with distributed inference and multi-GPU serving - Contributions to open-source inference frameworks - Experience with Docker and Kubernetes You don't need to tick every box. Curiosity and the ability to learn quickly matter more. ## Compensation We offer competitive compensation, equity in a high-growth startup, and comprehensive benefits. The base salary range for this role is $220,000 - $320,000, plus equity and benefits, depending on experience. ## Equal Opportunity [Inference.net](http://Inference.net) is an equal opportunity employer. We welcome applicants from all backgrounds and don't discriminate based on race, color, religion, gender, sexual orientation, national origin, genetics, disability, age, or veteran status. If you're excited about making AI inference faster for everyone, we'd love to hear from you. Please send your resume and GitHub to amar@inference.net and/or apply here on Ashby. ## About Inference ## Company Overview - **One-liner**: Inference.net provides a marketplace and infrastructure for AI-native teams to deploy, observe, evaluate, and train custom LLMs at dramatically reduced costs by utilizing otherwise wasted GPU capacity from data centers. - **Entity Type**: Private (Seed stage) - **Headquarters**: San Francisco, California, United States - **Founded**: 2023 - **Founders**: Amarjot Singh (Co-Founder), Ibrahim Ahmed (Co-Founder, CTO) ## Core Business - **Primary Industry**: AI Inference Infrastructure / Software Development - **Target Customers**: B2B; AI-native companies, startups, and enterprises spending over $50k/month on closed-source AI providers; digital banks; decentralized networks; and high-volume AI applications. - **Mission/Purpose**: "We believe efficient markets for AI inference will drive the widespread proliferation of artificial intelligence over the next decades, leading to unprecedented human flourishing on Earth and beyond. We aim to accelerate this process." ## Products & Services - **Inference.net API**: A pay-as-you-go, OpenAI-compatible API for serving open-source, custom, and fine-tuned LLMs. Offers 50-90% discounts compared to providers like OpenAI and Anthropic by aggregating spot compute from underutilized data center GPU capacity. - **Catalyst Deploy**: A deployment platform for hosting LLMs at massive scale across public cloud, private cloud, or hybrid environments, with a claimed 99.99% uptime. - **Catalyst Observe**: An LLM observability tool that traces every request path (prompts, tool calls, responses, downstream providers) and monitors latency, reliability, usage patterns, and quality signals. - **Catalyst Evaluate**: A model evaluation system that scores quality across any model or metric, using production traces to validate new model variants against baseline behavior before deployment. - **Catalyst Train**: Automatic fine-tuning workflows that turn production traces into training datasets. Allows users to train custom frontier-level language models fine-tuned to specific quality, cost, and latency targets in minutes. ## Market Standing - **Valuation/Market Cap**: Not publicly disclosed. - **Total Funding**: $11.8M (Series Seed, announced October 14, 2025). - **Notable Investors**: Led by Multicoin Capital and a16z CSX, with participation from Topology Ventures, Founders, Inc., and a group of angel investors. - **Growth Signals**: 100% headcount growth year-over-year (from 5 to 10 employees). LinkedIn follower growth of +173.6% year-over-year. Has deployed custom models for "some of the fastest-growing AI-native companies in the world," including a digital bank with 120M+ customers and a nutrition tracking app that scaled to 10M+ users. ## Competitive Advantages - **Unique Business Model**: Acts as a spot market for perishable GPU compute, purchasing underutilized data center capacity in small chunks. This creates an inherent cost advantage, passing 50-90% savings to customers. - **Custom Model Economics**: Their approach trains models up to 100x smaller than GPT-5-class systems that match or exceed frontier model performance for specific tasks, running 2-3x faster and costing up to 90% less. - **Differentiation Focus**: Pitching against "renting intelligence" from closed providers, arguing that custom models trained on proprietary data become a moat competitors cannot replicate. - **SOC 2 Type II Compliant**: Full compliance and operational oversight, enabling enterprise adoption. ## Strategic Focus - **Expand R&D**: Using seed funding to push the frontiers of model and infrastructure performance. - **Scale Customer Acquisition**: Targeting companies spending over $50k/month on closed-source AI, offering to cut costs and improve performance within 4 weeks. - **Continuous Improvement Loops**: Building systems that retrain models on fresh production data as use cases evolve, creating "models that get better every cycle." - **Multi-Model Platform**: Supporting integration with both provider-hosted models (OpenAI, Anthropic, Gemini) and open-source models on optimized infrastructure. ## Why Work Here - **High-Impact Role in AI Infrastructure**: Working at the intersection of cutting-edge LLM research and practical infrastructure engineering, directly enabling the economics of AI for other companies. - **Tiny, High-Caliber Team**: Only 10 employees, plus 4 active job openings, suggesting a lean, high-autonomy culture where individuals have outsized impact. - **Strong Backing**: Backed by top-tier investors including Multicoin Capital and a16z, providing stability and resources despite being an early-stage company. - **Office Policy**: On-site / In-Office in San Francisco, CA (HQ in SoMa area). Employees work from a physical office, with typical time on-site being "None" (indicating potential flexibility). - **Culture Signals**: Described as mission-driven ("human flourishing on Earth and beyond"), with a focus on technical excellence and economic efficiency. The company openly shares its philosophy and strategy in blog posts. - **Active Roles**: Looking for Machine Learning Researchers, Fullstack Engineers (Frontend Focus), Senior Software Engineers (Model Performance), and Applied Machine Learning Engineers – all of which touch core product and research. ## Sources 1. [Inference.net Website](https://inference.net/) 2. [Inference.net Company Page](https://inference.net/company/) 3. [LinkedIn Page](https://www.linkedin.com/company/inference-net) 4. [Built In Profile](https://builtin.com/company/inferencenet) 5. [Seed Round Announcement](https://inference.net/blog/seed-round/) ## Other roles at Inference - [Machine Learning Researcher](https://feeny.ai/job/machine-learning-researcher-inference-san-francisco-hm492bpfacsw) — San Francisco, CA - [Applied Machine Learning Engineer](https://feeny.ai/job/applied-machine-learning-engineer-inference-san-francisco-epe003kkzf34) — San Francisco, CA - [Fullstack Engineer - Frontend Focus](https://feeny.ai/job/fullstack-engineer-frontend-focus-inference-san-francisco-1s8ak7cryhnc) — San Francisco, CA - [Filmmaker / Storyteller](https://feeny.ai/job/filmmaker-storyteller-inference-san-francisco-ydsw1y7frf1a) — San Francisco, CA