--- title: 'Applied Machine Learning Engineer at Inference' canonical: 'https://feeny.ai/job/applied-machine-learning-engineer-inference-san-francisco-epe003kkzf34' type: 'job' last_seen: '2026-09-11' --- # Applied Machine Learning Engineer at Inference - **Company:** Inference - **Location:** San Francisco, CA - **Compensation:** $220k–$320k - **Employment:** full-time - **Work type:** onsite - **Posted:** 2026-01-05 - **Last confirmed live:** 2026-09-11 - **Apply:** https://jobs.ashbyhq.com/inference/b6798110-1218-4ac0-bcf4-9ee8b51810cb ## Job description Help us build the systems that train specialized AI models for the fastest-growing companies in the world. If you love taking cutting-edge ML techniques and turning them into products that ship, we'd love to meet you. About [Inference.net](http://Inference.net) [Inference.net](http://Inference.net) trains and hosts specialized language models for companies who want frontier-quality AI at a fraction of the cost. The models we train match GPT-5 accuracy but are smaller, faster, and up to 90% cheaper. Our platform handles everything end-to-end: distillation, training, evaluation, and planet-scale hosting. We are a well-funded ten-person team of engineers who work in-person in downtown San Francisco on difficult, high-impact engineering problems. Everyone on the team has been writing code for over 10 years, and has founded and run their own software companies. We are high-agency, adaptable, and collaborative. We value creativity alongside technical prowess and humility. We work hard, and deeply enjoy the work that we do. Most of us are in the office 4 days a week in SF; hybrid works for Bay Area candidates. ## About the Role You will be responsible for building and improving the core ML systems that power our custom model training platform, while also applying these systems directly for customers. Your role sits at the intersection of applied research and production engineering. You'll lead projects from data intake to trained model, building the infrastructure and tooling along the way. Your north star is model quality at scale, measured by how well our custom models match frontier performance, how efficiently we can train and serve them, and how smoothly we can deliver results to our customers. You'll own the full training lifecycle: processing data, creating dashboards for visibility, training models using our frameworks, running evaluations, and shipping results. This role reports directly to the founding team. You'll have the autonomy, a large compute budget / GPU reservation, and technical support to push the boundaries of what's possible in custom model training. ## Key Responsibilities - Lead projects from from data intake through the full training pipeline, including processing, cleaning, and preparing datasets for model training - Build and maintain data processing pipelines for aggregating, transforming, and validating training data - Create dashboards and visualization tools to display training metrics, data quality, and model performance - Train models using our internal frameworks and iterate based on evaluation results - Develop robust benchmarks and evaluation frameworks that ensure custom models match or exceed frontier performance - Build systems to automate portions of the training workflow, reducing manual intervention and improving consistency - Take research features and ship them into production settings - Apply the latest techniques in SFT, RL, and model optimization to improve training quality and efficiency - Collaborate with infrastructure engineers to scale training across our GPU fleet - Deeply understand customer use cases to inform training strategies and surface edge cases ## Requirements - 2+ years of experience training AI models using PyTorch - Hands-on experience with post-training LLMs using SFT or RL - Strong understanding of transformer architectures and how they're trained - Experience with LLM-specific training frameworks (e.g., Hugging Face Transformers, DeepSpeed, Axolotl, or similar) - Experience training on NVIDIA GPUs - Strong data processing skills and comfortable building ETL pipelines and working with large datasets - Track record of creating benchmarks and evaluations - Ability to take research techniques and apply them to production systems Nice-to-Have - Experience with model distillation or knowledge transfer - Experience building dashboards and data visualization tools - Familiarity with vision encoders and multimodal models - Experience with distributed training at scale - Contributions to open-source ML projects You don't need to tick every box. Curiosity and the ability to learn quickly matter more. ## Compensation We offer competitive compensation, equity in a high-growth startup, and comprehensive benefits. The base salary range for this role is $220,000 - $320,000, plus equity and benefits, depending on experience. ## Equal Opportunity [Inference.net](http://Inference.net) is an equal opportunity employer. We welcome applicants from all backgrounds and don't discriminate based on race, color, religion, gender, sexual orientation, national origin, genetics, disability, age, or veteran status. If you're excited about building the future of custom AI infrastructure, we'd love to hear from you. Please send your resume and GitHub to amar@inference.net and/or apply here on Ashby. ## About Inference ## Company Overview - **One-liner**: Inference.net provides a marketplace and infrastructure for AI-native teams to deploy, observe, evaluate, and train custom LLMs at dramatically reduced costs by utilizing otherwise wasted GPU capacity from data centers. - **Entity Type**: Private (Seed stage) - **Headquarters**: San Francisco, California, United States - **Founded**: 2023 - **Founders**: Amarjot Singh (Co-Founder), Ibrahim Ahmed (Co-Founder, CTO) ## Core Business - **Primary Industry**: AI Inference Infrastructure / Software Development - **Target Customers**: B2B; AI-native companies, startups, and enterprises spending over $50k/month on closed-source AI providers; digital banks; decentralized networks; and high-volume AI applications. - **Mission/Purpose**: "We believe efficient markets for AI inference will drive the widespread proliferation of artificial intelligence over the next decades, leading to unprecedented human flourishing on Earth and beyond. We aim to accelerate this process." ## Products & Services - **Inference.net API**: A pay-as-you-go, OpenAI-compatible API for serving open-source, custom, and fine-tuned LLMs. Offers 50-90% discounts compared to providers like OpenAI and Anthropic by aggregating spot compute from underutilized data center GPU capacity. - **Catalyst Deploy**: A deployment platform for hosting LLMs at massive scale across public cloud, private cloud, or hybrid environments, with a claimed 99.99% uptime. - **Catalyst Observe**: An LLM observability tool that traces every request path (prompts, tool calls, responses, downstream providers) and monitors latency, reliability, usage patterns, and quality signals. - **Catalyst Evaluate**: A model evaluation system that scores quality across any model or metric, using production traces to validate new model variants against baseline behavior before deployment. - **Catalyst Train**: Automatic fine-tuning workflows that turn production traces into training datasets. Allows users to train custom frontier-level language models fine-tuned to specific quality, cost, and latency targets in minutes. ## Market Standing - **Valuation/Market Cap**: Not publicly disclosed. - **Total Funding**: $11.8M (Series Seed, announced October 14, 2025). - **Notable Investors**: Led by Multicoin Capital and a16z CSX, with participation from Topology Ventures, Founders, Inc., and a group of angel investors. - **Growth Signals**: 100% headcount growth year-over-year (from 5 to 10 employees). LinkedIn follower growth of +173.6% year-over-year. Has deployed custom models for "some of the fastest-growing AI-native companies in the world," including a digital bank with 120M+ customers and a nutrition tracking app that scaled to 10M+ users. ## Competitive Advantages - **Unique Business Model**: Acts as a spot market for perishable GPU compute, purchasing underutilized data center capacity in small chunks. This creates an inherent cost advantage, passing 50-90% savings to customers. - **Custom Model Economics**: Their approach trains models up to 100x smaller than GPT-5-class systems that match or exceed frontier model performance for specific tasks, running 2-3x faster and costing up to 90% less. - **Differentiation Focus**: Pitching against "renting intelligence" from closed providers, arguing that custom models trained on proprietary data become a moat competitors cannot replicate. - **SOC 2 Type II Compliant**: Full compliance and operational oversight, enabling enterprise adoption. ## Strategic Focus - **Expand R&D**: Using seed funding to push the frontiers of model and infrastructure performance. - **Scale Customer Acquisition**: Targeting companies spending over $50k/month on closed-source AI, offering to cut costs and improve performance within 4 weeks. - **Continuous Improvement Loops**: Building systems that retrain models on fresh production data as use cases evolve, creating "models that get better every cycle." - **Multi-Model Platform**: Supporting integration with both provider-hosted models (OpenAI, Anthropic, Gemini) and open-source models on optimized infrastructure. ## Why Work Here - **High-Impact Role in AI Infrastructure**: Working at the intersection of cutting-edge LLM research and practical infrastructure engineering, directly enabling the economics of AI for other companies. - **Tiny, High-Caliber Team**: Only 10 employees, plus 4 active job openings, suggesting a lean, high-autonomy culture where individuals have outsized impact. - **Strong Backing**: Backed by top-tier investors including Multicoin Capital and a16z, providing stability and resources despite being an early-stage company. - **Office Policy**: On-site / In-Office in San Francisco, CA (HQ in SoMa area). Employees work from a physical office, with typical time on-site being "None" (indicating potential flexibility). - **Culture Signals**: Described as mission-driven ("human flourishing on Earth and beyond"), with a focus on technical excellence and economic efficiency. The company openly shares its philosophy and strategy in blog posts. - **Active Roles**: Looking for Machine Learning Researchers, Fullstack Engineers (Frontend Focus), Senior Software Engineers (Model Performance), and Applied Machine Learning Engineers – all of which touch core product and research. ## Sources 1. [Inference.net Website](https://inference.net/) 2. [Inference.net Company Page](https://inference.net/company/) 3. [LinkedIn Page](https://www.linkedin.com/company/inference-net) 4. [Built In Profile](https://builtin.com/company/inferencenet) 5. [Seed Round Announcement](https://inference.net/blog/seed-round/) ## Other roles at Inference - [Senior Software Engineer - Model Performance](https://feeny.ai/job/senior-software-engineer-model-performance-inference-san-francisco-ey1dbf876pf7) — San Francisco, CA - [Machine Learning Researcher](https://feeny.ai/job/machine-learning-researcher-inference-san-francisco-hm492bpfacsw) — San Francisco, CA - [Fullstack Engineer - Frontend Focus](https://feeny.ai/job/fullstack-engineer-frontend-focus-inference-san-francisco-1s8ak7cryhnc) — San Francisco, CA - [Filmmaker / Storyteller](https://feeny.ai/job/filmmaker-storyteller-inference-san-francisco-ydsw1y7frf1a) — San Francisco, CA - [Applied Machine Learning Engineer](https://feeny.ai/job/applied-machine-learning-engineer-vulcan-elements-research-triangle-park-22929pdy6we4) — Research Triangle Park, NC - [Applied Machine Learning Engineer](https://feeny.ai/job/applied-machine-learning-engineer-cohere-london-w8vgzjewysvs) — London, United Kingdom - [Applied Machine Learning Engineer](https://feeny.ai/job/applied-machine-learning-engineer-fireworks-ai-new-york-q13mytf7xf10) — New York, NY / San Mateo, CA