--- title: 'AI Researcher — Inference Optimization at Featherless AI' canonical: 'https://feeny.ai/job/ai-researcher-inference-optimization-featherless-ai-world-5q35t4rfgjve' type: 'job' last_seen: '2026-09-08' --- # AI Researcher — Inference Optimization at Featherless AI - **Company:** Featherless AI - **Location:** World - **Employment:** full-time - **Work type:** remote - **Posted:** 2026-01-23 - **Last confirmed live:** 2026-09-08 - **Apply:** https://jobs.ashbyhq.com/featherlessai/2cc24de9-a81d-4a67-a1f5-bfbdeb8322a6 ## Job description ## ROLE OVERVIEW We are seeking an AI Researcher with deep experience in inference optimization to design, evaluate, and deploy high-performance inference systems for large-scale machine learning models. You will work at the intersection of model architecture, systems engineering, and hardware-aware optimization, improving latency, throughput, and cost efficiency across real-world production environments. ## KEY RESPONSIBILITIES - Research and develop techniques to optimize inference performance for large neural networks. - Improve latency, throughput, memory efficiency, and cost per inference. - Design and evaluate model-level optimizations (quantization, pruning, KV-cache optimization, architecture-aware simplifications). - Implement systems-level optimizations (dynamic batching, kernel fusion, multi-GPU inference, prefill vs decode optimization). - Benchmark inference workloads across hardware accelerators. - Collaborate with engineering teams to deploy optimized inference pipelines. - Translate research insights into production-ready improvements. ## REQUIRED QUALIFICATIONS - Strong background in machine learning, deep learning, or AI systems. - Hands-on experience optimizing inference for large-scale models. - Proficiency in Python and modern ML frameworks (e.g., PyTorch). - Experience with inference tooling (e.g., Triton, TensorRT, vLLM, ONNX Runtime). - Ability to design experiments and communicate results clearly. ## PREFERRED / NICE-TO-HAVE QUALIFICATIONS - Experience deploying production inference systems at scale. - Familiarity with distributed and multi-GPU inference. - Experience contributing to open-source ML or inference frameworks. - Authorship or co-authorship of peer-reviewed research papers in machine learning, systems, or related fields. - Experience working close to hardware (CUDA, ROCm, profiling tools). ## WHAT SUCCESS LOOKS LIKE - Measurable gains in latency, throughput, and cost efficiency. - Optimized inference systems running reliably in production. - Research ideas successfully translated into deployable systems. - Clear benchmarks and documentation that inform product decisions. ## RELEVANT RESEARCH AREAS (BONUS) - Long-context inference optimization - Speculative decoding - KV-cache compression and paging - Efficient decoding strategies - Hardware-aware inference design ## About Featherless AI ## Company Overview - **One-liner**: Featherless AI provides a serverless platform that offers API access to over 40,000 open-weight AI models from a single endpoint, designed for developers and enterprises. - **Entity Type**: Private (Series A) - **Headquarters**: San Francisco, California, United States - **Founded**: 2023 - **Founders**: Eugene Cheah (CEO, Co-Founder) ## Core Business - **Primary industry**: Artificial Intelligence Infrastructure / Serverless LLM Hosting - **Target customers**: B2B, serving developers, AI startups, and enterprises seeking scalable, cost-effective inference for open-source models. - **Mission or purpose**: To democratize access to all AI models by making them available for serverless inference, eliminating the need for server setup and complex infrastructure management. ## Products & Services - **Featherless API**: A unified API gateway providing instant access to over 40,000 open-weight models (e.g., DeepSeek, Llama, Mistral, Qwen, RWKV, GLM, Kimi) without setup or hosting. Pricing is flat-rate with unlimited tokens, starting at $25/month for up to 4 concurrent connections and 32K context, scaling to $200/month for higher tiers. Agent-specific plans ($100/month) include sandbox environments and persistent storage. The service emphasizes low latency, dependable uptime, and predictable costs. ## Market Standing - **Valuation/Market Cap**: Not publicly disclosed. - **Key Metric**: **Total Funding of $25M** — raised $5M in a Seed round (April 2025, led by Airbus Ventures) and $20M in a Series A round (announced ~May 2026, details still emerging). - **Notable Investors/Partners**: Airbus Ventures, Kickstart Ventures, Panache Ventures, BMW i Ventures, AMD Ventures, and 11 other investors. The platform is built by researchers contributing to RWKV, a Linux Foundation project. - **Growth Signals**: The company is on a rapid growth trajectory, with headcount increasing 46.7% year-over-year to 14 employees. The website boasts over 2,100 "stars" for its Discord community. The company has a global presence, operating in 9 countries (including Singapore, Canada, Czechia, UK, Belgium, and Sweden). Website traffic is strong (73,508 monthly visits, growing +19.9% month-over-month), and there are 43 active job postings, a 65.4% quarterly increase in hiring. The Series A announcement signals significant investor confidence. ## Competitive Advantages - **Extensive Model Library & Zero-Friction Access**: A single API key provides access to the entire Hugging Face trending library, including models up to 229B parameters, with no need to manage infrastructure. - **Unlimited-Token, Flat-Rate Pricing**: A strong differentiator in the "per-token" pricing era, offering predictable costs suitable for scaling, with tiers from $25 to $200/month. - **Build for Reliability & Performance**: Architecture designed for real workloads with low latency and dependable uptime, utilizing proprietary GPU orchestration and model load-balancing. - **Open-Source Roots & Community**: Built by researchers contributing to RWKV (a Linux Foundation project), the company is deeply embedded in the open-source AI ecosystem, which fosters trust and community-driven development. ## Strategic Focus - **Scaling the Platform & Enterprise Adoption**: The Series A funding will be used to expand AI infrastructure and grow the platform. The company is actively hiring for senior roles like Founding Account Executives and Business Development Reps, signaling a shift toward aggressive go-to-market and enterprise sales. - **Expanding Global Presence**: With a distributed team across the US, Europe, and Asia, the company is building a global, remote-first workforce. - **Technical Innovation**: Focused on continuous improvement of inference performance and cost-efficiency through their GPU orchestration system and model load-balancing. ## Why Work Here - **High-Growth Stage**: As a Series A startup with strong investor backing, this is an opportunity to join a company experiencing rapid scaling, which offers significant career growth and impact potential. - **Impact & Ownership**: Employees are likely to have high autonomy and a direct impact on the company's trajectory, from building core infrastructure to driving revenue. - **Remote-First & Global Team**: Based on the distributed headcount across 9 countries (US, Singapore, Canada, UK, Belgium, etc.), the company is clearly remote-first, offering flexibility in where you work. Job postings reflect opportunities in the US and Europe (e.g., Paris, Berlin). - **Cutting-Edge Technical Challenge**: The core work involves solving complex problems in AI inference, GPU orchestration, and MLOps, making it a compelling place for engineers and researchers passionate about AI infrastructure. - **Culture & Values**: The company's deep ties to open-source AI communities and its "flat-rate, no-surprises" pricing philosophy likely translate into a transparent, developer-friendly internal culture. The small, highly-skilled team (14 people) suggests a close-knit, high-performing environment. ## Sources 1. [Featherless.ai Website](https://featherless.ai/) 2. [Featherless AI LinkedIn](https://www.linkedin.com/company/feather-serverless-ai) 3. [Featherless AI Docs](https://featherless.ai/docs/overview) 4. [CB Insights Profile](https://www.cbinsights.com/company/recursal-ai) 5. [Featherless AI Jobs](https://jobs.ashbyhq.com/featherlessai) ## Other roles at Featherless AI - [Founding Account Executive (AI Cloud)](https://feeny.ai/job/founding-account-executive-ai-cloud-featherless-ai-us-krryw5zpmt01) — US &, Canada - [Founding Business Development Rep (AI Cloud US/CA)](https://feeny.ai/job/founding-business-development-rep-ai-cloud-us-ca-featherless-ai-us-y876z86dw2y3) — US &, Canada - [Chief of Staff](https://feeny.ai/job/chief-of-staff-featherless-ai-san-francisco-15ydsmveqwsq) — San Francisco, CA - [Content Marketer](https://feeny.ai/job/content-marketer-featherless-ai-europe-xnsbz1evvkbs) — Europe - [Business Development Rep (AI Cloud)](https://feeny.ai/job/business-development-rep-ai-cloud-featherless-ai-europe-hd5fymsqsh5h) — Europe - [AI Researcher — Training Optimization](https://feeny.ai/job/ai-researcher-training-optimization-featherless-ai-world-hyc2csn67p2p) — World - [AI Researcher – Multilingual Data](https://feeny.ai/job/ai-researcher-multilingual-data-featherless-ai-world-aj24t2jw442p) — World - [AI Researcher — AI Architecture Research](https://feeny.ai/job/ai-researcher-ai-architecture-research-featherless-ai-world-dprg8203nt10) — World - [AI Researcher — Distillation](https://feeny.ai/job/ai-researcher-distillation-featherless-ai-world-zxs1tx4mwq1f) — World - [Machine Learning Engineer — AI Architecture Research](https://feeny.ai/job/machine-learning-engineer-ai-architecture-research-featherless-ai-world-88g5fbqva8pd) — World