Featherless AI

AI Researcher — Inference Optimization at Featherless AI (World)

Featherless AI· World·

Role details

Work type
Remote
Employment
Full-Time

Job description

ROLE OVERVIEW

We are seeking an AI Researcher with deep experience in inference optimization to design, evaluate, and deploy high-performance inference systems for large-scale machine learning models. You will work at the intersection of model architecture, systems engineering, and hardware-aware optimization, improving latency, throughput, and cost efficiency across real-world production environments.

KEY RESPONSIBILITIES

  • Research and develop techniques to optimize inference performance for large neural networks.
  • Improve latency, throughput, memory efficiency, and cost per inference.
  • Design and evaluate model-level optimizations (quantization, pruning, KV-cache optimization, architecture-aware simplifications).
  • Implement systems-level optimizations (dynamic batching, kernel fusion, multi-GPU inference, prefill vs decode optimization).
  • Benchmark inference workloads across hardware accelerators.
  • Collaborate with engineering teams to deploy optimized inference pipelines.
  • Translate research insights into production-ready improvements.

REQUIRED QUALIFICATIONS

  • Strong background in machine learning, deep learning, or AI systems.
  • Hands-on experience optimizing inference for large-scale models.
  • Proficiency in Python and modern ML frameworks (e.g., PyTorch).
  • Experience with inference tooling (e.g., Triton, TensorRT, vLLM, ONNX Runtime).
  • Ability to design experiments and communicate results clearly.

PREFERRED / NICE-TO-HAVE QUALIFICATIONS

  • Experience deploying production inference systems at scale.
  • Familiarity with distributed and multi-GPU inference.
  • Experience contributing to open-source ML or inference frameworks.
  • Authorship or co-authorship of peer-reviewed research papers in machine learning, systems, or related fields.
  • Experience working close to hardware (CUDA, ROCm, profiling tools).

WHAT SUCCESS LOOKS LIKE

  • Measurable gains in latency, throughput, and cost efficiency.
  • Optimized inference systems running reliably in production.
  • Research ideas successfully translated into deployable systems.
  • Clear benchmarks and documentation that inform product decisions.

RELEVANT RESEARCH AREAS (BONUS)

  • Long-context inference optimization
  • Speculative decoding
  • KV-cache compression and paging
  • Efficient decoding strategies
  • Hardware-aware inference design

Why work at Featherless AI

  • High-Growth Stage: As a Series A startup with strong investor backing, this is an opportunity to join a company experiencing rapid scaling, which offers significant career growth and impact potential.
  • Impact & Ownership: Employees are likely to have high autonomy and a direct impact on the company's trajectory, from building core infrastructure to driving revenue.
  • Remote-First & Global Team: Based on the distributed headcount across 9 countries (US, Singapore, Canada, UK, Belgium, etc.), the company is clearly remote-first, offering flexibility in where you work. Job postings reflect opportunities in the US and Europe (e.g., Paris, Berlin).
  • Cutting-Edge Technical Challenge: The core work involves solving complex problems in AI inference, GPU orchestration, and MLOps, making it a compelling place for engineers and researchers passionate about AI infrastructure.
  • Culture & Values: The company's deep ties to open-source AI communities and its "flat-rate, no-surprises" pricing philosophy likely translate into a transparent, developer-friendly internal culture. The small, highly-skilled team (14 people) suggests a close-knit, high-performing environment.

Application questions