soket.ai

Inference Engineer – LLM & Speech AI at soket.ai (Bengaluru, India)

soket.ai· Bengaluru, India·

Role details

Work type
Onsite
Employment
Full-Time

Job description

We are looking for an experienced Inference Engineer to build and optimize high-performance inference systems for Large Language Models (LLMs), Speech AI systems, and multimodal AI workloads. You will work on deploying production-grade AI systems with a strong focus on:

  • low latency,
  • high throughput,
  • GPU efficiency,
  • scalable serving infrastructure,
  • distributed inference,
  • and cost optimization. This role sits at the intersection of:
  • systems engineering,
  • deep learning infrastructure,
  • distributed computing,
  • and production AI deployment. You will collaborate closely with:
  • ML researchers,
  • platform engineers,
  • speech AI teams,
  • and product engineering teams.

Why work at soket.ai

  • Mission-Driven Culture: Opportunity to work on frontier AI (math, code, reasoning) with an explicit ethical mandate—rare in the for-profit AI lab space.
  • High-Impact Work: Directly contribute to India’s sovereign AI infrastructure, with work presented to national leaders and global tech giants.
  • Research-Heavy Environment: 38% of staff in technical roles; emphasis on deep research, open-source contributions, and building from scratch.
  • Hybrid Work Model: Offices in Gurugram (HQ) and Bengaluru; hybrid policy with a mix of remote and on-site work.
  • Growth Trajectory: Small, fast-growing team (30% YoY headcount growth) with active hiring in AI Research, Engineering, and Infrastructure—early-stage equity-like impact.
  • Notable Perks: Collaboration with Google Cloud, access to high-end compute (NVIDIA GPUs, Slurm clusters), and a seat at the table in India’s AI policy discussions.

Application questions