
Inference Engineer – LLM & Speech AI at soket.ai (Bengaluru, India)
soket.ai· Bengaluru, India·
Role details
Work type
Onsite
Employment
Full-Time
Job description
We are looking for an experienced Inference Engineer to build and optimize high-performance inference systems for Large Language Models (LLMs), Speech AI systems, and multimodal AI workloads. You will work on deploying production-grade AI systems with a strong focus on:
- low latency,
- high throughput,
- GPU efficiency,
- scalable serving infrastructure,
- distributed inference,
- and cost optimization. This role sits at the intersection of:
- systems engineering,
- deep learning infrastructure,
- distributed computing,
- and production AI deployment. You will collaborate closely with:
- ML researchers,
- platform engineers,
- speech AI teams,
- and product engineering teams.
Why work at soket.ai
- Mission-Driven Culture: Opportunity to work on frontier AI (math, code, reasoning) with an explicit ethical mandate—rare in the for-profit AI lab space.
- High-Impact Work: Directly contribute to India’s sovereign AI infrastructure, with work presented to national leaders and global tech giants.
- Research-Heavy Environment: 38% of staff in technical roles; emphasis on deep research, open-source contributions, and building from scratch.
- Hybrid Work Model: Offices in Gurugram (HQ) and Bengaluru; hybrid policy with a mix of remote and on-site work.
- Growth Trajectory: Small, fast-growing team (30% YoY headcount growth) with active hiring in AI Research, Engineering, and Infrastructure—early-stage equity-like impact.
- Notable Perks: Collaboration with Google Cloud, access to high-end compute (NVIDIA GPUs, Slurm clusters), and a seat at the table in India’s AI policy discussions.