--- title: 'AI engineer at statsby.ai' canonical: 'https://feeny.ai/job/ai-engineer-statsby-ai-pune-hmc4knbqx253' type: 'job' last_seen: '2026-09-08' --- # AI engineer at statsby.ai - **Company:** statsby.ai - **Location:** Pune, India - **Employment:** full-time - **Work type:** onsite - **Posted:** 2026-04-06 - **Last confirmed live:** 2026-09-08 - **Apply:** https://jobs.gem.com/statsby-ai/am9icG9zdDrodlZ2Qn6d3aCSnDCsjCV- ## Job description We are looking for an AI Engineer with ~2 years of hands-on experience in building, fine-tuning, or distilling language models. The ideal candidate has a strong foundation in Machine Learning and NLP, and is passionate about shipping production-grade AI systems. This role involves working across the full AI stack — from model development to deployment and observability. 🎓 Experience - 2+ years of professional experience in AI/ML engineering - Bachelor’s or Master’s degree in Computer Science, Machine Learning, Data Science, or a related field ✅ Core Requirements (Must-Have) - Proven experience in at least one of the following: Pre-training or training a small language model from scratch,  Fine-tuning large language models (LoRA, QLoRA, full fine-tuning) or Model distillation techniques - Hands-on experience building RAG pipelines, including vector databases (Pinecone, Weaviate, Qdrant, FAISS), embedding models, chunking strategies, and retrieval optimization - Strong proficiency in Python and ML frameworks like PyTorch, Hugging Face Transformers, and DeepSpeed or similar distributed training libraries - Solid understanding of transformer architecture, tokenization, attention mechanisms, and evaluation metrics (perplexity, BLEU, ROUGE, etc.) 📊 LLM Operations & Observability - Experience with LLM observability and evaluation tools (LangSmith, Weights & Biases, Arize, Helicone, or similar) - Familiarity with prompt engineering and systematic evaluation of LLM outputs (human-in-the-loop, automated benchmarks) - Understanding of LLM deployment considerations: latency optimization, caching strategies, token cost management, and rate limiting ✨ Nice to Have - Experience with agentic AI frameworks (LangChain, LlamaIndex, CrewAI, AutoGen) - Familiarity with model quantization (GGUF, GPTQ, AWQ) and serving frameworks (vLLM, TGI, Ollama, TensorRT-LLM) - Exposure to RLHF or DPO (Direct Preference Optimization) - Knowledge of MLOps practices: CI/CD, experiment tracking, model registries, Docker, Kubernetes - Experience with cloud AI services (AWS SageMaker, GCP Vertex AI, Azure ML) and GPU infrastructure management - Contributions to open-source AI/ML projects 🔍 Key Responsibilities - Design, train, fine-tune, and evaluate language models for production use cases - Build and maintain RAG pipelines and knowledge retrieval systems - Implement observability, monitoring, and evaluation frameworks for deployed LLM applications - Integrate AI into products through collaboration while staying ahead of AI trends and best practices ## About statsby.ai ## Company Overview - **One-liner**: Statsby Solutions is a pharma-native AI and data company that builds purpose-built platforms and provides consulting services to help pharmaceutical sponsors, biotechs, and CROs handle high-stakes clinical data and documentation in regulated environments. - **Entity Type**: Private (Bootstrapped/Service-based) - **Headquarters**: Pune, India - **Founded**: 2022 - **Founders**: Atul Pundhir (Co-Founder), Abhishek Kumar Mishra (Co-Founder & Chief Revenue Officer) ## Core Business - **Primary industry**: Pharmaceutical AI, Clinical Data Engineering, and Regulated AI/ML - **Target customers**: Pharma sponsors, biotechs, and Clinical Research Organizations (CROs) - **Mission/Purpose**: To create innovative solutions that empower organizations to harness their data, turning insights into action for greater efficiency, adaptability, and growth — with a focus on blending deep domain expertise with AI engineered for regulated environments where auditability, traceability, and scientific rigor are the baseline. ## Products & Services - **Revectra**: AI-powered Protocol-to-Platform Intelligence. Digitizes clinical trial protocols (new and legacy) into structured, version-controlled CDISC USDM 4.0 study definitions, exposes them via API, and generates downstream documents grounded in proprietary data, deployed on the client's cloud. - **Veractra**: AI-powered Clinical Study Report (CSR) generation. Produces ICH E3-compliant CSRs in days, featuring multi-stage numerical verification against source TLFs and sentence-level traceability, deployed entirely within the client's cloud environment. - **Data & AI Consulting**: End-to-end services including Clinical Data Platforms & Engineering (CDISC-aligned, GxP-ready cloud-native lakehouse architectures), Generative AI for Clinical Workflows (RAG systems, knowledge assistants, document automation), Agentic AI & Intelligent Automation (multi-agent systems for study start-up, PV case processing, submission assembly), and MLOps & Responsible AI (building auditable, compliant AI backbones for regulated environments). ## Market Standing - **Valuation/Funding**: Not publicly available (appears bootstrapped / service-funded) - **Key Metric**: Small, specialized team of 9 employees (as of latest data) - **Notable Clients/Partners**: Works with pharma sponsors, biotechs, and CROs; talent sourced from companies like Juniper Networks, Deloitte, Saama, Altair, and Triomics. - **Growth Signals**: -10% monthly headcount growth (likely reflects natural churn in a small consultancy); strong focus on deep regulatory compliance (GxP, ICH E6(R3), 21 CFR Part 11, FDA/EMA expectations) as a core differentiator; developing proprietary platforms (Revectra, Veractra) alongside consulting to scale impact. ## Competitive Advantages - **Pharma-native specialization**: Entirely focused on the pharmaceutical and clinical research industry, not a generalist AI consultancy. - **Regulatory engineering from day one**: Every platform is architected for GxP, ICH E6(R3), 21 CFR Part 11, and FDA/EMA expectations from the first architecture diagram — ensuring audit readiness and data provenance. - **Verification-first AI**: Veractra's multi-stage numerical verification against source TLFs and sentence-level traceability is a key differentiator for high-stakes regulatory document generation. - **Deployment flexibility**: All platforms are designed to be deployed entirely within the client's cloud environment, meeting enterprise security and compliance requirements. ## Strategic Focus - Deepening the "pharma-native" brand by building proprietary, purpose-built platforms (Revectra, Veractra) that move beyond consulting to product-led growth. - Expanding agentic AI capabilities for specific clinical workflows (study start-up, PV case processing, submission assembly, protocol deviation management). - Maintaining a "production-grade, regulated-first" engineering standard as the core value proposition. - Building long-term client relationships with a focus on what's running in the client's environment six months after delivery. ## Why Work Here - **Culture**: Strong emphasis on blending human ingenuity with advanced AI; described as a team of "passionate problem solvers and innovators" committed to pushing the boundaries of technology in a high-stakes industry. - **Work Environment**: Hybrid workspace based in Pune, India. Employees engage in a combination of remote and on-site work. - **Team**: Very small team (~9 people) — offers significant ownership and impact for early employees. Technical staff makes up 56% of the team. - **Reviews**: Rated 5.0/5.0 on LinkedIn (2 reviews) with perfect scores in Work-Life (4.5), Compensation (5.0), Culture (5.0), and Career (5.0). - **Engineering Focus**: Opportunity to work at the intersection of clinical science, AI, and regulated software engineering — a niche with high barriers to entry and strong career value. Tech stack includes Databricks, Scala, AWS Lambda, Apache Spark, Python, Power BI, Microsoft Azure, Kubernetes, Snowflake, and React. ## Sources 1. [statsby.ai](https://statsby.ai/) 2. [statsby.ai/about-us](https://statsby.ai/about-us/) 3. [statsby.ai/job-openings](https://statsby.ai/job-openings/) 4. [LinkedIn - Statsby Solutions](https://linkedin.com/company/statsby) 5. [Built In - Statsby Solutions](https://builtin.com/company/statsby-solutions) ## Other roles at statsby.ai - [Data Engineer ( Contract)](https://feeny.ai/job/data-engineer-contract-statsby-ai-pune-9xx9vyfj3k75) — Pune, India - [Databricks ( Lead)](https://feeny.ai/job/databricks-lead-statsby-ai-pune-bxrq9nta6q43) — Pune, India - [AI engineer](https://feeny.ai/job/ai-engineer-writer-san-francisco-59b8p2stqe3z) — San Francisco, CA - [AI engineer](https://feeny.ai/job/ai-engineer-lucis-paris-acp2wmr34xyv) — Paris, France - [AI engineer](https://feeny.ai/job/ai-engineer-symbiotic-security-morning-laffitte-83vvwh89s87n) — Morning Laffitte