--- title: 'Databricks ( Lead) at statsby.ai' canonical: 'https://feeny.ai/job/databricks-lead-statsby-ai-pune-bxrq9nta6q43' type: 'job' last_seen: '2026-09-08' --- # Databricks ( Lead) at statsby.ai - **Company:** statsby.ai - **Location:** Pune, India - **Employment:** full-time - **Work type:** onsite - **Posted:** 2026-07-27 - **Last confirmed live:** 2026-09-08 - **Apply:** https://jobs.gem.com/statsby-ai/am9icG9zdDqWwhvYjzoA8PesvNVdW4Sx ## Job description We are looking for an experienced Data Engineer (Lead) with strong expertise in Databricks to design, develop, and maintain scalable data pipelines and data processing solutions. ## Key Responsibilities - Design, develop, and maintain scalable data pipelines using Databricks. - Develop data processing solutions using PySpark and Python. - Work with Apache Spark for large-scale data processing. - Build and optimize ETL/ELT pipelines. - Work with Delta Lake for data storage and processing. - Develop and manage workflows using Databricks Workflows. - Perform data transformation, cleansing, and validation. - Optimize Spark jobs and improve pipeline performance. - Collaborate with data engineers, analysts, and other technical teams. - Troubleshoot data pipeline issues and ensure data quality. Required Skills - 3–5 years of experience in Data Engineering. - Strong hands-on experience with Databricks. - Good knowledge of PySpark and Python. - Strong understanding of Apache Spark. - Experience with Delta Lake and data lake architecture. - Good knowledge of SQL. - Experience in building ETL/ELT data pipelines. - Understanding of cloud platforms such as Azure, AWS, or GCP. - Experience with Git and CI/CD is an added advantage. Good to Have - Experience with Azure Databricks. - Knowledge of Azure Data Factory, ADLS, or similar cloud data services. - Experience with DBT, Snowflake, or other modern data platforms. - Databricks certification is a plus. ## What We Offer - Opportunity to work on real-world Data Engineering and AI projects. - Exposure to modern data technologies and cloud platforms. - Collaborative and learning-focused work environment. - Opportunity to work with a growing technology team. ## About statsby.ai ## Company Overview - **One-liner**: Statsby Solutions is a pharma-native AI and data company that builds purpose-built platforms and provides consulting services to help pharmaceutical sponsors, biotechs, and CROs handle high-stakes clinical data and documentation in regulated environments. - **Entity Type**: Private (Bootstrapped/Service-based) - **Headquarters**: Pune, India - **Founded**: 2022 - **Founders**: Atul Pundhir (Co-Founder), Abhishek Kumar Mishra (Co-Founder & Chief Revenue Officer) ## Core Business - **Primary industry**: Pharmaceutical AI, Clinical Data Engineering, and Regulated AI/ML - **Target customers**: Pharma sponsors, biotechs, and Clinical Research Organizations (CROs) - **Mission/Purpose**: To create innovative solutions that empower organizations to harness their data, turning insights into action for greater efficiency, adaptability, and growth — with a focus on blending deep domain expertise with AI engineered for regulated environments where auditability, traceability, and scientific rigor are the baseline. ## Products & Services - **Revectra**: AI-powered Protocol-to-Platform Intelligence. Digitizes clinical trial protocols (new and legacy) into structured, version-controlled CDISC USDM 4.0 study definitions, exposes them via API, and generates downstream documents grounded in proprietary data, deployed on the client's cloud. - **Veractra**: AI-powered Clinical Study Report (CSR) generation. Produces ICH E3-compliant CSRs in days, featuring multi-stage numerical verification against source TLFs and sentence-level traceability, deployed entirely within the client's cloud environment. - **Data & AI Consulting**: End-to-end services including Clinical Data Platforms & Engineering (CDISC-aligned, GxP-ready cloud-native lakehouse architectures), Generative AI for Clinical Workflows (RAG systems, knowledge assistants, document automation), Agentic AI & Intelligent Automation (multi-agent systems for study start-up, PV case processing, submission assembly), and MLOps & Responsible AI (building auditable, compliant AI backbones for regulated environments). ## Market Standing - **Valuation/Funding**: Not publicly available (appears bootstrapped / service-funded) - **Key Metric**: Small, specialized team of 9 employees (as of latest data) - **Notable Clients/Partners**: Works with pharma sponsors, biotechs, and CROs; talent sourced from companies like Juniper Networks, Deloitte, Saama, Altair, and Triomics. - **Growth Signals**: -10% monthly headcount growth (likely reflects natural churn in a small consultancy); strong focus on deep regulatory compliance (GxP, ICH E6(R3), 21 CFR Part 11, FDA/EMA expectations) as a core differentiator; developing proprietary platforms (Revectra, Veractra) alongside consulting to scale impact. ## Competitive Advantages - **Pharma-native specialization**: Entirely focused on the pharmaceutical and clinical research industry, not a generalist AI consultancy. - **Regulatory engineering from day one**: Every platform is architected for GxP, ICH E6(R3), 21 CFR Part 11, and FDA/EMA expectations from the first architecture diagram — ensuring audit readiness and data provenance. - **Verification-first AI**: Veractra's multi-stage numerical verification against source TLFs and sentence-level traceability is a key differentiator for high-stakes regulatory document generation. - **Deployment flexibility**: All platforms are designed to be deployed entirely within the client's cloud environment, meeting enterprise security and compliance requirements. ## Strategic Focus - Deepening the "pharma-native" brand by building proprietary, purpose-built platforms (Revectra, Veractra) that move beyond consulting to product-led growth. - Expanding agentic AI capabilities for specific clinical workflows (study start-up, PV case processing, submission assembly, protocol deviation management). - Maintaining a "production-grade, regulated-first" engineering standard as the core value proposition. - Building long-term client relationships with a focus on what's running in the client's environment six months after delivery. ## Why Work Here - **Culture**: Strong emphasis on blending human ingenuity with advanced AI; described as a team of "passionate problem solvers and innovators" committed to pushing the boundaries of technology in a high-stakes industry. - **Work Environment**: Hybrid workspace based in Pune, India. Employees engage in a combination of remote and on-site work. - **Team**: Very small team (~9 people) — offers significant ownership and impact for early employees. Technical staff makes up 56% of the team. - **Reviews**: Rated 5.0/5.0 on LinkedIn (2 reviews) with perfect scores in Work-Life (4.5), Compensation (5.0), Culture (5.0), and Career (5.0). - **Engineering Focus**: Opportunity to work at the intersection of clinical science, AI, and regulated software engineering — a niche with high barriers to entry and strong career value. Tech stack includes Databricks, Scala, AWS Lambda, Apache Spark, Python, Power BI, Microsoft Azure, Kubernetes, Snowflake, and React. ## Sources 1. [statsby.ai](https://statsby.ai/) 2. [statsby.ai/about-us](https://statsby.ai/about-us/) 3. [statsby.ai/job-openings](https://statsby.ai/job-openings/) 4. [LinkedIn - Statsby Solutions](https://linkedin.com/company/statsby) 5. [Built In - Statsby Solutions](https://builtin.com/company/statsby-solutions) ## Other roles at statsby.ai - [Data Engineer ( Contract)](https://feeny.ai/job/data-engineer-contract-statsby-ai-pune-9xx9vyfj3k75) — Pune, India - [AI engineer](https://feeny.ai/job/ai-engineer-statsby-ai-pune-hmc4knbqx253) — Pune, India