
Data Engineer ( Contract) at statsby.ai (Pune, India)
statsby.ai· Pune, India·
Role details
Work type
Onsite
Employment
Contract
Job description
We’re looking for a hands-on Lead Data Engineer to own and build the data architecture for an AI-driven product platform deployed across customer cloud environments.
This is an architect-who-still-ships role — you’ll design scalable data platforms, set engineering standards, review code, mentor engineers, and work hands-on with complex data and AI workloads.
What We’re Looking For
- 10+ years of experience in production data engineering
- 3+ years in Lead/Staff/Principal-level ownership
- Expert in Databricks, Delta Lake, Unity Catalog & Medallion Architecture
- Strong PySpark, Spark SQL, Python & SQL
- Expertise in Kimball dimensional modelling & dbt
- Working experience with Snowflake & Data Vault 2.0
- Experience across Azure & AWS
- Hands-on experience with LLM/RAG pipelines, embeddings & vector search
- Strong understanding of data governance, security, observability & cost optimization
- Experience with Kubernetes, Terraform/Bicep & CI/CD Experience in regulated/compliance-heavy environments is a plus
Why work at statsby.ai
- Culture: Strong emphasis on blending human ingenuity with advanced AI; described as a team of "passionate problem solvers and innovators" committed to pushing the boundaries of technology in a high-stakes industry.
- Work Environment: Hybrid workspace based in Pune, India. Employees engage in a combination of remote and on-site work.
- Team: Very small team (~9 people) — offers significant ownership and impact for early employees. Technical staff makes up 56% of the team.
- Reviews: Rated 5.0/5.0 on LinkedIn (2 reviews) with perfect scores in Work-Life (4.5), Compensation (5.0), Culture (5.0), and Career (5.0).
- Engineering Focus: Opportunity to work at the intersection of clinical science, AI, and regulated software engineering — a niche with high barriers to entry and strong career value. Tech stack includes Databricks, Scala, AWS Lambda, Apache Spark, Python, Power BI, Microsoft Azure, Kubernetes, Snowflake, and React.