statsby.ai

Data Engineer ( Contract) at statsby.ai (Pune, India)

statsby.ai· Pune, India·

Role details

Work type
Onsite
Employment
Contract

Job description

We’re looking for a hands-on Lead Data Engineer to own and build the data architecture for an AI-driven product platform deployed across customer cloud environments.

This is an architect-who-still-ships role — you’ll design scalable data platforms, set engineering standards, review code, mentor engineers, and work hands-on with complex data and AI workloads.

What We’re Looking For

  • 10+ years of experience in production data engineering
  • 3+ years in Lead/Staff/Principal-level ownership
  • Expert in Databricks, Delta Lake, Unity Catalog & Medallion Architecture
  • Strong PySpark, Spark SQL, Python & SQL
  • Expertise in Kimball dimensional modelling & dbt
  • Working experience with Snowflake & Data Vault 2.0
  • Experience across Azure & AWS
  • Hands-on experience with LLM/RAG pipelines, embeddings & vector search
  • Strong understanding of data governance, security, observability & cost optimization
  • Experience with Kubernetes, Terraform/Bicep & CI/CD Experience in regulated/compliance-heavy environments is a plus

Why work at statsby.ai

  • Culture: Strong emphasis on blending human ingenuity with advanced AI; described as a team of "passionate problem solvers and innovators" committed to pushing the boundaries of technology in a high-stakes industry.
  • Work Environment: Hybrid workspace based in Pune, India. Employees engage in a combination of remote and on-site work.
  • Team: Very small team (~9 people) — offers significant ownership and impact for early employees. Technical staff makes up 56% of the team.
  • Reviews: Rated 5.0/5.0 on LinkedIn (2 reviews) with perfect scores in Work-Life (4.5), Compensation (5.0), Culture (5.0), and Career (5.0).
  • Engineering Focus: Opportunity to work at the intersection of clinical science, AI, and regulated software engineering — a niche with high barriers to entry and strong career value. Tech stack includes Databricks, Scala, AWS Lambda, Apache Spark, Python, Power BI, Microsoft Azure, Kubernetes, Snowflake, and React.

Application questions