--- title: 'Principal Data Engineer at Gather AI' canonical: 'https://feeny.ai/job/principal-data-engineer-gather-ai-india-gb7hgwbveemp' type: 'job' last_seen: '2026-09-10' --- # Principal Data Engineer at Gather AI - **Company:** Gather AI - **Location:** India - **Work type:** remote - **Posted:** 2026-07-14 - **Last confirmed live:** 2026-09-10 - **Apply:** https://job-boards.greenhouse.io/gatherai/jobs/5186046007 ## Job description ## About Us Are you ready to build the future of the supply chain? At Gather AI, we're not just creating software; we're pioneering a new era of warehouse intelligence. We've developed a groundbreaking, vision-powered platform that uses autonomous drones and existing equipment to capture real-time data, completely digitizing workflows that have historically been manual and error-prone. This means facilities operate smarter, safer, and more efficiently, ultimately redefining "on-time, in full" delivery. If you're looking for an opportunity to contribute to truly transformative technology and make a significant impact in a vital industry, Gather AI is the place for you. We're leading the charge in the rapidly evolving robotics industry, and we invite you to join us in reshaping the global supply chain, one intelligent warehouse at a time. ## About the Team You'll join the Data Platform team at its inception, helping establish the foundation from day one . Today, production and analytical workloads share a single database, and every product team defines its own metrics. This team exists to fix that: designing the warehouse, the transformation layers, and the semantic model that every product and dashboard will build on going forward, in close partnership with product, engineering, and security. ## About the Role Most data platform roles ask you to extend something someone else has already built. This one starts with a blank canvas. As Principal Data Engineer, you'll architect Gather AI's data foundation from the ground up: separating analytical workloads from live production traffic, building a semantic layer so metrics are defined once and stay consistent everywhere, and linking structured records to real drone imagery and video with full traceability. You'll prove the model end to end on Gather's drone product, then generalize it so every new product extends the foundation instead of rebuilding it, all while working as a Principal-level individual contributor with real influence across engineering, product, and leadership. ## What You'll Do - Architect a greenfield, multi-layer data warehouse (raw, refined, serving) that separates analytical workloads from production OLTP traffic. - Deliver a governed, self-service data-access layer for internal consumers first (Product, CSM, Deployment/Operations, and Leadership) as Phase 1, ahead of customer-facing conversational analytics. - Build a semantic and metrics layer so every metric, such as "scan accuracy by site," is defined once in code and stays identical across every dashboard and product, making self-service safe from metric drift. - Own the quality bar: 99%+ availability SLA with freshness guarantees, 100% traceability, zero cross-tenant leakage, 99.5%+ pipeline success, and no data loss. - Design tenant isolation, per-tenant cost attribution, and schema and row-level RBAC to scale toward hundreds of tenants (300+ target), not today's fleet size. - Own data-ingestion correctness at the boundary with the integration/backend team, covering data contracts, schema validation, and pipeline quality, so WMS data lands in the right place, shape, and time across WMS versions. - Stand up a data catalog and lineage layer (Purview as the Azure-native fit, DataHub as the open-source alternative) so every consumer can find data, see ownership, and trace lineage when a metric looks wrong. - Prove the foundation end to end on Gather's drone product, then generalize it so each new product extends the model instead of rebuilding it - Act as the connective tissue between product and ML (3DCC, damage detection). Link structured records to unstructured drone imagery and video with full traceability, and stand up the data-infra readiness for feature stores and annotation pipelines on one trusted foundation. ## What You'll Need - 10+ years in data engineering, with 3+ years architecting data platforms for data products, analytics, or AI-driven products. - Proven experience building a greenfield data warehouse and leading an OLTP to OLAP transition, not just maintaining an existing one. - Deep expertise designing multi-layer transformation architectures and reusable frameworks that scale across multiple product areas. - Expert SQL and dbt, hands-on ELT and orchestration, and large-scale or streaming data experience. - Production experience on a major cloud (Azure preferred, AWS or GCP acceptable), plus infrastructure as code and CI/CD. - Track record with data quality, security, governance, and multi-tenancy in production environments. - Data transformation and modeling that turns raw multi-source data into refined, serving-ready datasets (raw to refined to serving). - Pipeline orchestration and workflow automation for scheduling, dependency management, and reliable execution across data flows. - Large-scale and distributed processing of high-volume batch data. - Real-time and streaming ingestion that captures and processes event data as it arrives. - Semantic and metrics-layer design that defines business metrics once and serves them consistently to every consumer. - Serving-layer optimization for fast, low-latency consumption through wide and flattened tables and pre-computed metrics. - Cloud data engineering and infrastructure automation that provisions, deploys, and operates the platform reproducibly (cloud-native, infrastructure as code, CI/CD). - Data quality, observability, and lineage that ensure trust, freshness, and end-to-end traceability. - Security, governance, and multi-tenancy including tenant isolation, access control, and resiliency. - Multimodal data integration that links structured records to unstructured image and video (drone captures) with traceability. Ways of Working - Treats data as a product for internal consumers, not just a pipeline feeding dashboards. - Comfortable making long-lead architecture calls (platform, isolation model) with incomplete consensus. - Strong cross-functional collaborator, works closely with integration/backend, ML, product, customer success teams and internal analytics consumers. ## Nice to Have - Experience modeling structured data linked to unstructured or blob data such as images, video, or sensor files - Experience with feature stores, annotation pipelines, or ML data infrastructure supporting computer vision products. - IoT, edge, or device-telemetry background - BI or presentation-layer and dashboard design experience - Warehousing, logistics, or supply-chain domain knowledge ## About Gather AI ## Company Overview - **One-liner**: Gather AI builds the Physical AI platform for logistics, using AI-powered vision on drones and material handling equipment to digitize warehouse operations in real time. - **Entity Type**: Private – Series B ($74M total funding) - **Headquarters**: Pittsburgh, United States - **Founded**: 2017 - **Founders**: Sankalp Arora (CEO), Geetesh Dubey (Chief Systems Officer), Daniel Maturana (Chief ML Scientist) ## Core Business - Primary industry: Logistics and warehouse management (Physical AI) - Target customers: B2B, enterprise-scale logistics, manufacturing, retail, aerospace, automotive, and government – customers include GEODIS, NFI Industries, Barrett Distribution Centers, Axon, and dnata - Mission: To give every facility in a network the same continuous picture of the floor, the intelligence to understand it, and the workflows to act on it – bringing the same intelligence to the physical world that exists in the digital one. ## Products & Services - **Gather AI Prana Platform**: An end-to-end Physical AI platform with three layers: - **SEE**: AI-powered vision on forklift-mounted cameras (for moving items) and drones (for static inventory) that captures ten-plus structured data points per frame, achieving 99.9%+ inventory accuracy. - **THINK**: AI reasoning that investigates root cause across inventory, labor, and fulfillment, surfacing ranked, reasoned priorities before each shift. - **ACT**: Workflows that route decisions to the right person or system with pre-drafted tasks, closing the loop from capture to action. ## Market Standing - **Valuation**: Not publicly disclosed (private company) - **Key Metric**: 2.5x year-over-year bookings growth (250% growth); total funding of $74M (including a $40M Series B raised in 2026); average time to ROI is 6 months. - **Notable Investors/Partners**: Backed by Tribeca Venture (noted on LinkedIn); deployed with GEODIS, NFI Industries, Barrett Distribution Centers, Axon, dnata, and the U.S. Department of Classified Antiquities. - **Growth Signals**: Doubled operational footprint in 2026; named to Fast Company’s World’s Most Innovative Companies list for 2026; employee headcount grew ~68% year over year to 80+ employees; expanded into government sector; active job postings increased 333% YoY. ## Competitive Advantages - **Closed-loop Physical AI**: Combines vision (SEE), reasoning (THINK), and action (ACT) in one platform – a unique integration that turns real-time visual data into prioritized tasks and automated workflows. - **Deep domain expertise**: Founded by Carnegie Mellon robotics PhDs; CTO Andrew Hoffman was a founding engineer at Kiva Systems (now Amazon Robotics); team has decades of warehouse automation experience. - **Network-wide consistency**: The platform enforces the same standard across every facility, enabling fixes from one building to become playbooks for the whole network. - **99.9%+ accuracy** down to the case level, maintained continuously, not just during audits. ## Strategic Focus - Scale platform across multi-facility enterprise networks and new verticals (government, aerospace, automotive). - Continue investing in AI/ML autonomy and computer vision to expand the “THINK” layer’s root-cause analysis and predictive capabilities. - Enhance network-wide outcomes by replicating operational playbooks across customer sites. ## Why Work Here - **Culture**: Radical transparency, candor, accountability, speed, and a people-first approach. Decisions are made openly; everyone has a voice in shaping the company’s future. - **Work model**: Fully remote or remote-flexible options available. - **Benefits**: Competitive salary and equity, comprehensive medical/dental/vision, 401K, unlimited PTO, monthly wellness benefit, continuing education/tuition support, voluntary life insurance. - **Engineering culture**: Opportunities to work on cutting-edge AI, ML, autonomy, computer vision, and robotics alongside former Amazon Robotics and CMU experts. Real customer impact at scale. ## Sources 1. [gather.ai](https://gather.ai) 2. [gather.ai/company](https://gather.ai/company) 3. [gather.ai/careers](https://gather.ai/careers) 4. [gather.ai/about](https://www.gather.ai/about) 5. [linkedin.com/company/gather-ai](https://www.linkedin.com/company/gather-ai) ## Other roles at Gather AI - [ML Annotation QA Engineer](https://feeny.ai/job/ml-annotation-qa-engineer-gather-ai-open-to-fz96nnrq0f5c) — Open To, India - [Staff Autonomy Engineer (MHE Vision)](https://feeny.ai/job/staff-autonomy-engineer-mhe-vision-gather-ai-open-to-xwrzam65zcy5) — Open To, United States / Pittsburgh, PA - [Staff Autonomy Engineer (Drone)](https://feeny.ai/job/staff-autonomy-engineer-drone-gather-ai-open-to-9q1w90fjkesc) — Open To, United States / Pittsburgh, PA - [Senior Robotics Platform Engineer (Robotics/Infra)](https://feeny.ai/job/senior-robotics-platform-engineer-robotics-infra-gather-ai-open-to-24kmb5w32ev7) — Open To, United States / Pittsburgh, PA - [Staff Full Stack Engineer (Hybrid)](https://feeny.ai/job/staff-full-stack-engineer-hybrid-gather-ai-pittsburgh-fxc39zh48y3e) — Pittsburgh, PA - [Senior Manager, Hardware](https://feeny.ai/job/senior-manager-hardware-gather-ai-pittsburgh-g0g1fht8vhym) — Pittsburgh, PA - [UX Designer](https://feeny.ai/job/ux-designer-gather-ai-open-to-6zhb7g3cbtkw) — Open To, United States / Pittsburgh, PA - [Senior QA Engineer](https://feeny.ai/job/senior-qa-engineer-gather-ai-hyderabad-tme00py0w0t3) — Hyderabad, India - [Principal Machine Learning Scientist](https://feeny.ai/job/principal-machine-learning-scientist-gather-ai-open-to-00wtke9aacqh) — Open To, United States / Pittsburgh, PA - [Director of Machine Learning](https://feeny.ai/job/director-of-machine-learning-gather-ai-pittsburgh-rhtp64m9hkez) — Pittsburgh, PA / Open To, United States