--- title: 'Research Engineer Intern, Evaluations at tensorstax.com' canonical: 'https://feeny.ai/job/research-engineer-intern-evaluations-tensorstax-com-san-francisco-hbe0m5zncybc' type: 'job' last_seen: '2026-09-09' --- # Research Engineer Intern, Evaluations at tensorstax.com - **Company:** tensorstax.com - **Location:** San Francisco, CA - **Employment:** full-time - **Work type:** onsite - **Posted:** 2025-03-18 - **Last confirmed live:** 2026-09-09 - **Apply:** https://jobs.gem.com/tensorstax-com/am9icG9zdDrJzq6tnUqO21FNSAnqLelJ ## Job description Research Engineer Intern, Evaluations & Benchmarks Location: San Francisco (Hybrid) About TensorStax: TensorStax is building fully autonomous AI systems to manage and optimize mission-critical data infrastructure. Our research integrates reinforcement learning and language models to enhance reasoning over large-scale data lakes and warehouses, detect failures in pipelines, and autonomously construct and optimize data workflows with high precision. We are looking for a Research Engineer Intern to design evaluation frameworks and benchmarks that assess the autonomy, adaptability, and reliability of AI agents in data engineering environments. This role is ideal for candidates passionate about AI evaluations, language model benchmarking, and autonomous data systems. What You’ll Do: - Develop evaluation environments to test AI agents' ability to reason, plan, and act autonomously within mission-critical data pipelines. - Design benchmarks to assess model capabilities in failure detection, pipeline optimization, and agentic decision-making in data workflows. - Implement automated assessment frameworks for language model-based agents operating over data lakes and warehouses. - Work with synthetic and real-world datasets to create robust testing environments for AI-driven data automation. - Collaborate with research engineers to refine reward shaping strategies, guiding models toward more efficient and agentic behaviors in data-intensive tasks. What We’re Looking For: - Experience in language model research, with a focus on benchmarking LLMs in mission-critical domains. - Strong background in AI evaluation methodologies, reinforcement learning, and RLHF techniques. - Familiarity with benchmarking language models for structured and unstructured data tasks. - Proficiency in Python and experience with ML frameworks like PyTorch or JAX. - Hands-on experience with data lakes, warehouses, and data engineering tools (Snowflake, BigQuery, dbt, Spark, Kafka). - High agency—proactive, resourceful, and comfortable working in a fast-paced research environment with minimal supervision. - Attention to detail—ability to design rigorous, reproducible experiments and evaluations. Bonus Points: - Contributions to open-source AI benchmarks (e.g., SweBench, BIRD, SPIDER). - Contributions to open-source agentic frameworks. - Experience developing custom RL environments for AI evaluation. - Strong understanding of ETL, ELT, and data transformation pipelines. Benefits: - Competitive internship stipend. - 100% employer-covered health, dental, and vision insurance (for eligible interns). - Access to Bay Club or Equinox in San Francisco. - Opportunity to work at the cutting edge of AI evaluations and autonomous data engineering research. ## About tensorstax.com ## Company Overview - **One-liner**: TensorStax provides an autonomous AI agentic platform that builds, validates, and maintains production-grade data pipelines directly on a company's existing infrastructure. - **Entity Type**: Private (Seed stage) - **Headquarters**: San Francisco, California, United States - **Founded**: Year not publicly specified, but active by 2025. - **Founders**: Not publicly listed on main sources. ## Core Business - **Primary industry**: Software Development / Artificial Intelligence / Data Engineering - **Target customers**: Enterprise data engineering teams, data platform teams, and any organization running complex data pipelines on modern data stacks (Snowflake, BigQuery, Redshift, Databricks, etc.). - **Mission**: To make advanced data infrastructure accessible to all companies by addressing the critical shortage of specialized talent, and to supercharge the rigid domain of data engineering with autonomous AI [globenewswire.com](https://www.globenewswire.com/news-release/2025/05/12/3078960/0/en/tensorstax-raises-5m-to-build-deterministic-ai-agents-for-data-engineers.html). ## Products & Services - **[Agentic Data OS / Platform]**: A SaaS platform that uses autonomous AI agents to plan, generate, and maintain production-grade data pipelines. It integrates directly with existing tooling (dbt, Airflow, Spark, Snowflake, BigQuery, Databricks) and runs in the customer's own cloud. Features include: - **Pipeline Generation**: AI generates structured pipeline plans and code based on infrastructure and schemas. - **LLM Compiler**: A proprietary deterministic control layer that validates syntax, resolves dependencies, and normalizes tool interfaces, claiming to boost agent success rates from 40-50% to 85-90% [globenewswire.com](https://www.globenewswire.com/news-release/2025/05/12/3078960/0/en/tensorstax-raises-5m-to-build-deterministic-ai-agents-for-data-engineers.html). - **Self-Healing**: Detects pipeline failures and automatically creates GitHub pull requests to fix failing code. - **Modeling & Testing**: Auto-generates dbt models, tests, and assertions with strong schema typing. - **Security**: SOC2 Type 2 compliant, integrates with HashiCorp Vault for credential management, supports self-hosted deployment, and adheres to GDPR and RBAC. The platform never stores or accesses raw data, only metadata and code [tensorstax.com](https://www.tensorstax.com/?tpcc=NL_Marketing). ## Market Standing - **Valuation/Market Cap**: Not disclosed. - **Key Metric**: **$5.3M** in total funding. - **Funding Round**: Seed Round (announced May 12, 2025), led by **Glasswing Ventures**, with participation from **Bee Partners** and **S3 Ventures** [globenewswire.com](https://www.globenewswire.com/news-release/2025/05/12/3078960/0/en/tensorstax-raises-5m-to-build-deterministic-ai-agents-for-data-engineers.html). - **Headcount**: ~4 employees [linkedin.com](https://linkedin.com/company/tensorstax). - **LinkedIn Followers**: Over 12,400 (YoY growth of +126%), indicating strong brand interest despite a tiny team [linkedin.com](https://linkedin.com/company/tensorstax). - **Growth Signals**: Very early-stage, post-Seed with a clear product-market fit hypothesis. High interest from the data community. The company’s technology (LLM Compiler) is a potential key differentiator in a hot market (AI for Data Engineering). ## Competitive Advantages - **Deterministic AI for a Strict Domain**: Unlike general-purpose coding assistants, TensorStax’s proprietary LLM Compiler is purpose-built for the constraints of data engineering, resulting in much higher success rates in production environments (85-90% vs. 40-50% for generic models) [globenewswire.com](https://www.globenewswire.com/news-release/2025/05/12/3078960/0/en/tensorstax-raises-5m-to-build-deterministic-ai-agents-for-data-engineers.html). - **Deep Native Integration**: The platform integrates directly with an enterprise's *existing* data stack, rather than requiring a new platform or re-architecting work. This is a major practical advantage for enterprise adoption. - **Security-First Architecture**: By never storing raw data and operating only on metadata, the platform solves the core security and compliance objection that kills many AI-for-data plays in larger enterprises. ## Strategic Focus - **Scaling the Engineering Team**: Actively hiring for key technical roles (Research Engineer, Frontend Engineer, Interns) to build out the core platform and AI capabilities [builtin.com](https://builtin.com/company/tensorstax/jobs). - **Accelerating Product Development**: The $5.3M seed round is explicitly earmarked for this purpose [globenewswire.com](https://www.globenewswire.com/news-release/2025/05/12/3078960/0/en/tensorstax-raises-5m-to-build-deterministic-ai-agents-for-data-engineers.html). - **Enterprise Adoption**: Focused on demonstrating value to early adopters and building a repeatable sales motion for enterprise data teams. ## Why Work Here - **High-Impact, Early-Stage Environment**: With only 4 employees, any new hire (especially in engineering/research) will have an outsized impact on the product, culture, and technical direction. - **Cutting-Edge Problem**: The work sits at the intersection of AI, systems engineering, and data infrastructure, offering significant learning opportunities and challenging technical problems. - **In-Office Culture**: All posted jobs are located in San Francisco, CA (HQ), suggesting a preference for in-person collaboration. - **Strong Backing**: Backed by top-tier VC Glasswing Ventures, providing financial runway and strategic support for the next phase of growth. - **Opportunity for Growth**: Joining a team with huge LinkedIn following growth and a hot product idea means potential for rapid career advancement as the company scales. ## Sources 1. [tensorstax.com](https://www.tensorstax.com/?tpcc=NL_Marketing) 2. [builtin.com](https://builtin.com/company/tensorstax/jobs) 3. [linkedin.com](https://linkedin.com/company/tensorstax) 4. [globenewswire.com](https://www.globenewswire.com/news-release/2025/05/12/3078960/0/en/tensorstax-raises-5m-to-build-deterministic-ai-agents-for-data-engineers.html) ## Other roles at tensorstax.com - [Frontend Engineer Internship](https://feeny.ai/job/frontend-engineer-internship-tensorstax-com-san-francisco-p2fsjzkacnkr) — San Francisco, CA - [Backend Engineer, Data & Agent Platform](https://feeny.ai/job/backend-engineer-data-agent-platform-tensorstax-com-san-francisco-c5k22xqjdww2) — San Francisco, CA - [Research Engineer, Reinforcement Learning](https://feeny.ai/job/research-engineer-reinforcement-learning-tensorstax-com-san-francisco-gys442sy11x4) — San Francisco, CA