--- title: 'Member of Technical Staff (Evals & Post-Training) at Ambral' canonical: 'https://feeny.ai/job/member-of-technical-staff-evals-post-training-ambral-new-york-eyrabnspqwa3' type: 'job' last_seen: '2026-09-13' --- # Member of Technical Staff (Evals & Post-Training) at Ambral - **Company:** Ambral - **Location:** New York, NY - **Compensation:** $165k–$325k - **Employment:** full-time - **Work type:** onsite - **Posted:** 2026-09-01 - **Last confirmed live:** 2026-09-13 - **Apply:** https://jobs.ashbyhq.com/ambral/56dfa5cb-f18a-474f-88aa-adea46605c38 ## Job description ## What we do Ambral Labs helps enterprises own the intelligence behind their most important workflows. Every company has years of historical evidence showing how work gets done: the context people had, the decisions they made, the actions they took, and the outcomes that followed. Today, most of that history is inert. It isn’t structured in a way that companies can use to evaluate models and improve agent behavior. Ambral turns this history into replayable environments and eval sets grounded in real workflows and observed outcomes. We use those environments to improve model performance through reinforcement learning and other post-training techniques, alongside context engineering, harness design, and agent engineering. The result is better, more cost-efficient AI for each enterprise’s specific work, powered by open-weight models that the company owns and controls. This allows each company to retain ownership of its core intelligence instead of outsourcing it to a model provider. We graduated from YC S2025, raised millions in funding, and are already deployed within multi-billion dollar enterprises. Now we're growing the founding team ## What you’ll do At the center of Ambral Labs is a replayable environment engine for the enterpirse. The system reconstructs a company’s world as it existed at a particular moment in the past then exposes that state through the same tools an agent would use in production. This lets us place new policies and agent configurations inside real historical environments, observe how they reason and act, and grade their performance against real outcomes. You'll work across research, infrastructure, and production systems including: - Building an environment factory that converts recorded enterprise data and task definitions into runnable environments - Designing graders that turn ambiguous business objectives into verifiable rewards - Developing methods for mining useful tasks, trajectories, and evaluation cases from historical workflows - Creating eval sets that are representative, reproducible, and resistant to overfitting - Finding the right combinations of models, tools, context, and policies to maximize performance while reducing inference cost - Training and evaluating agents that operate over long horizons, incomplete information, and large tool spaces - Building replay and observability systems that make agent behavior explainable and measurable - Scaling from individual environments to thousands of concurrent training and evaluation runs These problems are wide open. You’ll have significant ownership over both the research direction and the production systems that make it real. You’ll work directly with the CTO, deploy into real enterprise workflows, and see your research tested against consequential problems and observable outcomes. ## Who you are - You have 1-7 years of experience building production software or machine-learning systems (we're hiring at multiple levels for this role). - Bonus points for working on reinforcement-learning environments, LLM post-training, evaluation infrastructure, agent harnesses, or closely related systems - You understand how environment design, reward design, context, tooling, and policy behavior interact - You’re comfortable turning fuzzy business objectives into tasks and signals that can be evaluated reliably - You can diagnose whether a model’s limitations come from the model itself, its context, its tools, its harness, or its training - You can move between research questions and production implementation without treating them as separate jobs - You write strong software and can build systems that process large, messy datasets at scale - You care about reproducibility, observability, and understanding why a model behaves the way it does - You’re looking to do the best work of your life and build something you’ll be proud of for decades We care much more about what you’ve built and how you think than credentials or conventional career paths. ## Benefits - Significant equity and ownership - Equinox membership - Free meals, coffee, and snacks - Health insurance - Unlimited PTO ## About Ambral ## Company Overview - **One-liner**: Ambral builds AI agents that unify customer data across enterprise systems to autonomously manage account expansions, prevent churn, and surface revenue opportunities. - **Entity Type**: Private (Startup – Y Combinator S25, raised millions in seed funding) - **Headquarters**: New York City, NY, USA - **Founded**: 2025 - **Founders**: Sam Brickman (CEO) and Jack Stettner (CTO) ## Core Business - **Primary industry**: AI-powered enterprise account management and customer success - **Target customers**: B2B enterprises and scale-ups with complex, multi-system customer data - **Mission**: Make deep customer intimacy scalable by giving every account an AI account manager that continuously learns and acts ## Products & Services - **[Ambral Platform]**: Integrates with CRMs, data warehouses, support tickets, meeting transcripts, and 20+ internal systems to create living digital twins of each customer account. Uses AI to simulate decisions, surface risks/opportunities, and autonomously execute actions (e.g., expansion calls, renewal follow-ups). - **[Ambral Cortex]**: Enterprise-grade AI Account Manager that handles conversational onboarding, manual reporting, driving expansions, and securing renewals – designed to serve the "long tail" of accounts that are typically underserved. ## Market Standing - **Valuation**: Not disclosed (private startup) - **Key Metric**: Total funding – "millions" raised (exact amount not public); Y Combinator S25 batch graduate - **Notable Investors/Partners**: Y Combinator (Harj Taggar as primary partner); founders have previous exits and experience at SpaceX, Everlywell, Wonder ($7B valuation) - **Growth Signals**: Already deployed within multi-billion dollar enterprises; attributable expansion revenue in the hundreds of millions; team of 3 with active hiring for Senior Founding Engineer ## Competitive Advantages - **Technical moat**: Founders have deep experience in mission-critical systems (SpaceX flight software) and scaling AI at healthcare unicorns; platform back-tests models against actual risk and growth drivers for each customer - **Data unification**: Resolves identity across 20+ sources into a single source of truth, enabling autonomous action without manual data prep - **Simulation capability**: "Ambral Simulate" lets companies run hypothetical scenarios before making structural decisions (pricing, product launches, GTM pivots) ## Strategic Focus - **Immediate priority**: Scale the platform from initial enterprise deployments to broader market, hire founding engineering team, and expand autonomous capabilities for account management - **Long-term vision**: Become the coordination layer for enterprise operations – where AI agents synthesize data, execute routine work, and escalate critical signals to humans ## Why Work Here - **Culture**: Early-stage startup with a small, high-caliber team (3 people); philosophy of "elite groups of operators acting on critical signals" – you'll have outsized impact - **Work environment**: Hybrid (employees engage in a mix of remote and on-site work); office in New York City - **Notable perks**: Competitive salary range for founding engineer ($185K–$245K); equity (2%–4%); opportunity to shape the product and architecture from the ground up - **Engineering culture**: Building at the intersection of AI agents and enterprise data – similar to "sitting in mission control" – with a focus on reliability and real-time decision-making ## Sources 1. [ambral.com](https://www.ambral.com/) 2. [ycombinator.com/companies/ambral](https://www.ycombinator.com/companies/ambral) 3. [ycombinator.com/companies/ambral/jobs](https://www.ycombinator.com/companies/ambral/jobs) 4. [builtin.com/company/ambral](https://builtin.com/company/ambral) 5. [jobs.ashbyhq.com/Ambral](https://jobs.ashbyhq.com/Ambral) ## Other roles at Ambral - [Head of Research](https://feeny.ai/job/head-of-research-ambral-new-york-8dcndf068dr2) — New York, NY - [Member of Technical Staff (New Grad)](https://feeny.ai/job/member-of-technical-staff-new-grad-ambral-new-york-9rawgxyzrt6m) — New York, NY