--- title: 'RL Environments Engineer at Bespoke Labs' canonical: 'https://feeny.ai/job/rl-environments-engineer-bespoke-labs-mountain-view-r1brn5r4z8ex' type: 'job' last_seen: '2026-09-10' --- # RL Environments Engineer at Bespoke Labs - **Company:** Bespoke Labs - **Location:** Mountain View, CA - **Employment:** full-time - **Work type:** onsite - **Posted:** 2026-08-27 - **Last confirmed live:** 2026-09-10 - **Apply:** https://jobs.ashbyhq.com/bespokelabs/f8db1def-3295-4209-bcb2-2f42ed16ac6c ## Job description ## About Bespoke Labs Bespoke Labs is an applied AI research lab pioneering data and RL environment curation for training and evaluating agents. Recently, we curated [Open Thoughts](https://open-thoughts.ai/), one of the best open reasoning datasets used by multiple frontier labs, trained SOTA specialized models such as [Bespoke-MiniChart-7B](https://www.bespokelabs.ai/blog/bespoke-minichart-7b) and [Bespoke-MiniCheck](https://www.bespokelabs.ai/bespoke-minicheck), and [taught](https://www.bespokelabs.ai/blog/improving-multi-turn-tool-use-with-reinforcement-learning) agents to do multi-turn tool-calling with reinforcement learning. Bespoke is uniquely positioned to capture a large market share of data and RL environment curation. ## About the Role This is a delivery role. We want an engineer who has built the machinery that turns environment ideas into hundreds or thousands of validated agentic coding tasks, and who can do it here, fast. You will not be studying environments in the abstract. You will build the pipelines that mass-produce them, design the complex coding worlds agents train inside, and keep pushing throughput: more environments, higher quality, less manual work per task. We will measure you on the volume and quality of environments you ship, not on papers. The thing we care about most is whether you have done this before. If you have stood up an environment-generation pipeline, scaled agentic task creation into the hundreds or thousands, and shipped it, we want to talk. ## What You'll Do - Build environment-generation pipelines. Own the systems that produce RL environments programmatically, including templating, automated grading, verification, and QA, so the team ships environments at scale instead of one at a time. - Create complex coding worlds. Build high-fidelity environments around real codebases, with the conventions, dependencies, tooling, and technical debt that real software actually has. - Scale agentic task creation to thousands. Take task generation from handfuls to hundreds and thousands of validated agentic coding tasks, with automation doing the heavy lifting. - Build tools that raise throughput. Find the bottlenecks in environment production and remove them. Build the internal tooling and infrastructure that makes everyone on the team faster. - Own the full task lifecycle. Prompt, environment, grader, running frontier models against the task, failure analysis, and iteration, until each task is rigorous, fair, and hard to game. - Defend quality at scale. Catch reward hacking and grader loopholes, and build the verification and standards that hold the bar as volume grows. - Direct coding agents heavily. Use frontier coding agents to build and validate environments faster, judging their output and catching the subtle failures. Direct frontier coding agents heavily to build and validate environments, judging their output and catching the quiet failures they produce. ## What We're Looking For A record of shipped volume. You have built agentic coding tasks or environments and can show us how many you personally drove and what they cost to produce. Experience scaling that output through automation rather than through more people doing more manual work. Strong software engineering fundamentals and fluency in several languages that holds up in production code. Real experience with production software. Large codebases, build systems, testing, deployment, on-call, and root cause analysis. You know what real engineering work feels like because you have done it. An adversarial mindset. You look at a grader and ask how a model would cheat it, and then you fix that. A clear sense of what frontier coding agents can and cannot do, and where they cut corners. Ownership. You build, debug, and ship without much supervision. You May Be a Good Fit If You Also - Have worked on RL training systems, post-training, verifiers, or tool-use harnesses - Come from developer tooling, CI/CD, sandboxes, or code execution infrastructure - Have built large-scale automated test generation, fuzzing harnesses, or benchmark suites, which is close cousin work even if it was never called an RL environment - Have contributed to a public agentic benchmark such as Terminal-Bench - Have open-source work that other people depend on ## What We Offer - Location: Mountain View, CA (Onsite) - Base Salary: $250,000 – $300,000 USD / year - Additional Comp: 25% performance-based bonus + equity ## Benefits & Perks - Health, dental, and vision coverage - 401(k) - Daily onsite lunch provided - Visa sponsorship and relocation support available - Direct impact on how the industry trains and evaluates agents We value different backgrounds and paths into this work. If this role excites you but you do not check every box, apply anyway. ## About Bespoke Labs ## Company Overview - **One-liner**: Bespoke Labs is an applied AI research lab building company-scale RL environments and data curation infrastructure to train reliable, production-grade AI agents. - **Entity Type**: Private (Seed stage) - **Headquarters**: Mountain View, California, United States (with an office in Santa Clara, CA) - **Founded**: 2024 - **Founders**: Mahesh Sathiamoorthy (CEO, formerly Google DeepMind) and Alex Dimakis (CSO, Professor at UC Berkeley) ## Core Business - **Primary industry**: Applied AI Research / Software Development (Agent Infrastructure) - **Target customers**: B2B — Frontier AI labs (e.g., Anthropic, OpenAI, Google DeepMind) and enterprises needing to train, evaluate, and optimize complex, long-horizon AI agents for production. - **Mission**: To build the environment infrastructure for the agent revolution, making AI agents dependable in production by creating entire digital worlds for training. ## Products & Services - **Company-Scale RL Environments**: Infrastructure that simulates real codebases and microservices, allowing agents to master complex, long-horizon workflows required for production. - **Agent Evaluation & Optimization (GEPA)**: An evolutionary algorithm (GEPA optimizer) that automates prompt and policy searches, achieving superior accuracy faster than manual prompt engineering. Over 200 teams use GEPA in production. - **Production-Grade RL & Benchmarking**: Collaborative research in RL, data curation, and benchmarking to ensure training environments keep pace with the advancing frontier of model capabilities. - **OpenThoughts Dataset**: One of the best open reasoning datasets (10k+ monthly downloads on Hugging Face), cited in ICLR 2026. - **Terminal-Bench**: The first environment-based benchmark for agentic systems, cited by Anthropic, OpenAI, and Google DeepMind, published at ICLR 2026. - **Bespoke-MiniCheck & Bespoke-MiniChart**: SOTA models developed in-house for fact-checking and chart understanding. ## Market Standing - **Valuation**: Not publicly disclosed. - **Total Funding**: $8.25M (Seed Round, June 2024). - **Notable Investors/Partners**: Lead investor is **8VC**. Advisors include Tasso Argyros (VP of Eng at Databricks, ex-CEO of ActionIQ), Joseph Gonzalez (Associate Professor at UC Berkeley, creator of LMSYS/vLLM), and Greg Durrett (Associate Professor at NYU). Cited by Anthropic, OpenAI, and Google DeepMind. - **Growth Signals**: Headcount grew **+238.5% year-over-year** (from ~8 to 28 employees). Monthly headcount growth is **+33.3%**. Active job postings increased **+183.3% quarterly** (34 open positions). Over 200 teams use GEPA in production. OpenThoughts sees 10k+ monthly downloads. ## Competitive Advantages - **First-Mover in Agent Environments**: They identified the bottleneck for agent reliability is the environment, not the model, and are building the foundational infrastructure for this new paradigm. - **SOTA Research Output**: Published work at ICLR 2026 (Terminal-Bench, GEPA, OpenThoughts) that is being cited and used by frontier labs (Anthropic, OpenAI, Google DeepMind). - **Strong Academic & Industry Ties**: Founded by a Google DeepMind alum and a UC Berkeley professor, with advisors from Databricks, UC Berkeley, and NYU. Team includes alumni from Google, Scale AI, Microsoft, and AI2. - **Open Source Influence**: OpenThoughts is a widely used open reasoning dataset, building community credibility and attracting top talent. ## Strategic Focus - **Scale the Environment Infrastructure**: Build increasingly complex and realistic digital worlds for training agents, moving beyond demos to production-grade reliability. - **Expand Enterprise Adoption**: Help enterprises evaluate and optimize agents in environments that mirror their specific systems and processes. - **Continue Frontier Research**: Push the boundaries of RL, data curation, and benchmarking to maintain a lead in the agent training space. - **Talent Acquisition**: Aggressively hiring researchers and engineers (34 open roles) to scale the team and the product. ## Why Work Here - **High-Impact, Cutting-Edge AI Work**: Opportunity to work on one of the most important problems in AI — making agents reliable. The team is described as "exceptional researchers and engineers." - **Strong Research Culture**: The lab publishes at top venues (ICLR) and contributes to open source (OpenThoughts). The work is a blend of product engineering and fundamental research. - **Rapid Growth Stage**: With 238% YoY headcount growth and a recent seed round, this is an early-stage opportunity to shape the company's culture and technical direction. - **Office-First, Collaborative Environment**: Based in Mountain View/Santa Clara with an emphasis on in-person collaboration. The team is small (28 people) but growing fast. - **Notable Alumni & Hiring Sources**: The team attracts talent from Scale AI, Berkeley RISE Lab, Google DeepMind, and Grammarly, indicating a high-caliber peer group. - **Open Roles**: Actively hiring for GPU/CUDA Engineers, Product Engineers, Scientific ML Engineers (PhD Interns), and Quantitative Financial Specialists, among others. ## Sources 1. [bespokelabs.ai](https://www.bespokelabs.ai/) 2. [bespokelabs.ai/about-us](https://www.bespokelabs.ai/about-us) 3. [linkedin.com/company/bespokelabsai](https://www.linkedin.com/company/bespokelabsai) 4. [pitchbook.com/profiles/company/651710-80](https://pitchbook.com/profiles/company/651710-80) 5. [jobs.ashbyhq.com/bespokelabs](https://jobs.ashbyhq.com/bespokelabs) ## Other roles at Bespoke Labs - [Desktop/System Administrator](https://feeny.ai/job/desktop-system-administrator-bespoke-labs-bengaluru-f4e04cka1jkn) — Bengaluru, India - [Backend Engineer](https://feeny.ai/job/backend-engineer-bespoke-labs-mountain-view-2wfkmwxgk7se) — Mountain View, CA - [Full Stack Engineer - MTV](https://feeny.ai/job/full-stack-engineer-mtv-bespoke-labs-mountain-view-mv3kv8nrts8f) — Mountain View, CA - [Full Stack Engineer - BLR, India](https://feeny.ai/job/full-stack-engineer-blr-india-bespoke-labs-bengaluru-pazpmr8p8dp8) — Bengaluru, India - [Engagement Manager](https://feeny.ai/job/engagement-manager-bespoke-labs-mountain-view-4aevsp00m946) — Mountain View, CA - [Founding Recruiter](https://feeny.ai/job/founding-recruiter-bespoke-labs-remote-hnw3zbanavm6) - [Research Engineer](https://feeny.ai/job/research-engineer-bespoke-labs-mountain-view-vv4mp69f8z1b) — Mountain View, CA - [Product Engineer (Bangalore)](https://feeny.ai/job/product-engineer-bangalore-bespoke-labs-bengaluru-zg7eead9kgg5) — Bengaluru, India - [Design and Brand Storyteller](https://feeny.ai/job/design-and-brand-storyteller-bespoke-labs-mountain-view-asg657e8gnwr) — Mountain View, CA - [AI Enterprise Engineer](https://feeny.ai/job/ai-enterprise-engineer-bespoke-labs-mountain-view-cvams3d3p42m) — Mountain View, CA