--- title: 'Product engineer, Agent at Judgment Labs' canonical: 'https://feeny.ai/job/product-engineer-agent-judgment-labs-san-francisco-t4ryveqq2yns' type: 'job' last_seen: '2026-09-10' --- # Product engineer, Agent at Judgment Labs - **Company:** Judgment Labs - **Location:** San Francisco, CA - **Employment:** full-time - **Work type:** onsite - **Posted:** 2026-08-13 - **Last confirmed live:** 2026-09-10 - **Apply:** https://jobs.ashbyhq.com/judgmentlabs/dbd14605-7d13-4a56-8c85-119b09bb0ba3 ## Job description Product Engineer - Agents Job Description ## The Role Judgment is the learning infrastructure for AI agents. Agents in production don't improve from prompts alone. They improve from experience: the tasks they attempt, the mistakes they make, the edge cases they hit. Here's how it works: - We ingest everything your agents do in production: traces, tool calls, decisions, outcomes - Judgment turns that raw experience into structured signals: failure modes, behaviors, rubrics, evals - Teams close the loop, shipping agent improvements validated against real production evidence You'll build the product experiences that make this loop legible, and you'll build the agents that run it. This is not a role where you implement specs handed down. You'll own problems end-to-end: talking to customers, defining what to build, building it, and iterating until it's great. ## What You Will Accomplish - Judgment Agent: Shape how the Judgment Agent runs large-scale investigations: parallel investigators working across thousands of production traces, each covering a different dimension (failure modes, tool errors, regressions, drift), merging results into one answer. - Verification: Build the platform for verifying agent changes: hosted simulated environments for stateful agent evals, trajectory replay against changed agents, and monitors for unintended behavior changes. - Agent investigation interfaces: Design how engineers understand what their agents did and why. Long traces, tool calls, decisions, failures. What does debugging look like when the "program" is a reasoning loop? How do you make a thousand-step trajectory legible in minutes? - Swarm UX: A hundred parallel investigations is useless if engineers can't follow them. Design how humans watch a swarm work, redirect investigators that go down the wrong path, and consume findings without reading a hundred reports. - The improvement loop: Build the workflows that turn production trajectories into datasets, judges, and regression checks, so the path from "found a problem" to "verified a fix" feels like one motion. - The platform underneath: Workspaces, roles, permissions, billing, usage, and limits for teams running many agents across many environments. - Judgment everywhere agents are built: An SDK and terminal-first experience so Claude Code, Codex, and OpenCode sessions can summon Judgment as a subagent mid-development. ## What You'll Bring - Experience building and scaling end-to-end production systems, from data layer to UI - Strong technical problem-solving skills, especially in fast-changing, ambiguous environments - A builder and tinkerer's mindset with high agency - you find creative ways to overcome obstacles and ship - Hands-on experience building with LLMs or agents, or the drive to get there fast - Comfort working directly with customers to understand their needs and solve real-world problems - Excellent communication skills - clear, direct, and persuasive across technical and non-technical audiences ## About Judgment Labs ## Company Overview - **One-liner**: Judgment Labs builds a continuous-improvement stack for AI agents, helping teams monitor, evaluate, and improve agent behavior in production. - **Entity Type**: Private (Seed + Series A funded) - **Headquarters**: San Francisco, California, United States - **Founded**: 2025 - **Founders**: Alex Shan (CEO), Andrew Li (Chief Scientist), Joseph Camyre (CTO) ## Core Business - **Primary industry**: AI Infrastructure / Agent Observability - **Target customers**: B2B — AI-native companies and engineering teams building and deploying autonomous AI agents - **Mission or purpose**: Provide the infrastructure for a future where billions of AI agents act autonomously, giving teams the tools to make their products better with every interaction. ## Products & Services - **Behavior Discovery**: Automatically constructs and refines evaluation rubrics from verifiable signals, surfacing failure modes and usage patterns from unlabeled production trajectories. - **AutoRubrics**: Automates the creation of evaluation criteria for agent behavior. - **Agent Search**: Enables querying across agent trajectories at a behavioral level, beyond simple input/output keyword search. - **Agent Judge**: Builds cheaper and more accurate trajectory-level evaluators using harnesses. - **Slack & MCP Integrations**: Allows teams to investigate issues, run tests, and take action directly from Slack, Claude, Codex, Cursor, or any MCP client. ## Market Standing - **Valuation/Market Cap**: Not disclosed - **Key Metric**: $32M in total funding (Seed + Series A, announced May 2026) - **Notable Investors/Partners**: Lightspeed Venture Partners (led both rounds), Nova Global, SV Angel, Valor Equity Partners, Dynamic. Key partner: James Alcorn (Partner at Lightspeed) sits on the board. - **Growth Signals**: +400% YoY headcount growth (19 employees as of mid-2026); platform already in production at a growing list of agent-native companies; company describes customer traction as “extraordinary.” ## Competitive Advantages - First-mover advantage in a new category: agent behavior monitoring (ABM) infrastructure, built specifically for deep, multi-step agent trajectories rather than simple chatbot input/output evals. - Founders have deep domain expertise: CEO Alex Shan was an AI researcher at Stanford’s NLP group under Chris Manning; Chief Scientist Andrew Li was an early research hire at TogetherAI; CTO Joseph Camyre built large-scale infrastructure at Datadog. - Product is already standardizing within agent-native startups, creating a network effect as more production data flows through the platform. ## Strategic Focus - Aggressively hiring AI researchers and engineers in San Francisco. - Expanding the forward-deployed engineering team to serve a growing customer base. - Mission: “Give every team building agents the tools to make their products better with every interaction.” ## Why Work Here - **Culture**: Described as fast-moving, high-agency, and intellectually ambitious — values critical thinking, questioning assumptions, and taking bold bets. - **Work Policy**: On-site in San Francisco (all roles require physical presence at HQ). - **Notable Perks**: Work at the frontier of AI agent infrastructure; join a small but rapidly growing team (19 people) backed by top-tier VCs; early employees will define the company’s trajectory and culture. - **Engineering Culture**: Emphasis on research engineering, rapid prototyping, and productionizing cutting-edge AI evaluation methods. ## Sources 1. [judgmentlabs.ai](https://judgmentlabs.ai/) 2. [judgmentlabs.ai/careers](https://www.judgmentlabs.ai/careers) 3. [linkedin.com/company/judgmentlabs](https://www.linkedin.com/company/judgmentlabs) 4. [builtin.com](https://builtin.com/company/judgment-labs) 5. [businesswire.com](https://www.businesswire.com/news/home/20260512621556/en/Judgment-Labs-Closes-%2432M-in-Seed-and-Series-A-Funding-to-Build-the-Continuous-Improvement-Layer-for-AI-Agents) ## Other roles at Judgment Labs - [Founding Account Executive](https://feeny.ai/job/founding-account-executive-judgment-labs-san-francisco-973j3geb8a8e) — San Francisco, CA - [Product engineer, full stack](https://feeny.ai/job/product-engineer-full-stack-judgment-labs-san-francisco-rzn6c0km6mxc) — San Francisco, CA - [Backend/Infra Engineer](https://feeny.ai/job/backend-infra-engineer-judgment-labs-san-francisco-dpt1kesgk6d9) — San Francisco, CA - [Applied AI Engineer](https://feeny.ai/job/applied-ai-engineer-judgment-labs-san-francisco-2c68arqvx3sp) — San Francisco, CA