--- title: 'Applied AI Engineer at Judgment Labs' canonical: 'https://feeny.ai/job/applied-ai-engineer-judgment-labs-san-francisco-2c68arqvx3sp' type: 'job' last_seen: '2026-09-10' --- # Applied AI Engineer at Judgment Labs - **Company:** Judgment Labs - **Location:** San Francisco, CA - **Employment:** full-time - **Work type:** onsite - **Posted:** 2026-01-11 - **Last confirmed live:** 2026-09-10 - **Apply:** https://jobs.ashbyhq.com/judgmentlabs/26ab8af7-8b33-43aa-9063-1ff782b36beb ## Job description The Role: We are looking for Applied AI Engineers to build AI systems that use agent interaction data to understand how agents behave, evaluate them at scale, and improve them through learning and feedback. Your research will not live on a whiteboard. You’ll work directly with real-world agent data, apply frontier methods in production, and see your work ship into the product. By making agent behavior measurable and debuggable, your systems will support teams deploying agents across finance, legal, operations, and other high-stakes workflows. You will own projects end-to-end, with significant autonomy, and work closely with the team to build self-improving agent systems. What You'll Do: - Build AI systems to aggregate, index, and analyze large-scale long-running agent interaction data in order to extract meaningful signals - Design and implement post-training and optimization workflows to improve agents, both internally and for customers - Build agent platform infrastructure, including orchestration, runtimes, and developer tools that help teams define, test, deploy, and iterate on complex agent workflows - Build internal tools and infrastructure that support rapid experimentation, analysis, and training - Work closely with product to integrate agents into customer-facing workflows - Collaborate with external companies and research partners on frontier AI research ## What We're Looking For Every hire clears three bars, no exceptions: - Agency. You are intellectually curious, self-directed, and stay up to date with the latest research, blogs, trends, and ideas. - Depth of thought. You can reason clearly about abstract systems, and ideally have experience working on agents, RL, or the infrastructure that supports them. - Ownership. You own outcomes, not just tasks. You use freedom to experiment responsibly, make business-driven decisions, and focus first on work that moves the company forward. More specifically, you should bring strength in at least one of the following areas: - Data quality, evaluation, benchmarking, and hands-on work with messy production data - Agent systems built or evaluated in real-world or production settings - Reinforcement learning, post-training, agents, or machine learning fundamentals - Infrastructure and systems work across training, data pipelines, evaluation, or model serving - Translating research into product while balancing customer constraints, technical tradeoffs, and business impact - Turning ambiguous problems into clear, well-designed plans ## About Judgment Labs ## Company Overview - **One-liner**: Judgment Labs builds a continuous-improvement stack for AI agents, helping teams monitor, evaluate, and improve agent behavior in production. - **Entity Type**: Private (Seed + Series A funded) - **Headquarters**: San Francisco, California, United States - **Founded**: 2025 - **Founders**: Alex Shan (CEO), Andrew Li (Chief Scientist), Joseph Camyre (CTO) ## Core Business - **Primary industry**: AI Infrastructure / Agent Observability - **Target customers**: B2B — AI-native companies and engineering teams building and deploying autonomous AI agents - **Mission or purpose**: Provide the infrastructure for a future where billions of AI agents act autonomously, giving teams the tools to make their products better with every interaction. ## Products & Services - **Behavior Discovery**: Automatically constructs and refines evaluation rubrics from verifiable signals, surfacing failure modes and usage patterns from unlabeled production trajectories. - **AutoRubrics**: Automates the creation of evaluation criteria for agent behavior. - **Agent Search**: Enables querying across agent trajectories at a behavioral level, beyond simple input/output keyword search. - **Agent Judge**: Builds cheaper and more accurate trajectory-level evaluators using harnesses. - **Slack & MCP Integrations**: Allows teams to investigate issues, run tests, and take action directly from Slack, Claude, Codex, Cursor, or any MCP client. ## Market Standing - **Valuation/Market Cap**: Not disclosed - **Key Metric**: $32M in total funding (Seed + Series A, announced May 2026) - **Notable Investors/Partners**: Lightspeed Venture Partners (led both rounds), Nova Global, SV Angel, Valor Equity Partners, Dynamic. Key partner: James Alcorn (Partner at Lightspeed) sits on the board. - **Growth Signals**: +400% YoY headcount growth (19 employees as of mid-2026); platform already in production at a growing list of agent-native companies; company describes customer traction as “extraordinary.” ## Competitive Advantages - First-mover advantage in a new category: agent behavior monitoring (ABM) infrastructure, built specifically for deep, multi-step agent trajectories rather than simple chatbot input/output evals. - Founders have deep domain expertise: CEO Alex Shan was an AI researcher at Stanford’s NLP group under Chris Manning; Chief Scientist Andrew Li was an early research hire at TogetherAI; CTO Joseph Camyre built large-scale infrastructure at Datadog. - Product is already standardizing within agent-native startups, creating a network effect as more production data flows through the platform. ## Strategic Focus - Aggressively hiring AI researchers and engineers in San Francisco. - Expanding the forward-deployed engineering team to serve a growing customer base. - Mission: “Give every team building agents the tools to make their products better with every interaction.” ## Why Work Here - **Culture**: Described as fast-moving, high-agency, and intellectually ambitious — values critical thinking, questioning assumptions, and taking bold bets. - **Work Policy**: On-site in San Francisco (all roles require physical presence at HQ). - **Notable Perks**: Work at the frontier of AI agent infrastructure; join a small but rapidly growing team (19 people) backed by top-tier VCs; early employees will define the company’s trajectory and culture. - **Engineering Culture**: Emphasis on research engineering, rapid prototyping, and productionizing cutting-edge AI evaluation methods. ## Sources 1. [judgmentlabs.ai](https://judgmentlabs.ai/) 2. [judgmentlabs.ai/careers](https://www.judgmentlabs.ai/careers) 3. [linkedin.com/company/judgmentlabs](https://www.linkedin.com/company/judgmentlabs) 4. [builtin.com](https://builtin.com/company/judgment-labs) 5. [businesswire.com](https://www.businesswire.com/news/home/20260512621556/en/Judgment-Labs-Closes-%2432M-in-Seed-and-Series-A-Funding-to-Build-the-Continuous-Improvement-Layer-for-AI-Agents) ## Other roles at Judgment Labs - [Product engineer, Agent](https://feeny.ai/job/product-engineer-agent-judgment-labs-san-francisco-t4ryveqq2yns) — San Francisco, CA - [Founding Account Executive](https://feeny.ai/job/founding-account-executive-judgment-labs-san-francisco-973j3geb8a8e) — San Francisco, CA - [Product engineer, full stack](https://feeny.ai/job/product-engineer-full-stack-judgment-labs-san-francisco-rzn6c0km6mxc) — San Francisco, CA - [Backend/Infra Engineer](https://feeny.ai/job/backend-infra-engineer-judgment-labs-san-francisco-dpt1kesgk6d9) — San Francisco, CA - [Applied AI Engineer](https://feeny.ai/job/applied-ai-engineer-arena-intelligence-inc-bay-area-zm7z9vzxbxb6) — Bay Area - [Applied AI Engineer](https://feeny.ai/job/applied-ai-engineer-maintainx-san-francisco-qdq8pd728ahj) — San Francisco, CA - [Applied AI Engineer](https://feeny.ai/job/applied-ai-engineer-redpine-stockholm-vaahy9qp7cfc) — Stockholm, Sweden - [Applied AI Engineer](https://feeny.ai/job/applied-ai-engineer-moss-warsaw-ghb9v3ctc8y6) — Warsaw, Poland - [Applied AI Engineer](https://feeny.ai/job/applied-ai-engineer-monte-carlo-americas-385312hg9b1q) — Americas - [Applied AI Engineer](https://feeny.ai/job/applied-ai-engineer-block-labs-portugal-bkr1smvtqxqa) — Portugal