--- title: 'Backend/Infra Engineer at Judgment Labs' canonical: 'https://feeny.ai/job/backend-infra-engineer-judgment-labs-san-francisco-dpt1kesgk6d9' type: 'job' last_seen: '2026-09-10' --- # Backend/Infra Engineer at Judgment Labs - **Company:** Judgment Labs - **Location:** San Francisco, CA - **Employment:** full-time - **Work type:** onsite - **Posted:** 2026-06-15 - **Last confirmed live:** 2026-09-10 - **Apply:** https://jobs.ashbyhq.com/judgmentlabs/e11f2965-ef1a-4fe9-8bed-76fe3827350a ## Job description Senior Backend Engineer San Francisco · On Site · Full Time Judgment Labs is building the infrastructure for continual learning in long-horizon AI agents. The next generation of agents will not improve from prompts alone. They will improve from experience: the tasks they attempt, the tools they use, the mistakes they make, the edge cases they encounter, and the outcomes they produce in production. The hard part is turning that raw experience into high-quality data that can actually improve the system. Judgment builds the infrastructure to do that. We turn long agent trajectories into clean, structured data for evals, labeling, rubric generation, context engineering, and RL workflows. Instead of only showing teams what happened, Judgment helps decide what matters, what should be learned from, and how that learning should flow back into the agent. Databricks built the data infrastructure for analytics. Judgment is building the learning infrastructure for agents. We’ve raised $30M+ from Lightspeed, SV Angel, Valor Equity Partners, and others. ## The Role We’re looking for a Senior Backend Engineer to own the systems that ingest, structure, evaluate, and serve agent experience data at production scale. This role includes the backend and data infrastructure surface area: high-throughput telemetry ingestion, ClickHouse-backed OLAP performance, evaluation pipelines, RabbitMQ/Temporal workflows, multi-tenant scheduling, and product-facing APIs. Some weeks you’ll be deep in distributed systems and query performance. Other weeks you’ll ship a customer-facing feature end to end across backend, frontend, and the data layer. This is not a narrow API role. The backend is where raw agent trajectories become structured learning data. Interesting Technical Challenges - High-throughput telemetry ingestion. Parse and persist OTEL traces across protobuf and JSON formats at hundreds of thousands of spans per second, writing to ClickHouse while keeping ingest latency low and backpressure graceful as customer traffic spikes. - Petabyte-scale OLAP performance. Design schemas, partitioning, indexes, storage layouts, and query paths so behavioral queries over billions of spans stay fast. Turn real access-pattern analysis into concrete data modeling decisions. - Long-horizon trajectory modeling. Agent workflows are messy: multi-step tasks, tool calls, retries, partial failures, context changes, and unclear outcomes. Build the abstractions that turn those trajectories into structured data for evals, labeling, rubric generation, context engineering, and RL workflows. - Queue- and workflow-driven evaluation. Evaluations fan out across RabbitMQ and Temporal workflows. Getting this right means reasoning about retries, timeouts, idempotent state transitions, exactly-once-ish semantics, and reconciling runs that fail partway so nothing is silently orphaned. - Multi-tenant fairness at scale. A single large customer should not be able to starve everyone else. Build scheduling and execution systems so latency stays predictable across hundreds of teams sharing the same evaluation pipeline. - Near-real-time scoring. Behavioral scorers and agent judges call LLM APIs at scale, so batching, rate-limit management, retry/backoff, failure handling, and cost control are core backend systems problems. - Learning loops for agents. Build the product and systems layer that helps teams decide what matters, what should be learned from, and how that learning flows back into the agent. ## What You’ll Do - Design and build backend systems for trace ingestion, trajectory processing, evaluation orchestration, scoring, labeling, rubric generation, and customer-facing analytics. - Own the API surface used by the Judgment platform UI, SDKs, JudgmentHub libraries, MCP server, Slack agent, and customer integrations. - Build and operate the RabbitMQ / Temporal evaluation pipeline, including retry semantics, failure recovery, state reconciliation, and tenant-level scheduling. - Optimize the ClickHouse OLAP layer: schema design, partitioning, skip indexes, full-text-search pruning, query rewrites, deduplication, pagination correctness, and storage growth. - Turn raw spans, conversations, tool calls, scorer outputs, and agent-judge results into clean data models customers can use for evals, labeling, context engineering, and RL workflows. - Ship features end to end, often across Next.js, backend APIs, queues/workflows, and the data layer. - Work directly with customers to understand where their agents fail, what data is useful, and how Judgment should structure that experience for learning. - Roll out safely with feature flags, design docs, code reviews, tests, observability, and production debugging. - Raise the engineering bar through clear interfaces, maintainable systems, thoughtful reviews, and strong ownership. ## What We’re Looking For - Strong backend engineering experience building and operating production systems under real load. - Excellent fundamentals in distributed systems, API design, data modeling, reliability, and performance. - Experience working with high-volume event, trace, log, metric, or telemetry data. - Strong intuition for data systems: query patterns, storage layout, indexing, partitioning, latency, correctness, and cost. - Comfort owning systems beyond initial launch: debugging production issues, improving observability, scaling bottlenecks, and cleaning up abstractions as the product evolves. - Ability to work across backend, data, product, and infrastructure boundaries rather than treating them as separate silos. - Product judgment and willingness to ship across the stack when needed. - Clear communication. You can write a design doc, review a diff, explain a tradeoff, and unblock others without turning everything into process. ## Nice to Have - Experience with ClickHouse, OLAP systems, distributed query engines, or large-scale analytical databases. - Experience with RabbitMQ, Temporal, Kafka, Spark, Flink, Ray, Airflow, Dagster, Prefect, or similar queue/workflow/data systems. - Experience with OTEL, observability products, tracing, logging, or monitoring infrastructure. - Experience building systems that call LLM APIs at scale, including rate-limit management, retries, batching, and cost control. - Experience with LLM evaluation, labeling systems, rubric generation, context engineering, RL data pipelines, embedding pipelines, vector search, clustering, or anomaly detection. - Experience building developer-facing products, SDK-backed platforms, or customer-facing infrastructure. Why Judgment? - We’re building the learning infrastructure for agents. As agents move from demos to production, the bottleneck is no longer just better prompts. It is turning real production experience into high-quality data for evals, labeling, rubric generation, context engineering, and RL workflows. - The technical problems are foundational. Long agent trajectories are messy, high-volume, and hard to reason about. We’re building the systems that ingest them, structure them, evaluate them, surface what matters, and feed that learning back into the agent. - This is a Databricks-scale infrastructure opportunity. Databricks built the data infrastructure for analytics. Judgment is building the learning infrastructure for agents. - You’ll work on problems customers actually feel. Engineers talk directly to teams building production agents, see where their systems fail, and turn those failures into product and infrastructure. - Small team, high ownership. You will not own a narrow slice. You’ll shape core systems early, ship quickly, and work across product, data, backend, infra, and customer environments. - In person in San Francisco. We work together in person because the problems are hard, the product is moving fast, and the feedback loops matter. ## About Judgment Labs ## Company Overview - **One-liner**: Judgment Labs builds a continuous-improvement stack for AI agents, helping teams monitor, evaluate, and improve agent behavior in production. - **Entity Type**: Private (Seed + Series A funded) - **Headquarters**: San Francisco, California, United States - **Founded**: 2025 - **Founders**: Alex Shan (CEO), Andrew Li (Chief Scientist), Joseph Camyre (CTO) ## Core Business - **Primary industry**: AI Infrastructure / Agent Observability - **Target customers**: B2B — AI-native companies and engineering teams building and deploying autonomous AI agents - **Mission or purpose**: Provide the infrastructure for a future where billions of AI agents act autonomously, giving teams the tools to make their products better with every interaction. ## Products & Services - **Behavior Discovery**: Automatically constructs and refines evaluation rubrics from verifiable signals, surfacing failure modes and usage patterns from unlabeled production trajectories. - **AutoRubrics**: Automates the creation of evaluation criteria for agent behavior. - **Agent Search**: Enables querying across agent trajectories at a behavioral level, beyond simple input/output keyword search. - **Agent Judge**: Builds cheaper and more accurate trajectory-level evaluators using harnesses. - **Slack & MCP Integrations**: Allows teams to investigate issues, run tests, and take action directly from Slack, Claude, Codex, Cursor, or any MCP client. ## Market Standing - **Valuation/Market Cap**: Not disclosed - **Key Metric**: $32M in total funding (Seed + Series A, announced May 2026) - **Notable Investors/Partners**: Lightspeed Venture Partners (led both rounds), Nova Global, SV Angel, Valor Equity Partners, Dynamic. Key partner: James Alcorn (Partner at Lightspeed) sits on the board. - **Growth Signals**: +400% YoY headcount growth (19 employees as of mid-2026); platform already in production at a growing list of agent-native companies; company describes customer traction as “extraordinary.” ## Competitive Advantages - First-mover advantage in a new category: agent behavior monitoring (ABM) infrastructure, built specifically for deep, multi-step agent trajectories rather than simple chatbot input/output evals. - Founders have deep domain expertise: CEO Alex Shan was an AI researcher at Stanford’s NLP group under Chris Manning; Chief Scientist Andrew Li was an early research hire at TogetherAI; CTO Joseph Camyre built large-scale infrastructure at Datadog. - Product is already standardizing within agent-native startups, creating a network effect as more production data flows through the platform. ## Strategic Focus - Aggressively hiring AI researchers and engineers in San Francisco. - Expanding the forward-deployed engineering team to serve a growing customer base. - Mission: “Give every team building agents the tools to make their products better with every interaction.” ## Why Work Here - **Culture**: Described as fast-moving, high-agency, and intellectually ambitious — values critical thinking, questioning assumptions, and taking bold bets. - **Work Policy**: On-site in San Francisco (all roles require physical presence at HQ). - **Notable Perks**: Work at the frontier of AI agent infrastructure; join a small but rapidly growing team (19 people) backed by top-tier VCs; early employees will define the company’s trajectory and culture. - **Engineering Culture**: Emphasis on research engineering, rapid prototyping, and productionizing cutting-edge AI evaluation methods. ## Sources 1. [judgmentlabs.ai](https://judgmentlabs.ai/) 2. [judgmentlabs.ai/careers](https://www.judgmentlabs.ai/careers) 3. [linkedin.com/company/judgmentlabs](https://www.linkedin.com/company/judgmentlabs) 4. [builtin.com](https://builtin.com/company/judgment-labs) 5. [businesswire.com](https://www.businesswire.com/news/home/20260512621556/en/Judgment-Labs-Closes-%2432M-in-Seed-and-Series-A-Funding-to-Build-the-Continuous-Improvement-Layer-for-AI-Agents) ## Other roles at Judgment Labs - [Product engineer, Agent](https://feeny.ai/job/product-engineer-agent-judgment-labs-san-francisco-t4ryveqq2yns) — San Francisco, CA - [Founding Account Executive](https://feeny.ai/job/founding-account-executive-judgment-labs-san-francisco-973j3geb8a8e) — San Francisco, CA - [Product engineer, full stack](https://feeny.ai/job/product-engineer-full-stack-judgment-labs-san-francisco-rzn6c0km6mxc) — San Francisco, CA - [Applied AI Engineer](https://feeny.ai/job/applied-ai-engineer-judgment-labs-san-francisco-2c68arqvx3sp) — San Francisco, CA