--- title: 'Senior Software Engineer, RL Environments at Pareto' canonical: 'https://feeny.ai/job/senior-software-engineer-rl-environments-pareto-san-francisco-jx9sfz7cyqn3' type: 'job' last_seen: '2026-09-12' --- # Senior Software Engineer, RL Environments at Pareto - **Company:** Pareto - **Location:** San Francisco, CA - **Compensation:** $245k–$300k - **Employment:** full-time - **Work type:** hybrid - **Posted:** 2026-08-20 - **Last confirmed live:** 2026-09-12 - **Apply:** https://jobs.ashbyhq.com/pareto-ai/8e05ddf2-4876-478a-9a36-a4b65a3d588f ## Job description ## About Pareto Humanity is in a virtuous cycle: human insight improves AI, and better AI expands what people can do. Sustaining it depends on the one input that can't be automated: [expert human judgment](https://pareto.ai/blog/debating-persuasive-llms-truthful-answers). At Pareto, we build the platform that turns that judgment into the [data](https://pareto.ai/blog/community-driven-knowledge-resource-ai), [evals](https://pareto.ai/blog/introducing-attunebench), and RL environments [frontier models](https://pareto.ai/blog/llm-metacognition-shared-and-shallow) learn from. We work with leading frontier labs like Anthropic and GDM, and we give skilled people everywhere a way to shape the future of AI and share in what it creates. This RL environment and human-data infrastructure is already in production. Our job now is to scale it. ## The Role You'll own the RL environments frontier labs train on, end to end. Scope the problem with the requester, build the image and the tools inside it, write the graders that score it, ship it into the customer's platform, and keep it healthy once it's running. You sit between Pareto's engineering team and the researchers at the labs we work with, close enough to both that you can tell when a training goal and a buildable spec have drifted apart. Nobody will hand you a finished spec. You'll get a research problem, define what gets built, and stay with it after it lands. In your first year, good looks like environments that ship faster than the last one did, because you invested in the build and release path instead of hand-rolling each delivery. What you build becomes training signal. That's the reason the ownership runs all the way through production. ## What You'll Own - The environment, end to end. Build industry-leading RL environments and MCP tools that power post-training loops for frontier labs. - Scoping with the requester. Partner closely with research labs, turn a rough training goal into a spec you can build against, and push back early when the ask won't produce usable signal. - Production health. Investigate failed tasks and jobs and take the lead when a delivery pipeline degrades. - Platform leverage. Automated image builds and release notes, spec-first design, CI gates, review harnesses. The next environment should cost a fraction of the last one. - The signal back into product. When the platform falls short of what a lab needs, you're the first to know. Getting that gap onto the roadmap is part of the job, not someone else's follow-up. ## What We're Looking For - You’ve been building production systems for 7+ years. Enough range to tell which problems need a careful design and which just need shipping. - You write production Python or TypeScript. Strong functional-language background counts too. You'll be reading unfamiliar code and shipping changes to it in the same week. - You've built and shipped containerized services. Docker layering, dependency pinning, images that behave the same on the tenth run as the first. Reproducibility is the whole game here. - You get real leverage from coding agents, and you review what they produce. Fluent use is table stakes. The judgment to catch what the agent got wrong is the differentiator. - You own things after they ship. You've been on the hook for something in production, worked an incident to root cause, and written the doc that kept it from happening twice. - You know your way around cloud infrastructure. Containers, managed databases, and the deploy path, on AWS or any major cloud. You don't need to be an infrastructure specialist, but you do need to debug your own deploys. - You're based in the US and can get to the Bay Area as needed. You Probably Aren't the Right Fit If You - Need requirements locked before you start. Specs here evolve with the research, and that's expected. - Prefer working at arm's length from stakeholders. This role is high-contact by design. - Want to hand off at launch. This role owns environments in production, including the bad weeks. - Only want to build one kind of thing. This role moves between image builds, grader design, and client work in the same week. ## Compensation Base salary $245K–$300K, plus equity. Final offer depends on experience and level, and we share level-specific ranges early in the process. ## Why Pareto The environments you build become the training signal for models at Anthropic and GDM. Not adjacent to that work. Inside it. Equity is part of the package at every level. Apply even if you don't match every line above. We care more about how you think than we do about a clean résumé. ## About Pareto ## Company Overview - **One-liner**: Pareto provides expert-curated data and human judgment signals to train and evaluate frontier AI models. - **Entity Type**: Private (Seed-stage; total funding $5.1M) - **Headquarters**: San Francisco, California, United States - **Founded**: 2020 - **Founders**: Phoebe Yao ## Core Business - **Industry**: AI training data and reinforcement learning infrastructure; Software Development - **Target Customers**: B2B – primarily AI research labs, frontier model companies (Anthropic, Character AI, MATS, METR), and enterprise teams building advanced AI systems. - **Mission/Purpose**: To transform how expert knowledge is captured, scaled, and integrated into AI development – building the infrastructure for training next‑generation models with humans at the center. ## Products & Services - **End‑to‑End Rater Management**: Rapid sourcing, onboarding, and training of expert raters at scale for on‑demand data work. - **Expert‑Led Evaluation**: Deep domain‑specific assessments to benchmark and improve model performance (e.g., finance, healthcare, engineering). - **Iterative Experiment Design**: Collaborative design of experiments from prompt engineering to gold‑standard evaluation setups. - **Global, High‑Quality Data Collection**: Task‑specific datasets designed to meet research goals, with multilingual and regional diversity. - **Modular Research Infrastructure**: End‑to‑end support for novel or complex research workflows, including feedback loops and real‑time refinement. ## Market Standing - **Valuation/Market Cap**: Not publicly available (private company) - **Key Metric**: Total Funding – $5.1M (Seed rounds: $600K in 2020, $4.5M led by 8 investors in 2022; additional seed investors in 2021) - **Notable Investors/Partners**: Investors not explicitly named in search results, but partners include **Anthropic**, **Character AI**, **MATS**, **METR** (all cited as clients/references on the solutions page) - **Growth Signals**: Headcount of 406 employees (+27% YoY); operates in 45 countries; LinkedIn followers grew +164% yearly; trusted by leading AI labs for high‑stakes evaluation work. ## Competitive Advantages - **Specialized Human Signal**: The company turns “nondeterministic expert judgment into durable reward signals” – a difficult, unsolved problem that many competitors avoid. - **Calibrated Tasking**: Measures each model’s capability frontier and calibrates tasks to where the model will learn the most, making expert knowledge trainable. - **Elite Expert Network**: Maintains a global, vetted pool of prompt engineers, annotators, and evaluators with deep domain expertise. - **Trust from Frontier Labs**: Strong testimonials from Anthropic, Character AI, MATS, and METR verify the quality and impact of their work. ## Strategic Focus - **Infrastructure for Next‑Gen AI**: Building the verification layer for reinforcement learning on real‑world expertise. - **Scaling Expert Knowledge**: Moving beyond simple labeling to capture complex judgment (taste, reasoning, specialist instinct) as training signal. - **Global, On‑Demand Operations**: Continual expansion of the expert network across 45+ countries to provide 24/7 coverage and linguistic diversity. ## Why Work Here - **Culture & Values**: Emphasizes “Default to open” (honesty and clarity), “Stay curious” (deep answers), “Build from connection” (human‑centered), and “Push for excellence” (continuous growth). - **Work Environment**: Global remote‑first operation with team members in 45 countries; headquarters in San Francisco (office in Stanford, CA). Collaboration across time zones encouraged. - **Engineering & Research Culture**: Head of Engineering, Chief Scientist, and Head of Applied AI roles signal strong technical leadership. The career page describes “the most interesting unsolved problem we know” – appealing to those passionate about AI research. - **Growth Trajectory**: Rapid headcount growth (+27% YoY) and high LinkedIn follower growth suggest strong momentum and career development opportunities. ## Sources 1. [pareto.ai/about-us](https://pareto.ai/about-us) 2. [pareto.ai/careers](https://pareto.ai/careers) 3. [pareto.ai/solutions](https://pareto.ai/solutions) 4. [linkedin.com/company/hellopareto](https://www.linkedin.com/company/hellopareto) ## Other roles at Pareto - [AI Engagement Manager](https://feeny.ai/job/ai-engagement-manager-pareto-united-states-9715330zxvwv) — United States