--- title: 'Research Engineer – Benchmarking at Mercor' canonical: 'https://feeny.ai/job/research-engineer-benchmarking-mercor-san-francisco-2myynbnw8b7v' type: 'job' last_seen: '2026-09-05' --- # Research Engineer – Benchmarking at Mercor - **Company:** Mercor - **Location:** San Francisco, CA - **Compensation:** $180k–$500k - **Employment:** full-time - **Work type:** onsite - **Posted:** 2026-08-18 - **Last confirmed live:** 2026-09-05 - **Apply:** https://jobs.ashbyhq.com/mercor/40cc6334-3a6c-42f8-a4ec-d490486d00e8 ## Job description ## ABOUT MERCOR Mercor's mission is to organize human intelligence to power the AI economy. We're a leading AI data company, building the layer between human expertise and frontier models. Millions of domain experts on the platform are paid over $4 million per day to train frontier AI models. Mercor's APEX benchmark family measures AI's real-world impact on professional work. Mercor Enterprise brings this same infrastructure to Fortune 500 companies: helping companies capture how their best people actually work, translating that expertise directly back into agents. Mercor is creating a new category of work where expertise powers AI advancement. Achieving this requires an ambitious, fast-paced and deeply committed team. You’ll work alongside researchers, operators, and AI companies at the forefront of shaping the systems that are redefining society. Mercor is a profitable Series C company valued at $10 billion. We work in-person five days a week in our San Francisco, NYC, or London offices. ## ABOUT THE ROLE As a Research Engineer at Mercor, you’ll work at the intersection of engineering and applied AI research. You’ll own benchmarking pipelines, evaluation systems, and failure analysis workflows that directly inform how we train and improve frontier language models. Your work will define how we measure tool use, agentic behavior, and real-world reasoning. You’ll design and run evals, build rubrics and scorers, and turn failure analysis into actionable improvements for post-training, RLVR, and data pipelines. ## WHAT YOU’LL DO - Benchmarking: Design, implement, and maintain benchmarks and metrics for tool use, agentic behavior, and real-world reasoning; ensure benchmarks scale with training and stay aligned with product and research goals. - Evaluation systems: Build and operate LLM evaluation systems end-to-end runs, scoring, dashboards, and reporting, so researchers and applied AI teams can track model performance and compare runs at scale. - Failure analysis: Run systematic failure analysis on model outputs (e.g., wrong tool use, reasoning errors, safety/alignment issues); categorize failure modes, quantify prevalence, and feed findings into reward design, data curation, and benchmark design. - Rubrics and evaluators: Create and refine rubrics, automated evaluators, and scoring frameworks that drive training and evaluation decisions; balance rigor with scalability (human vs. model-as-judge, calibration, agreement). - Data quality and usability: Quantify data usability, quality, and impact on key benchmarks; use evals and failure analysis to guide data generation, augmentation, and curation. - Cross-team collaboration: Work with AI researchers, applied AI teams, and data producers to align evals with training objectives and to prioritize benchmarks and failure analyses that matter most. - Ownership in a fast-paced environment: Operate in a high-iteration research setting with strong ownership of benchmarks, evals, and failure-analysis workflows. ## WHAT WE’RE LOOKING FOR - Strong applied research background, with focus on model evaluation, benchmarking, and/or failure analysis. - Strong coding skills and hands-on experience with ML models and evaluation code. - Solid grasp of data structures, algorithms, and backend systems. - Comfort with APIs, SQL/NoSQL, and cloud platforms for running and storing eval results. - Ability to reason about model behavior, experimental results, and data quality from evals and failure analyses. - Excitement to work in person in San Francisco five days a week in a high-intensity, high-ownership environment. ## NICE TO HAVE - Industry experience on a post-training or evaluation/benchmarking team (highest priority). - Publications at top-tier venues (NeurIPS, ICML, ACL), especially in evaluation or benchmarking. - Experience building or running LLM evaluations, benchmarks, or failure-analysis pipelines. - Experience with synthetic data generation, rubric design, or RL-style workflows that use evals for reward shaping. - Work samples or code (e.g., eval frameworks, benchmark suites, failure-analysis reports or tooling) that demonstrate relevant skills. ## BENEFITS - Bi-annual performance bonus structure - Generous equity grant vested over 4 years - Up to $15k Relocation bonus - $10K housing bonus (if you live within 0.5 miles of our office) - $1.5K monthly stipend for meals - Free Equinox membership - $200 monthly laundry reimbursement - $200 monthly personal wellness reimbursement - Health, Dental, Vision insurance ## About Mercor ## Company Overview - **One-liner**: Mercor organizes human intelligence to power the AI economy by connecting domain experts with frontier AI labs and enterprises for model training, evaluation, and deployment. - **Entity Type**: Private (Series C) - **Headquarters**: San Francisco, California, USA (with offices in New York City and London) - **Founded**: 2023 - **Founders**: Brendan Foody & Adarsh Hiremath (Co-founders & Co-CEOs) ## Core Business - Primary industry: AI Data & Infrastructure, Human-in-the-Loop AI, Enterprise AI - Target customers: Frontier AI labs (OpenAI, Anthropic, etc.), Fortune 500 enterprises deploying AI in production, and domain experts seeking high-value AI work - Mission: Organizing human intelligence to power the AI economy — building the layer between human expertise and frontier models. ## Products & Services - **Mercor Platform**: A global network of millions of domain experts across 300+ professional fields, paid over $3 million per day to train frontier AI models. Experts are found and vetted using dynamic, role-specific AI interviews. - **APEX Benchmark**: A benchmark family that measures AI's real-world impact on professional work, setting a standard for whether AI can perform economically valuable work. - **Mercor Enterprise**: Brings Mercor's expert network and evaluation infrastructure to Fortune 500 companies, deploying custom AI agents, staffing teams with vetted domain experts, and helping organizations encode their own knowledge into AI systems. ## Market Standing - **Valuation**: $10 billion valuation (as of Series C) - **Key Metric**: Total Funding — $350 million Series C (recently raised); company is profitable - **Notable Investors/Partners**: Not explicitly named in provided sources, but the company works with leading AI labs and Fortune 500 enterprises - **Growth Signals**: - Named to Bloomberg's "2026 24 AI Startups to Watch" - Forbes AI 50 (2025) - Forbes Cloud 100 (2025) - LinkedIn Top Startup (2025) - Described as "the fastest-growing company in the world" - Tens of thousands of skilled professionals earning an average of more than $95/hour on the platform ## Competitive Advantages - **Massive, vetted expert network**: Millions of domain experts across 300+ fields, found and vetted via AI-driven interviews — a hard-to-replicate moat. - **Direct frontier model exposure**: Every team works directly with frontier models, training them and building around them — most companies haven't imagined this level of integration. - **Profitability at scale**: Profitable at Series C with a $10B valuation, indicating strong unit economics. - **End-to-end AI infrastructure**: From expert sourcing and vetting to model evaluation (APEX) to enterprise deployment — a vertically integrated stack. ## Strategic Focus - **Scaling the expert network**: Continuing to onboard millions of domain experts globally to meet the insatiable demand for high-quality human training data. - **Enterprise AI deployment**: Expanding Mercor Enterprise to help Fortune 500 companies encode organizational knowledge into AI systems and deploy custom agents. - **Benchmarking AI's economic impact**: Growing the APEX benchmark as the standard for measuring whether AI can perform economically valuable work. - **Preparing for AGI**: The company's founding thesis is about answering what role humans will play in the economy as AGI approaches. ## Why Work Here - **Culture**: "Built for people who thrive in high-velocity environments." Values include Intensity, Simplicity, User Obsession, High Standards, Can-Do Attitude, and Time Travelers (thinking ahead). - **Work model**: In-person collaboration from offices in San Francisco, New York City, and London. No remote option — the company believes in-person collaboration helps move faster and solve harder problems. - **Engineering culture**: Direct exposure to frontier AI research, labs, and startups. Every team works with frontier models daily. Described as "the fastest-growing company in the world." - **Compensation**: Generous equity grants, competitive salaries (ranges from $100K–$500K depending on role), and a "proximity bonus" (generous annual housing support). - **Benefits**: - Unlimited paid vacation and sick days - Free Equinox membership (premium fitness access) - Monthly wellness stipend and free mental health services - Monthly food stipend - 401k with employer match - Parental leave for both birthing and non-birthing parents - **Team**: "A team of ambitious, diverse people who want to work on what matters most." Open roles across Engineering (31 roles), Operations (15 roles), Enterprise (5 roles), Finance (4 roles), and more. ## Sources 1. [mercor.com](https://www.mercor.com/) 2. [mercor.com/careers](https://www.mercor.com/careers/) 3. [mercor.com/mission](https://www.mercor.com/mission/) 4. [jobs.ashbyhq.com/mercor](https://jobs.ashbyhq.com/mercor) 5. [linkedin.com/company/mercor-ai](https://www.linkedin.com/company/mercor-ai) ## Other roles at Mercor - [Infrastructure Engineer](https://feeny.ai/job/infrastructure-engineer-mercor-san-francisco-1t8t2mejenj2) — San Francisco, CA - [Security Engineer, Application Security](https://feeny.ai/job/security-engineer-application-security-mercor-san-francisco-5z2pbdwzrw0n) — San Francisco, CA - [Technology Operations Specialist (L3)](https://feeny.ai/job/technology-operations-specialist-l3-mercor-san-francisco-evpqqw1fmc47) — San Francisco, CA - [Cloud Platform Engineer (SF)](https://feeny.ai/job/cloud-platform-engineer-sf-mercor-san-francisco-0c9b2zbpmk8g) — San Francisco, CA - [Software Engineer, Applied AI](https://feeny.ai/job/software-engineer-applied-ai-mercor-new-york-29sf9gcj7j2s) — New York, NY - [Software Engineer, Platform](https://feeny.ai/job/software-engineer-platform-mercor-new-york-8rn1qgb4yxmm) — New York, NY - [GTM Recruiter](https://feeny.ai/job/gtm-recruiter-mercor-mercor-hq-69mmwz3epxe8) — Mercor Hq, 181 Fremont Street - [Recruiter](https://feeny.ai/job/recruiter-mercor-san-francisco-n76yb0v38rbp) — San Francisco, CA - [Revenue Operations](https://feeny.ai/job/revenue-operations-mercor-san-francisco-jgq6v727w1n0) — San Francisco, CA - [Member of Technical Staff, Agentic Systems](https://feeny.ai/job/member-of-technical-staff-agentic-systems-mercor-san-francisco-et81vk2eeq51) — San Francisco, CA