--- title: 'Task Development Engineer at METR' canonical: 'https://feeny.ai/job/task-development-engineer-metr-remote-5cbf872b6d7w' type: 'job' last_seen: '2026-09-11' --- # Task Development Engineer at METR - **Company:** METR - **Location:** Remote - **Employment:** contract - **Work type:** remote - **Posted:** 2026-08-01 - **Last confirmed live:** 2026-09-11 - **Apply:** https://jobs.lever.co/metr/b4812bf4-c259-406b-8ffa-4a463fff34f7 ## Job description ## About METR We are a nonprofit research organization that develops scientific methods to assess AI capabilities, risks, and mitigations, with a specific focus on threats related to AI R&D automation and misalignment. We believe it is robustly good for policymakers and civil society to have a clear understanding of risks from AI systems, and we are extremely excited to build a team of ambitious, excellent people to tackle one of the most important challenges of our time. ## About the role - As part of informing the world about risk from frontier AI systems, METR often runs and [p](https://metr.org/blog/2026-05-19-frontier-risk-report/)[ublishes](https://metr.org/blog/2026-05-19-frontier-risk-report/) evaluations of frontier models. - [Time Horizons](https://metr.org/time-horizons/) is a central tool the world uses to understand AI progress. Our methodology has been included in [system](https://www-cdn.anthropic.com/08ab9158070959f88f296514c21b7facce6f52bc.pdf)[cards](https://arxiv.org/pdf/2601.03267), called an ["obsession" by the NYT](https://www.nytimes.com/2026/04/17/technology/how-do-you-measure-an-ai-boom.html), [has wide reach](https://x.com/METR_Evals/status/2024923422867030027) online, and is [used by governments](https://www.aisi.gov.uk/frontier-ai-trends-report) to inform national policy. It is essential to our [broader risk assessment work](https://thezvi.substack.com/p/the-most-important-charts-in-the) to have good capability evaluations. - Task Development Engineers contribute to METR’s expanding ambition of our evaluations with high quality tasks supporting the Time Horizons methodology. We expect our results to be seen by policymakers, frontier labs, national security stakeholders, and other key decisionmakers influencing society’s response to AI progress. What this role looks like - (Primarily, and most importantly) Developing difficult, novel tasks for models. You will build well-scoped tasks that remain challenging as model time horizons grow, potentially to hundreds of hours. - Quality assurance for existing tasks. Once a task has been developed, you will verify that it's actually solvable as specified, and that the model is given (only) the information it needs. - Baselining and scoring tasks. Where helpful, you may be asked to baseline tasks within your domain of expertise, and/or score task completions from AIs or human baseliners. - Improving task development infrastructure. We're always improving our processes. Strong candidates will notice when existing workflows are inefficient or produce low-quality output, and take responsibility for improving them. ## Skills we're looking for - Software engineering: You have several years of experience working on complex projects and codebases. - Evaluations: You have experience building hard (ideally agent-based) AI evaluations (e.g. RE-Bench, HCAST, SWE-bench Verified, Cybench, GPQA), ideally using the [Inspect](https://inspect.aisi.org.uk/) framework. - High attention to detail: You read closely, spot misspecifications and ambiguity, and pay attention to fiddly minutiae. - (Nice to have) Familiarity with METR infrastructure: Prior experience with [Hawk](https://hawk.metr.org/), and familiarity with the methodology behind our [Time Horizons](https://metr.org/time-horizons/) work, is a plus. Job details and compensation - Location: Remote (worldwide) - Hours: 20-40 hours per week (flexible schedule determined by you) - Timezone Requirements: A minimum of 1 hour (and ideally 4 hours) of overlap with the Pacific Coast Time workday, but you determine your exact work schedule. - Employment type: Contract / freelance - You decide the manner in which you complete your tasks to a standard that matches other professionals in this field. - Compensation: $150-300/hour. - Top of this range is reserved for exceptional candidates. - Individuals who contribute >80 hours will be acknowledged in the final research output (if desired). Our Culture METR is a mission-driven organization. We believe our work can meaningfully shape humanity's future for the better, and we want to be the best people in the world doing this work. We have a tight-knit, collaborative research culture rooted in truth-seeking and integrity. We're fiercely committed to producing high-quality, trustworthy science. We're honest and transparent about our results, especially when they may go against the grain. We've earned trust as reliable partners who handle confidential information with care. We maintain a low-ego, drama-free environment focused on what matters. We encourage you to apply even if your background may not seem like the perfect fit! We would rather review a larger pool of applications than risk missing out on a promising candidate for the position. We are committed to diversity and equal opportunity in all aspects of our hiring process. We do not discriminate on the basis of race, religion, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status. We welcome and encourage all qualified candidates to apply for our open positions. ## About METR ## Company Overview - **One-liner**: METR (pronounced "meter") is a non-profit research organization that develops scientific methods to assess catastrophic risks from advanced AI systems by evaluating their autonomous capabilities. - **Entity Type**: Non-profit (funded by donations, not a typical private company; no equity) - **Headquarters**: Berkeley, California, USA - **Founded**: Not explicitly stated, but founded by Beth Barnes; likely ~2021–2022 - **Founders**: Beth Barnes (Founder, CEO) ## Core Business - **Primary industry**: AI safety research and evaluation - **Target customers**: AI developers (e.g., OpenAI, Anthropic, Google DeepMind, Meta, Amazon), governments, and the broader public - **Mission or purpose**: "Develop scientific methods to assess catastrophic risks stemming from AI systems’ autonomous capabilities and enable good decision-making about their development." ## Products & Services - **Frontier AI Capability Evaluations**: Systematic assessments of how autonomously AI systems can perform tasks (e.g., conducting research, developing apps, cyberattacks, self‑hardening). Published as open research. - **Risk Assessments & Safety Policies**: Advises AI developers and governments on risk assessment methodologies, including the "Responsible Scaling Policies" approach adopted by nine leading AI labs. - **Time‑Horizon Research**: Open‑source analysis showing that the length of tasks AI agents can complete doubles every ~7 months, a key input for forecasting transformative AI timelines. - **Evaluation Platform**: An open‑source platform built on Inspect AI for running AI agent evaluations at scale. - **Monitorability Evaluations**: Research on detecting AI agents attempting to evade monitoring or perform side tasks, including datasets of "reward hacking" and "sandbagging" behaviors. - **Productivity RCT**: A randomized controlled trial with experienced open‑source developers measuring how much AI tools actually boost productivity (finding systematic overestimation). ## Market Standing - **Valuation/Market Cap**: Not applicable (non‑profit) - **Key Metric**: Funding – METR is supported by donations from major foundations and individuals. Notable donors include The Audacious Project (TED), Jane Street, Sijbrandij Foundation, Pew Charitable Trusts, Schmidt Sciences, Packard Foundation, and others. Small part of income from a technical assistance contract with the European AI Office. **No funding from AI companies** (maintains independence). - **Notable Investors/Partners**: Partners with OpenAI, Anthropic, Google DeepMind, Meta, Amazon (pilot risk assessments); member of NIST AI Safety Institute Consortium, California Cybersecurity Task Force, UK AI Security Institute; technical assistance to European AI Office. - **Growth Signals**: Growing team (multiple open roles), expanding into cyberforensics and embedded assessments, research cited widely in AI safety policy, and adoption of their Responsible Scaling Policies by major developers. ## Competitive Advantages - **Independence**: No funding from AI companies, enabling unbiased, transparent research. - **Scientific Rigor**: Publishes all research openly; uses empirical methods (RCTs, time‑horizon analysis) rather than speculation. - **Policy Influence**: Their frameworks (e.g., Responsible Scaling Policies) are now industry standards, and they advise governments globally. - **Deep Technical Expertise**: Team of researchers and engineers building state‑of‑the‑art evaluations for autonomous capabilities, including security‑relevant evaluations. ## Strategic Focus - **Current priorities**: Developing methodologies to track AI loss‑of‑control risk, improving monitorability evaluations, expanding cyberforensics capabilities, and scaling the team to meet growing demand for independent evaluations. - **Direction for growth**: Increasing influence on AI governance, deepening partnerships with governments and companies, and building tools for continuous risk assessment. ## Why Work Here - **Mission‑driven**: Direct contribution to aligning AI development with public safety, working on one of the most critical problems of our time. - **Compensation**: Highly competitive with top AI labs (salary ranges $328K–$687K for technical staff, $150–$300/hr for contractors), plus benefits like medical/dental/vision, wellness ($1,500/yr), mental health ($6,000/yr), professional development ($5,250/yr), unlimited PTO. - **Work environment**: Small, fast‑moving, mission‑driven team based in Berkeley. On‑site preferred for technical roles (at least a few days a week), but hybrid and remote (including international) can be accommodated. Operations roles require in‑person. - **Hiring process**: Unique focus on work tests (1–3 take‑home tasks) and a 1–2 day paid work trial (travel and lodging reimbursed). Interviews are given substantially less weight. Commitment to diversity and equal opportunity. - **Culture**: Open, transparent, and empirical; values rigor and independence. Visa sponsorship available for technical roles. ## Sources 1. [metr.org](https://metr.org/) 2. [metr.org/about](https://metr.org/about#our-team) 3. [metr.org/careers](https://metr.org/careers) 4. [metr.org/hiring](https://metr.org/hiring) 5. [jobs.lever.co/metr](https://jobs.lever.co/metr) ## Other roles at METR - [General Counsel](https://feeny.ai/job/general-counsel-metr-berkeley-4s59mgfhjagg) — Berkeley, CA - [Member of Technical Staff, Cyberforensics](https://feeny.ai/job/member-of-technical-staff-cyberforensics-metr-berkeley-kxzz3emzx9gj) — Berkeley, CA - [System Administrator](https://feeny.ai/job/system-administrator-metr-berkeley-87zqbgq2ffsv) — Berkeley, CA - [Member of Technical Staff, Embedded Assessments](https://feeny.ai/job/member-of-technical-staff-embedded-assessments-metr-berkeley-jsprvy2fdmqy) — Berkeley, CA - [Member of Technical Staff, Security Engineering](https://feeny.ai/job/member-of-technical-staff-security-engineering-metr-berkeley-w44zdwnsrdma) — Berkeley, CA - [Member of Technical Staff, Evaluation Execution](https://feeny.ai/job/member-of-technical-staff-evaluation-execution-metr-berkeley-rtt64jwna58y) — Berkeley, CA - [General Expression of Interest](https://feeny.ai/job/general-expression-of-interest-metr-berkeley-avs3pft13sf0) — Berkeley, CA