--- title: 'Software Engineer, Benchmarking at Epoch AI' canonical: 'https://feeny.ai/job/software-engineer-benchmarking-epoch-ai-remote-8taq5p4eccd2' type: 'job' last_seen: '2026-09-16' --- # Software Engineer, Benchmarking at Epoch AI - **Company:** Epoch AI - **Location:** Remote - **Employment:** full-time - **Work type:** remote - **Posted:** 2026-06-29 - **Last confirmed live:** 2026-09-16 - **Apply:** https://jobs.lever.co/epoch-ai/d172645e-a11f-44a0-88d0-7a989e0a28f6 ## Job description Epoch AI is looking for a Software Engineer who will help us evaluate frontier AI models, enabling researchers, developers, and policymakers to better understand AI development. The role will involve running and maintaining our benchmarking infrastructure as well as contributing to the development of brand new benchmarks. ## About the role Please do not include a cover letter, photograph, or headshot of yourself, or any personal information that is not relevant to the role for which you're applying (including marital status, age, identity traits, etc.). We are looking for a Software Engineer to help us expand and develop our AI Benchmarking Hub. You will work closely with the rest of the benchmarking team to run and maintain benchmarks, integrate with AI providers, set up existing benchmarks to run on our infrastructure, help design and develop brand new benchmarks, and facilitate internal experiments. This role is fully remote, and we are able to hire in many countries. We invite anyone who is interested to apply, regardless of background, experience, or credentials. Applications are rolling. ## Key Responsibilities - Implement benchmarks: Implement AI benchmarks within our evaluation infrastructure (primarily using the Inspect library) to expand the suite of capabilities we track. Develop our existing suite of benchmarks so we can quickly and painlessly evaluate new model releases. - Develop new benchmarks: Contribute to the development of brand new benchmarks. You will have the opportunity to pitch and prototype your own ideas in addition to helping out with existing projects. - Collaborate: Work closely with researchers, analysts, and other engineers at Epoch AI to ensure evaluation data and outputs are accurate, insightful, and effectively integrated into our research products and publications. ## What we are looking for - Solid engineering skills: A strong software engineering background with more than two years of professional experience building and maintaining complex systems. You are expected to regularly contribute high-quality, robust, and maintainable code and be comfortable diving deep into existing codebases and infrastructure. - Ideas and creativity: Candidates should be able to generate their own ideas for new benchmarks, experiments, novel things to try, and other projects. - Mission-driven: You’re motivated by Epoch AI’s mission to provide rigorous, independent insight into key trends in AI. You want to deliver public, trustworthy evaluations of AI capabilities on challenging benchmarks, empowering researchers, policymakers, and the wider public to make well-informed decisions about AI. - Ability to travel: Attendance at our three annual team retreats is strongly encouraged for all staff. AI domain expertise is a strong plus but not required. (This includes hands-on experience running LLM evaluations, familiarity with evaluation frameworks like [Inspect](https://inspect.aisi.org.uk/), as well as a solid grasp of current AI trends.) Solid engineering skills and an ability to learn quickly matter more than direct background in these areas. ## Compensation & Benefits - Annual salary between $150,000 and $325,000 USD, depending on experience, seniority, and location. - Salaries are not restricted to USD, and contracts and payments are usually in local currencies. Conversions are based on one-year average exchange rates. - Fully remote environment, including flexible work hours and schedules for most roles. - Competitive global benefits program, including a comprehensive health insurance program—including supplemental benefits specific to a local country, as available and mandated by local law—and life insurance and a pension plan, if applicable in your country. - Generous paid time off (PTO), including no specific limit on PTO with 30 days per year protected, unlimited personal and sick leave, and up to 6 months (combination of paid + unpaid) parental leave for permanent staff. - A flexible and generous expense policy for you to spend on equipment and a large range of productivity tools or learning/development opportunities you might find valuable, subject to regulations and manager approval. - Access to our very well-equipped offices in Berkeley, California, including paid meals, snacks, gym, and more. Additional Information While we welcome applicants from all time zones, we prefer candidates who can overlap with UTC–8 (Pacific Time) and UTC (Greenwich Mean Time), as most of our staff work in this range of time zones. We also prefer candidates who can travel: we hold three retreats per year, during which we record podcast episodes and other communication efforts. Please submit all of your application materials in English and note that we require professional level English proficiency. Epoch is committed to building an inclusive, equitable, and supportive community for you to thrive and do your best work. We’re committed to finding the best people for our team, so please don’t hesitate to apply for a role regardless of your age, gender identity/expression, political identity, personal preferences, physical abilities, veteran status, neurodiversity or any other background. Please emailcareers@epoch.ai if you have any questions about this role, accessibility requests, or if you want to request an extension to the application deadline. However, we will not review applications submitted to this email address; please submit your application through the link on this page. ## About Epoch AI Epoch AI is a research institute that investigates trends in machine learning and the economic consequences of AI. Our mission is to develop a comprehensive, publicly accessible knowledge base on AI that informs policymakers, industry leaders, and society at large. We strive to achieve both rigor and accessibility to our work, as exemplified by some of our most successful projects, including our database of AI models and our AI trends dashboard. Our body of research includes our work on compute trends (IJCN 2022), data scarcity (ICML 2024), and algorithmic progress (NeurIPS 2024). You can read more about our work and mission on our website and in this Time profile. ## About Epoch AI ## Company Overview - **One-liner**: Epoch AI is a data-first research nonprofit that investigates the trajectory, progress, and impact of artificial intelligence through empirical analysis, open databases, and independent model benchmarks. - **Entity Type**: Private nonprofit research organization (relies on philanthropic funding and mission-aligned paid services) - **Headquarters**: Remote-first, global team; some roles listed in the Bay Area. No single public headquarters specified. - **Founded**: 2021 (began as a group of volunteers; formally organized as Epoch AI following the success of their 2022 training-compute paper) - **Founders**: Not publicly available in the sources reviewed ## Core Business - **Primary industry**: AI research, data analytics, benchmarking, and policy-oriented intelligence - **Target customers**: Policymakers, journalists, AI developers, researchers, government agencies, large AI companies, nonprofits, and organizations affected by AI (including hardware, energy, and investment firms) - **Mission**: "To improve society's understanding of the drivers, progress, and impact of artificial intelligence" — building a shared scientific foundation for thinking about AI, neutral to any specific agenda, so decisions about the technology are informed by the best available data ## Products & Services - **FrontierMath**: A state-of-the-art private benchmark for testing frontier mathematical capabilities of AI models, commissioned by OpenAI. Epoch AI also runs pilots for software engineering and remote-work benchmarks. - **AI Model Evaluations**: Independent, regular evaluations of publicly available models across domains, published on a public dashboard to illustrate capability trends. - **AI Trend Databases**: Open-source intelligence program tracking AI models, training compute estimates, ML hardware, GPU clusters, AI data centers, and investment — the foundation of their research and public explainers. - **Commissioned Research & Consulting**: Paid, mission-aligned services for companies, nonprofits, and government bodies, including data collection, benchmark development, model evaluations, and private strategic reports (e.g., compute trend analysis for UK ARIA, biological sequence model analysis for Sentinel Bio). - **Interactive Visualizations & Explainers**: Public tools and analyses that distill complex AI trends into accessible formats for a broad audience. ## Market Standing - **Valuation/Market Cap**: Not applicable — independent nonprofit; not disclosed - **Key Metric**: Total funding not publicly disclosed; supported by philanthropic donations, revenue from commissioned research, and public-good funding - **Notable Partners/Clients**: OpenAI, Google, Stanford's AI Index, UK Department for Science, Innovation & Technology (DSIT), UK ARIA, Sentinel Bio, Electric Power Research Institute - **Growth Signals**: Grew from a volunteer collective (2021) into a multidisciplinary, global team; authored one of the most-cited papers in the field, *Compute Trends Across Three Eras of Machine Learning*; gained recognition for FrontierMath; expanded into model evaluations and public dashboards; routinely engaged by major AI labs and government agencies ## Competitive Advantages - **Neutral, data-first approach**: Explicitly non-partisan and "agenda-free," producing evidence that serious AI discussions start from — a differentiator in a hype-heavy field - **Pioneering compute-trend tracking**: Among the first to systematically monitor AI training compute, inference-compute scaling, and scaling feasibility through 2030 - **Trusted by both sides**: Works with frontier AI labs (OpenAI, Google) and government bodies (UK DSIT, UK ARIA), indicating credibility across industry and policy - **Open by default**: Publishes most research and data publicly, building a compounding reputation and research ecosystem - **FrontierMath benchmark**: The state-of-the-art benchmark for frontier math capabilities, giving Epoch AI unique leverage in model evaluation ## Strategic Focus - Tracking AI capabilities over time and maintaining independent evaluations of new models - Investigating the drivers and bottlenecks of AI progress, including whether scaling can continue through 2030 and the feasibility of decentralized training - Measuring and forecasting AI's economic and societal impact (e.g., automation of remote work, economy-wide modeling) - Expanding benchmark development and partnerships with AI companies, governments, and adjacent sectors (energy, hardware, investment) - Building out the public dashboard and interactive tools to make research accessible ## Why Work Here - **Remote-first, global team**: Roles are listed as full-time or part-time remote, with some positions in the Bay Area — flexibility across research, engineering, and operations - **Mission-driven work**: Directly contribute to society's understanding of AI, with a "curiosity-driven" and multidisciplinary culture - **High-impact exposure**: Work alongside researchers, engineers, and policymakers shaping AI discourse; past partners include OpenAI and UK DSIT - **Culture**: The team explicitly values caring for each other and taking wellbeing seriously; transparency and open communication are core principles - **Variety of roles**: Openings span research (Data Scientist, Researcher), engineering (Software Engineer, Benchmarking), operations (People Ops, IT & Security), and communications — plus rolling expressions of interest for interdisciplinary talent ## Sources 1. [epoch.ai](https://epoch.ai/) 2. [epoch.ai/about](https://epoch.ai/about) 3. [epoch.ai/about/careers](https://epoch.ai/about/careers) 4. [epoch.ai/latest/what-is-epoch](https://epoch.ai/latest/what-is-epoch) 5. [jobs.lever.co/epoch-ai](https://jobs.lever.co/epoch-ai) ## Other roles at Epoch AI - [Writer & Editor](https://feeny.ai/job/writer-editor-epoch-ai-bay-area-wsb2bp38fyd0) — Bay Area - [Social Video Producer](https://feeny.ai/job/social-video-producer-epoch-ai-bay-area-tf0x02gpnzcb) — Bay Area - [Expression of Interest: Special Projects Associate](https://feeny.ai/job/expression-of-interest-special-projects-associate-epoch-ai-remote-h9t4142pd5xe) - [IT and Security Specialist](https://feeny.ai/job/it-and-security-specialist-epoch-ai-remote-bgjrekpdryb4) - [Recruiter / Senior Recruiter](https://feeny.ai/job/recruiter-senior-recruiter-epoch-ai-remote-njc0gpzapvbh) - [Senior Product Designer](https://feeny.ai/job/senior-product-designer-epoch-ai-remote-vp76qx0k5ekr) - [Researcher / Senior Researcher](https://feeny.ai/job/researcher-senior-researcher-epoch-ai-remote-6r750xw78dnc) - [Data Scientist](https://feeny.ai/job/data-scientist-epoch-ai-remote-0s1wpwz0qx9c) - [Expression of Interest](https://feeny.ai/job/expression-of-interest-epoch-ai-remote-20n4q62x68fx)