--- title: 'Eval360 - Error Analysis Engineer at Institute of Foundation Models' canonical: 'https://feeny.ai/job/eval360-error-analysis-engineer-institute-of-foundation-models-sunnyvale-4yt0zm9r4y3a' type: 'job' last_seen: '2026-09-14' --- # Eval360 - Error Analysis Engineer at Institute of Foundation Models - **Company:** Institute of Foundation Models - **Location:** Sunnyvale, CA - **Compensation:** $150k–$450k - **Employment:** full-time - **Work type:** onsite - **Posted:** 2026-06-19 - **Last confirmed live:** 2026-09-14 - **Apply:** https://jobs.lever.co/ifm-us/adf911e6-d3d0-41ed-a7bf-553e1c50e684 ## Job description ## About the Institute of Foundation Models The Institute of Foundation Models is a dedicated research lab focused on building, understanding, using, and risk-managing foundation models. Our mission is to advance AI research, support the next generation of AI builders, and develop impactful systems that improve how frontier models are trained, evaluated, deployed, and governed. As part of our team, you will work closely with researchers, machine learning engineers, data scientists, software engineers, and product teams on some of the most important challenges in AI development. You will contribute to systems that help measure model quality, identify failure modes, and improve the reliability, safety, and readiness of model releases. ## The Role We are looking for an Eval360 - Error Analysis Engineer to help build, improve, and operate Eval360, an evaluation service that serves as a quality gate for AI models. This person will focus specifically on error analysis: understanding where models fail, why they fail, how those failures should be categorized, and how evaluation systems can better detect, measure, and prevent these issues before models are released. You will collaborate with researchers, machine learning engineers, product managers, data scientists, and platform teams to develop AI evaluation applications and internal tools based on next-generation AI research. You will be part of a cross-functional team responsible for the full software development lifecycle, from requirements gathering and system design to implementation, deployment, monitoring, debugging, documentation, and continuous improvement. The ideal candidate is comfortable working across the stack, including front-end interfaces for reviewing errors, back-end evaluation pipelines, data analysis workflows, model evaluation infrastructure, databases, dashboards, and APIs. This person should have strong software engineering skills, excellent analytical judgment, and the ability to turn ambiguous model failures into structured insights that improve evaluation quality. ## Key Responsibilities - Collaborate with researchers, machine learning engineers, data scientists, product managers, and internal stakeholders to implement innovative software solutions for Eval360 and related model evaluation workflows. - Build and improve Eval360 as an evaluation service that acts as a quality gate for model development, model comparison, and model release decisions. - Perform deep error analysis on model outputs, including identifying failure patterns, categorizing issues, tracing root causes, and proposing improvements to evaluation methodology. - Develop tools, workflows, and dashboards that make it easier for researchers and engineers to inspect model failures, compare model behavior, and understand quality regressions. - Design and implement client-side and server-side architecture for evaluation review systems, error analysis interfaces, reporting tools, and internal evaluation applications. - Develop responsive, usable interfaces that support error triage, annotation review, evaluation debugging, and model quality investigation. - Build and maintain back-end services, APIs, data pipelines, and integrations that support evaluation execution, results storage, analysis, and reporting. - Test software to ensure responsiveness, correctness, reliability, and efficiency across evaluation workflows. - Troubleshoot, debug, and upgrade evaluation systems, including identifying issues in data processing, evaluation metrics, model output handling, job orchestration, and user-facing analysis tools. - Create and maintain security, access control, and data protection settings for evaluation data, model outputs, annotations, and internal tooling. - Write clear technical documentation for Eval360 systems, error taxonomies, evaluation workflows, debugging procedures, and user-facing tools. - Work with researchers, data scientists, analysts, and machine learning engineers to improve evaluation quality, model diagnostics, and failure-mode visibility. - Keep track of new development tools, evaluation frameworks, model analysis methods, data quality techniques, and architectures relevant to AI evaluation systems. - Contribute to the design of error taxonomies, evaluation rubrics, quality thresholds, regression detection methods, and model readiness criteria. - Help ensure Eval360 produces reliable, interpretable, and actionable signals for model quality gates. - Contribute to research publications, technical reports, internal knowledge sharing, and external presentations where appropriate. - Contribute to intellectual property and thought leadership in AI evaluation, error analysis, model quality measurement, and evaluation infrastructure. - Perform all other duties as reasonably directed by the line manager that are aligned with these functional objectives. Academic Qualifications - Bachelor's degree in Computer Science, Machine Learning, Data Science, Software Engineering, Statistics, or a related technical field required. - Master's or Ph.D. in Computer Science, Machine Learning, Artificial Intelligence, Data Science, or a related field preferred. Professional Experience - Proven experience as a Software Engineer, Full Stack Developer, Machine Learning Evaluation Engineer, Data Scientist, AI Engineer, or similar role. - Experience building software systems for AI, machine learning, data analysis, evaluation, annotation, experimentation, or model monitoring. - Experience working with AI algorithms and the ability to develop systems that accommodate AI-related requirements. - Experience performing error analysis, model evaluation, data quality analysis, or failure-mode investigation for machine learning or language model systems. - Experience developing internal applications, dashboards, review tools, or web-based workflows for technical users. - Familiarity with common software stacks, including front-end frameworks, back-end services, databases, APIs, and cloud or internal infrastructure. - Familiarity with GitHub, Git, CI/CD workflows, and collaborative software development practices. - Knowledge of front-end languages and libraries such as HTML, CSS, JavaScript, TypeScript, React, Angular, or similar technologies. - Knowledge of back-end languages and frameworks such as Python, Java, C#, Node.js, FastAPI, Flask, Django, or similar technologies. - Familiarity with databases such as MySQL, PostgreSQL, MongoDB, or other structured and unstructured data stores. - Familiarity with evaluation frameworks, experiment tracking systems, data pipelines, or machine learning infrastructure is strongly preferred. - Ability to analyze complex model outputs and translate qualitative failures into structured, measurable categories. - Strong problem-solving and troubleshooting skills, especially for ambiguous technical issues involving models, data, metrics, and software systems. - Effective communication and collaboration skills, with the ability to work across research, engineering, data, and product teams. - Strong attention to detail and a high bar for evaluation quality, reliability, and interpretability. ## Preferred Qualifications - Experience with large language models, foundation models, multimodal models, or model evaluation systems. - Experience designing or using error taxonomies, evaluation rubrics, benchmark datasets, human evaluation workflows, or automated grading systems. - Experience with Python-based data analysis tools such as pandas, NumPy, Jupyter, or similar. - Experience with visualization or dashboarding tools for model quality analysis. - Experience with distributed systems, job queues, workflow orchestration, or large-scale data processing. - Experience working in a research environment or with fast-moving AI product and model teams. Visa Sponsorship This position is eligible for visa sponsorship. ## Benefits Include - Comprehensive medical, dental, and vision benefits - Bonus - 401K plan - Generous paid time off, sick leave, and holidays - Paid parental leave - Employee assistance program - Life insurance and disability insurance ## About Institute of Foundation Models ## Company Overview - **One-liner**: The Institute of Foundation Models (IFM) is a global AI research lab dedicated to the open and independent development of frontier-class foundation models, operating under the Mohamed bin Zayed University of Artificial Intelligence (MBZUAI). - **Entity Type**: Academic Research Institute (part of MBZUAI, a university) - **Headquarters**: Abu Dhabi, United Arab Emirates (with labs in Sunnyvale, CA, USA and Paris, France) - **Founded**: 2026 (launched with Silicon Valley lab) - **Founders**: Part of MBZUAI, led by President Eric Xing ## Core Business - **Primary industry**: Artificial Intelligence research, foundation model development, open-source AI - **Target customers**: Researchers, academia, industry partners, and the global AI community (B2B/Research/Public sector) - **Mission**: To produce the world’s leading models across modalities while ensuring responsible, open, and socially meaningful impact. ## Products & Services For each major offering: - **JAIS Series**: State-of-the-art open-source Arabic LLM covering Modern Standard Arabic and regional dialects, including Moroccan Arabic (Darija). Fully documented training lifecycle. - **Vicuna**: Lightweight open chatbot surpassing 90% of ChatGPT and Bard quality. - **K2 (K2-65B)**: Large language model with advanced reasoning capabilities, focusing on sustainable performance. Updated version soon to be released. - **PAN World Model**: Next-gen foundation model for embodied reasoning and physical-world simulation, integrating multimodal inputs (language, video, spatial data, physical actions). Includes PAN-Agent for multimodal reasoning tasks. - **FM For Bio – GET**: Domain-specific model for biology. - **LLM360**: Open-source framework enhancing transparency and collaboration in LLM research, providing training code, datasets, and model checkpoints. - **Research partnerships**: Collaboration with academia, startups, and enterprises via shared compute, co-publishing, and co-development. ## Market Standing - **Valuation/Market Cap**: Not applicable (non-commercial research institute) - **Key Metric**: Funded by MBZUAI; specific budget not publicly disclosed. Notable open-source model releases (JAIS, Vicuna, K2, PAN). - **Notable Investors/Partners**: MBZUAI, with international advisory boards and peer review processes. Partnerships with industry leaders, academic institutions, and public organizations. - **Growth Signals**: Launch of Silicon Valley lab in Sunnyvale, CA (2026); expansion into Paris and Abu Dhabi nodes; active hiring for research interns, distributed ML engineers, HPC engineers; release of flagship models (PAN, K2 update). ## Competitive Advantages - **Openness & transparency**: Full release of training code, datasets, and model checkpoints through LLM360 – one of the most transparent approaches in AI. - **Global research network**: Three hubs (Abu Dhabi, Silicon Valley, Paris) combining academic rigor with startup agility. - **Access to large-scale compute**: High-performance computing infrastructure for training frontier models. - **Multilingual & cultural focus**: JAIS addresses underrepresentation of Arabic and other languages, preserving cultural authenticity. - **World model innovation**: PAN differentiates from text-only models by predicting comprehensive world states for advanced reasoning and simulation. ## Strategic Focus - Continue building open, powerful foundation models across language, vision, multimodal, and domain-specific systems. - Expand global collaboration and talent acquisition, especially in Silicon Valley. - Advance responsible AI with safety systems and international advisory boards. - Drive real-world impact through partnerships in scientific discovery, human-AI interaction, and public good. ## Why Work Here - **Culture**: Collaborative, cutting-edge academic research environment with a mission to open-source AI for global benefit. Teams span Abu Dhabi, Paris, and Silicon Valley. - **Remote/hybrid/office**: All listed roles are **on-site** (Sunnyvale, CA or Abu Dhabi). No remote policy indicated. - **Notable perks**: Work with world-class researchers, access to large-scale compute, opportunity to publish and contribute to open-source models. The institute combines the agility of a startup with the resources of an established university. - **Current openings**: AI Research Internship (LLM), Distributed Machine Learning Engineer, Eval360 Error Analysis Engineer, High Performance Computing Software Engineer (Supercomputing) – all on-site in Sunnyvale or Abu Dhabi. ## Sources 1. [ifm.ai](https://ifm.ai/) 2. [ifm.ai/about](https://ifm.ai/about/) 3. [jobs.lever.co/ifm-us](https://jobs.lever.co/ifm-us) 4. [LinkedIn – Institute of Foundation Models](https://www.linkedin.com/company/institute-of-foundation-models) 5. [MBZUAI News – Launch of IFM and Silicon Valley Lab](https://mbzuai.ac.ae/news/mbzuai-launches-institute-of-foundation-models-and-establishes-silicon-valley-ai-lab/) ## Other roles at Institute of Foundation Models - [AI Engineer Internship – LLM Data](https://feeny.ai/job/ai-engineer-internship-llm-data-institute-of-foundation-models-abu-dhabi-ryr1w6wpmy2b) — Abu Dhabi, United Arab Emirates - [Research Scientist – World Modeling, Data](https://feeny.ai/job/research-scientist-world-modeling-data-institute-of-foundation-models-sunnyvale-bch8rt52tddk) — Sunnyvale, CA - [Community Development Manager](https://feeny.ai/job/community-development-manager-institute-of-foundation-models-sunnyvale-1ns2xzrzec79) — Sunnyvale, CA - [Social Media Manager](https://feeny.ai/job/social-media-manager-institute-of-foundation-models-sunnyvale-nppzja97rfm9) — Sunnyvale, CA - [Content Marketing & Editorial, Sr. Manager](https://feeny.ai/job/content-marketing-editorial-sr-manager-institute-of-foundation-models-sunnyvale-9yeba8pbf82f) — Sunnyvale, CA - [Senior Communications Consultant](https://feeny.ai/job/senior-communications-consultant-institute-of-foundation-models-sunnyvale-q3waewdeghjx) — Sunnyvale, CA - [Admin Operations Coordinator](https://feeny.ai/job/admin-operations-coordinator-institute-of-foundation-models-sunnyvale-299qtkyk8h27) — Sunnyvale, CA - [Inference Optimization Intern – Performance Modeling](https://feeny.ai/job/inference-optimization-intern-performance-modeling-institute-of-foundation-6dp7c44ge7e3) — Sunnyvale, CA - [AI Research Internship - WM](https://feeny.ai/job/ai-research-internship-wm-institute-of-foundation-models-sunnyvale-59229gv1vw1g) — Sunnyvale, CA - [Research Scientist, Agentic Data & Benchmarking](https://feeny.ai/job/research-scientist-agentic-data-benchmarking-institute-of-foundation-models-knfx9e7pzxej) — Sunnyvale, CA