--- title: 'Research Engineer - Evals at AGI, Inc.' canonical: 'https://feeny.ai/job/research-engineer-evals-agi-inc-san-francisco-h3fm9b42ytfx' type: 'job' last_seen: '2026-09-10' --- # Research Engineer - Evals at AGI, Inc. - **Company:** AGI, Inc. - **Location:** San Francisco, CA - **Employment:** full-time - **Posted:** 2026-05-27 - **Last confirmed live:** 2026-09-10 - **Apply:** https://jobs.ashbyhq.com/agi-inc/d9465f93-b206-4d76-b736-02b3c639edde ## Job description Think Different. Build the Future. 🚀 ## Our Mission Build everyday AGI. Trustworthy, consumer-grade agents that redefine human–AI collaboration for millions. Software shouldn’t wait for commands; it should partner with you, amplifying what you can do every single day. Why AGI, Inc. We’re a stealth team of elite founders and AI researchers, with backgrounds spanning Stanford, OpenAI, and DeepMind. We’re industry leaders in mobile and computer-use agents, bringing these capabilities to consumer scale. Grounded in years of agent research, our AI is designed with trustworthiness and reliability as core pillars, not afterthoughts. We are supported by tier-1 investors who funded the first generation of AI giants; now they’re backing us to build the next: everyday AGI. (Watch the [demo](https://drive.google.com/file/d/1ZydjdMeMh3x-QItUPQFbJUbhBW-4-XHa/view?usp=sharing)) If you see possibility where others see limits, read on. You decide what "better" means. Models, agents, and product features all ship behind one question: did this actually get better? Without a strong evals function, the lab ships vibes. With one, every training run, every prompt change, every agent capability moves a number we trust — and the team makes decisions on real signal, not the loudest opinion in the room. You'll build the eval harness for AGI — across model capability, agentic behavior, on-device performance, and end-user experience. You'll set the bar for what counts as "shipped" and protect it from the gravity of product deadlines. 🤩 Tasks you will own - The eval suites that gate every model and agent release — capability, behavior, regressions, and human-rated rubrics that catch what automated evals miss - The dashboards and tooling that make researcher experiment loops fast and leadership decisions easy - The bar — what counts as ready to ship, and how we know 🤚 Areas where you will assist - Research, by making sure what we measure is what we want - Product engineers, by instrumenting real-user behavior on real devices - Partnerships, by translating "did it get better" into language an OEM partner can hold us to 📚 Skills you'll be expected to teach - How to measure non-deterministic systems — agent eval, tool use, long-horizon tasks, multilingual behavior - How to push back on a metric that's being gamed without breaking the team 🧑‍🎓 Skills you'll be expected to learn - On-device perf trade-offs and how they show up in real-user evals - What QA-ing AI at OEM scale actually looks like - The realities of shipping consumer agents to production partners 🏆 Timeline of success After 30 days — You've audited every eval we run today and produced a sharp doc on what's good, what's noise, and what's missing. You've fixed the most embarrassing gap. After 60 days — You've stood up a new eval surface — agentic, on-device, or behavioral — and the team is making real decisions on its output. Researchers come to you before launching a run, not after. After 90 days — Releases now ship against your eval bar, not a vibe-check. You've caught a regression that would have shipped, and cleared a launch the team was nervous about. You're shaping the research roadmap by surfacing where we're flat, where we're climbing, and where we're lying to ourselves. 💰 Compensation Competitive cash and meaningful equity. Top-tier relocation and immigration support. SF, in person. How to apply Send a link to an eval, benchmark, or measurement system you built — and one paragraph on what decision it changed. Plus your resume or LinkedIn. Every exceptional candidate hears back within 48 hours. ## About AGI, Inc. ## Company Overview - **One-liner**: AGI, Inc. is an applied AI lab bringing superintelligence to the edge by building fully agentic, on-device AI that runs locally without cloud dependencies. - **Entity Type**: Private (Seed stage – $20M raised) - **Headquarters**: San Francisco, California, United States (also offices in India, Japan, Brazil) - **Founded**: 2025 - **Founders**: Div Garg (CEO; co-founder background inferred from leadership), Steve Frey (Cofounder, Product) ## Core Business - **Primary industry/industries**: Agentic AI / On-device Artificial Intelligence / Edge Computing - **Target customers**: B2C (individuals using mobile devices) and B2B (partnerships with hardware makers like Qualcomm, Lenovo, Visa for agentic commerce) - **Mission/purpose statement**: “Bring genuinely useful AGI into everyday life” – making superintelligence 100% secure, private, and accessible on the devices people already own. ## Products & Services - **[AGI-0]**: A mobile-use agent that proactively handles tasks (booking taxis, ordering food, replying to messages, finding flights) by operating a user’s apps locally and personally. Fully agentic, no cloud round-trips. - **[On-Device Foundation Models]**: AGI, Inc. runs its own foundation models on-device, enabling agents that act rather than just answer. Up to 96% faster processing-near-memory without hardware changes. - **[Research & Developer Tools]**: Offers a blog, deeplearning.ai courses, and an evaluation framework for web AI agents. ## Market Standing - **Valuation/Market Cap**: Not disclosed (private) - **Key Metric**: Total funding $20M (Seed round, September 2025, led by 11 investors including Techstars, individual angels) - **Notable Investors/Partners**: Qualcomm Technologies (collaboration to bring agentic AI to Snapdragon-powered devices), Visa (transforming agentic commerce), Lenovo (proof-of-concept integration demonstrated at MWC), Anand V Lalwani (angel) - **Growth Signals**: 26 employees, monthly headcount growth +4.3%, LinkedIn followers 22,530 (+3.5% monthly), 6 active job postings (+50% quarterly). Demonstrated prototype at MWC (March 2026) and a research preview released in October 2025. Monthly website visits ~38,700. ## Competitive Advantages - **True On-Device AI**: No cloud round-trips – all data stays on the device, ensuring privacy, security, and low latency. Differentiates from cloud-dependent assistants. - **Full Agentic Capability**: Agents act proactively on the user’s behalf, not just answer queries. Integrated with phone apps similarly to a human user. - **Hardware Partnerships**: Early collaborations with Qualcomm (Snapdragon), Lenovo, and Visa give route-to-market and credibility. - **Processing-Near-Memory Innovation**: Claims 96% faster processing without hardware changes, a significant efficiency moat. ## Strategic Focus - **Current priorities**: Scaling the on-device agent platform, deepening hardware partnerships (Qualcomm, Lenovo), expanding agentic commerce (Visa), and hiring top AI/engineering talent. Growing presence in US, India, Japan, Brazil. - **Direction for growth**: Deploy “superintelligence in your hands” to billions of devices – moving AI from data centers to pockets, driveways, and living rooms. ## Why Work Here - **Culture**: Described as a place for “passionate builders, innovators, and dreamers”. High-impact mission to bring AGI to everyday life. Employee rating on LinkedIn 5.0/5.0 (1 review) – strong marks for work-life, compensation, culture, career. - **Remote/hybrid/office**: Presence in San Francisco (HQ plus two other office locations), with additional team members in India, Japan, Brazil. Likely a flexible or hybrid model given global distribution. - **Notable perks/engineering culture**: Flat structure with 26 employees; department breakdown shows technical roles (12%) and product (6%) plus research. Open positions: AI Researcher, ML Platform Engineer, iOS Engineer, Backend Engineer, Product Designer – indicating strong engineering and research focus. - **Growth trajectory**: Early-stage startup with $20M funding and rapid hiring (+50% job postings quarterly). Opportunity to shape foundational AI product from early days. ## Sources 1. [theagi.company](https://theagi.company/) 2. [linkedin.com](https://www.linkedin.com/company/the-agi-company) 3. [theagi.company/build](https://www.theagi.company/build) 4. [jobs.ashbyhq.com/agi-inc/792da3fd-c1ae-4e06-9398-f40092f1d612](https://jobs.ashbyhq.com/agi-inc/792da3fd-c1ae-4e06-9398-f40092f1d612) 5. [jobs.ashbyhq.com/agi-inc/442fee6d-2c44-4a2d-9daa-d209a22a14d5](https://jobs.ashbyhq.com/agi-inc/442fee6d-2c44-4a2d-9daa-d209a22a14d5) ## Other roles at AGI, Inc. - [iOS Engineer](https://feeny.ai/job/ios-engineer-agi-inc-san-francisco-t1e8e58hm2yc) — San Francisco, CA - [AI Researcher](https://feeny.ai/job/ai-researcher-agi-inc-san-francisco-3yc662yahsp8) — San Francisco, CA - [AI Engineer - Backend](https://feeny.ai/job/ai-engineer-backend-agi-inc-san-francisco-v6qp81n5hys5) — San Francisco, CA - [Product Designer](https://feeny.ai/job/product-designer-agi-inc-san-francisco-qfhz9cmpca8z) — San Francisco, CA - [AI Product FDE](https://feeny.ai/job/ai-product-fde-agi-inc-san-francisco-q8zs4r41qxgx) — San Francisco, CA - [Research Engineer - Evals](https://feeny.ai/job/research-engineer-evals-firecrawl-san-francisco-4dhwy2dz84ay) — San Francisco, CA / Toronto, Canada