--- title: 'LLM / Agentic Evaluation Rig Engineer at Phizenix' canonical: 'https://feeny.ai/job/llm-agentic-evaluation-rig-engineer-phizenix-hyderabad-x3nrtpps5f6d' type: 'job' last_seen: '2026-09-10' --- # LLM / Agentic Evaluation Rig Engineer at Phizenix - **Company:** Phizenix - **Location:** Hyderabad, India - **Work type:** hybrid - **Posted:** 2026-08-21 - **Last confirmed live:** 2026-09-10 - **Apply:** https://job-boards.greenhouse.io/phizenix/jobs/5398766008 ## Job description We are looking for an LLM / Agentic Evaluation Rig Engineer to build the system that decides whether our AI output is good enough to ship. Because our commentary sits next to externally reported financials, we cannot rely on vibes — grounding, faithfulness, and hallucination have to be measured, tracked, and gated before anything reaches a customer. You own the evaluation infrastructure: the datasets, the scorers, the harnesses, and the CI gates that hold the AI and agentic layers to a hard quality bar. You are the team's source of truth on whether a model, prompt, or agent change is actually an improvement — and the one who blocks it if it isn't. What makes this role different - You define "good enough to ship" — your gates block regressions in grounding and faithfulness from reaching production. - Evidence over vibes — every claim is checked against the verified source data it must be grounded in. - Agentic evaluation — you evaluate multi-step reasoning flows, not just single prompts. - Real leverage — your rig is how the whole AI team moves fast without breaking trust. ## Responsibilities Datasets & Scorers (35%) - Build and curate evaluation datasets, including adversarial and edge-case sets with ground-truth labels - Build scorers for grounding, faithfulness, hallucination, factual consistency, and structured-output validity - Combine rule-based checks, reference-based metrics, and LLM-as-judge where appropriate - Verify generated claims map to verified source data — no unsupported statements Harnesses & CI Gates (30%) - Build harnesses that run evaluations reproducibly across model, prompt, and agent versions - Wire evaluation into CI so grounding / faithfulness regressions block releases - Track quality over time with dashboards and clear pass / fail thresholds Agentic Evaluation (25%) - Evaluate multi-step / agentic flows — routing, tool-use, verification, confirmation - Build trace capture and step-level scoring for agent runs - Detect where a flow silently degrades Collaboration (10%) - Partner with the Staff AI Engineer to turn findings into model / prompt / orchestration improvements - Partner with QA to integrate AI evaluation into the broader release process Technical Stack Evaluation - LLM eval frameworks (promptfoo, DeepEval, Ragas, LangSmith) - LLM-as-judge, reference-based metrics - Dataset / ground-truth curation AI & Orchestration - LLM APIs & managed LLMs (Bedrock / Vertex / Azure OpenAI) - RAG & agentic patterns (LangGraph) - Structured-output validation Engineering - Python - CI/CD (GitHub Actions) - Dashboards & metrics tracking ## What You'll Build in Year One - A labeled evaluation dataset suite (including adversarial cases) for the generation and agentic layers. - A scorer library for grounding, faithfulness, hallucination, and structured-output validity. - A reproducible harness wired into CI that blocks releases on quality regressions. - Step-level trace capture and scoring for agentic flows, with dashboards leadership can trust. Required Qualifications Core - 4+ years in software / ML engineering, with hands-on work building LLM evaluation or quality tooling. - Real understanding of grounding, faithfulness, and hallucination — and how to measure them rigorously. Technical - Strong Python and solid engineering practices (reproducibility, CI/CD). - Comfort designing evaluation for non-deterministic systems without producing flaky or meaningless metrics. - Familiarity with LLM eval frameworks and LLM-as-judge patterns. Nice-to-Have - Experience evaluating agentic / multi-step LLM systems. - Familiarity with RAG, structured output, and managed LLMs in-VPC. - FinTech / financial-services domain or other high-stakes, correctness-critical AI. - Background in statistics or measurement / metrics design. ## About Phizenix ## Company Overview - **One-liner**: Phizenix provides workforce solutions for the AI era, including top tech talent placement, AI strategy & development, and AI workforce training. - **Entity Type**: Private (Pre-Seed / Bootstrapped – no public funding rounds disclosed) - **Headquarters**: Livermore, California, United States - **Founded**: 2025 - **Founders**: Khursheed Irani ## Core Business - **Primary industry**: IT Services and IT Consulting, AI Consulting, Workforce Staffing - **Target customers**: B2B – businesses seeking AI transformation, technical talent, and AI upskilling (mid-market to enterprise) - **Mission or purpose statement**: “We’re a team of tech founders, operators, and builders on a mission to help businesses thrive in an AI-native world. We believe AI should empower people, not replace them.” ## Products & Services - **[Talent Staffing]**: Executive recruiting, permanent hire, contract staffing and management for technical roles (e.g., AI/ML engineers, DevOps, firmware engineers). - **[AI Strategy & Development]**: Consulting to discover automation opportunities, design AI solutions, and build production AI systems. - **[AI Workforce Training]**: Live workshops and online certification programs to increase teams’ AI skills. ## Market Standing - **Valuation/Market Cap**: Not publicly available (private company, no funding rounds disclosed) - **Key Metric**: Not disclosed (revenue not public). Employees: 12 as of early 2026. - **Notable Investors/Partners**: No named investors; leadership includes Khursheed Irani (CEO), Raashid Mehasanewala (Head of Solutions), Jumana Ghadiali (VP Operations), Sanjay Dayal (Senior Technology Advisor), Aziz Ghadiali (Head of AI and Marketing). - **Growth Signals**: Headcount growth of +7.7% monthly (LinkedIn estimate); active job postings on Greenhouse (3 open roles as of 2026); LinkedIn follower count 3,587 with +3.9% monthly growth. ## Competitive Advantages - **AI-native focus**: Entire business model centered on AI transformation, not a general staffing firm. - **Silicon Valley presence**: HQ in Livermore, CA, with additional presence in India and Canada, allowing access to diverse talent pools. - **Experienced leadership team**: Founders and VPs bring 20+ years of combined experience in tech, operations, and AI. ## Strategic Focus - **Current priorities**: Expand AI consulting engagements, scale talent staffing pipeline, and grow the training certification arm. Emphasis on “AI-native future of work” – helping companies integrate AI alongside human teams rather than replacing them. ## Why Work Here - **Culture**: Small, mission-driven team focused on AI empowerment. Emphasis on “AI should empower people, not replace them.” - **Remote/Hybrid/Office**: Many roles are remote; some are hybrid in San Francisco/Bay Area (e.g., Director of Engineering – Shopify & ERP). Offices in Livermore, CA and Bengaluru, India. - **Notable perks/engineering culture**: Opportunity to work at the intersection of AI and business transformation; hands-on involvement in building AI solutions for clients; growth potential in a young, rapidly expanding company. ## Sources 1. [phizenix.com](https://www.phizenix.com/) 2. [phizenix.com/careers](https://www.phizenix.com/careers) 3. [phizenix.com/about](https://www.phizenix.com/about) 4. [LinkedIn/Phizenix](https://www.linkedin.com/company/phizenix) 5. [Greenhouse/Phizenix](https://job-boards.greenhouse.io/phizenix/jobs/5197144008) ## Other roles at Phizenix - [Sr. Staff Backend Engineer](https://feeny.ai/job/sr-staff-backend-engineer-phizenix-hyderabad-psrdjq79ha1t) — Hyderabad, India - [QA Automation Engineer](https://feeny.ai/job/qa-automation-engineer-phizenix-hyderabad-emn3ryp3sk2z) — Hyderabad, India - [Product Manager- Workflows](https://feeny.ai/job/product-manager-workflows-phizenix-remote-kp0tpcrhcbpd) - [Senior Analog & Mixed-Signal IC Designer](https://feeny.ai/job/senior-analog-mixed-signal-ic-designer-phizenix-santa-clara-22hs598ryxs0) — Santa Clara, CA / Irvine, CA - [ASIC Physical Design Engineer](https://feeny.ai/job/asic-physical-design-engineer-phizenix-bengaluru-sd4fja8bp25t) — Bengaluru, India - [Digital Design Lead](https://feeny.ai/job/digital-design-lead-phizenix-santa-clara-gb92mzn0r7vw) — Santa Clara, CA - [Senior Analog/RF Layout Engineer](https://feeny.ai/job/senior-analog-rf-layout-engineer-phizenix-bengaluru-gen055rfqf07) — Bengaluru, India - [Senior Analog & Mixed-Signal IC Designer](https://feeny.ai/job/senior-analog-mixed-signal-ic-designer-phizenix-bengaluru-gzxnw70c7jwj) — Bengaluru, India - [Analog & Mixed-Signal IC Design Lead Engineer](https://feeny.ai/job/analog-mixed-signal-ic-design-lead-engineer-phizenix-bengaluru-hpvp5jt5b4qa) — Bengaluru, India - [Future Opportunities at Phizenix – Let's Stay Connected!](https://feeny.ai/job/future-opportunities-at-phizenix-let-s-stay-connected-phizenix-remote-ymyeg9621244)