--- title: 'LLM Red Team Intern (Evaluation Systems) at Elloe AI' canonical: 'https://feeny.ai/job/llm-red-team-intern-evaluation-systems-elloe-ai-austin-q8t54b7ywt5w' type: 'job' last_seen: '2026-09-10' --- # LLM Red Team Intern (Evaluation Systems) at Elloe AI - **Company:** Elloe AI - **Location:** Austin, TX - **Employment:** internship - **Work type:** remote - **Posted:** 2025-07-11 - **Last confirmed live:** 2026-09-10 - **Apply:** https://jobs.gem.com/elloe-ai/am9icG9zdDr5AR3gDAm7kAvWtCaXeGWv ## Job description Internship | Remote | LLM Evaluation | Reports to CTO or Safety Lead ## About Elloe Elloe is the immune system for AI. We don’t train models — we protect their outputs. We trace every hallucination, enforce every policy boundary, and create an audit trail for every critical LLM interaction. Our modules (TruthChecker™, AutoRAG™, Autopsy™) are embedded in hospitals, banks, and regulatory sandboxes. Our job is to make sure these systems are safe before anything hits production. This role will help us break, stress-test, and harden the models used by governments and enterprises alike. ## About the Role You’ll red team real-world LLM deployments, design eval harnesses, and help scale Elloe’s output-level safety layer. This isn’t just prompt tuning — it’s forensic risk mapping. You’ll work directly with product and safety leads to uncover failure patterns and codify guardrails for GenAI systems under real-world scrutiny. ## What You’ll Own 1. Red Teaming & Risk Testing - Create prompts to trigger hallucinations, policy violations, or failure scenarios - Stress test Elloe-protected deployments using open and proprietary models - Document behavioral exploits across use cases (healthcare, compliance, gov) 1. Evaluation Design - Build truthsets and scoring rubrics tied to factuality, policy, or ethical standards - Benchmark Elloe’s modules across model types (Claude, GPT-4, Gemini, open models) - Collaborate with product to refine and expand our eval harnesses 1. Safety Intelligence - Identify blind spots in current detection logic - Recommend scoring methods or red flag thresholds for deployment - Support internal model comparison reports or customer safety audits ## Who You Are - ML/AI researcher or engineer (undergrad, grad, or early career) - Experience working with LLMs, eval sets, and prompt design - Strong attention to detail, grounded in safety and adversarial thinking - Bonus: exposure to safety benchmarks like TruthfulQA, MMLU, or red teaming tools ## Why This Matters This is real-world alignment, not research theater. You’ll be helping define how AI gets deployed responsibly — with traceability, transparency, and real-time protection. You’ll leave this role with: - Exposure to high-stakes LLM safety deployments - Published frameworks or scoring methods used by enterprises - Mentorship from technical founders operating at the bleeding edge of AI safety Logistics & Application - Start Date: Rolling - Duration: 12–16 weeks - Compensation: Research stipend - Location: Remote-first; flexible for global candidates - To Apply: Share a jailbreak or eval idea you’d love to run against GPT-4. ## About Elloe AI ## Company Overview - **One-liner**: Elloe AI provides a real-time compliance layer that acts as an “immune system” for AI, making large language models safe to deploy across regulated industries. - **Entity Type**: Private (Pre-Seed, $1M total funding) - **Headquarters**: Austin, Texas, United States - **Founded**: 2022 (Pre-Seed round closed April 2022) - **Founders**: Owen Sakawa (CEO & Co-Founder), Jackson Mwaniki (CTO & Co-Founder), Aaron Madolora (CPO & Co-Founder) ## Core Business - **Primary industry**: AI Safety & Governance, Enterprise Compliance, Cybersecurity - **Target customers**: Governments, hospitals, financial institutions, and other regulated enterprises (B2B) - **Mission statement**: “Build the immune system for AI—systems that make AI safe, truthful, and compliant in the real world.” ## Products & Services - **Elloe AI Platform**: An API/SDK that sits on top of an LLM’s output layer, providing three “anchors”: 1. **Fact-checking anchor** – verifies responses against trusted sources. 2. **Regulatory compliance anchor** – checks for violations of HIPAA, GDPR, PII exposure, and other regulations. 3. **Audit trail anchor** – logs every decision, source, and confidence score for full traceability. The platform is not built on another LLM, but uses machine learning and human-in-the-loop oversight. ## Market Standing - **Valuation/Market Cap**: Not publicly available - **Key Metric**: Annual Revenue ~$870K (company-reported); Total Funding $1M (Pre-Seed, led by Mad Ventures in April 2022) - **Notable Investors/Partners**: Mad Ventures (sole lead investor in Pre-Seed round); also acquired Saada (2022) - **Growth Signals**: Finalist (Top 20) in TechCrunch Disrupt 2025 Startup Battlefield; headcount growth trending negative (-16.7% YoY) but still hiring for multiple roles as of late 2025; offices opened in Nairobi and London; claims to have prevented 100,000+ medical errors, blocked $20M+ in biased loans, and protected 12M+ AI-driven risks. ## Competitive Advantages - **Not an LLM judging LLMs**: Elloe AI’s system uses its own ML models and human analysts, avoiding the “Band-Aid on a wound” problem of having one LLM check another. - **Multi-layer anchor system**: Fact-checking, regulatory compliance, and full audit trail are built-in, not bolted on. - **Regulatory expertise**: The company keeps up with evolving global regulations (HIPAA, GDPR, etc.) and employs humans-in-the-loop to stay current. - **Early trust in regulated verticals**: Already deployed by governments, hospitals, and enterprises, giving them domain-specific credibility. ## Strategic Focus - **Deepen compliance capabilities**: Expand coverage of regulatory frameworks and jurisdictions. - **Global expansion**: Offices in Austin, Nairobi, and London suggest a push into North America, Africa, and Europe. - **Product development**: Building out “anchors” for more use cases, and possibly edge inference for GPU environments (as seen in job postings for Principal Engineer – Distributed Systems). - **Talent acquisition**: Actively hiring for roles in engineering, design, product, policy, and communications – indicating a growth phase focused on scaling the team despite recent headcount dip. ## Why Work Here - **Mission-driven culture**: The team describes itself as “mission-first, values-led,” focused on making AI safe and trustworthy in critical domains. - **Small, global team**: With ~10 employees across three continents, new hires can have outsized impact and direct exposure to founders. - **Flexible work options**: Job postings list “In-Office or Remote” for most roles, with the Austin HQ operating on-site; typical time on-site is listed as “OnSite Workspace” for HQ, but remote arrangements appear available. - **High-profile visibility**: Being a Top 20 finalist at TechCrunch Disrupt 2025 gives the company industry recognition and a startup event pedigree. - **Meaningful impact**: The company’s own metrics (100K+ medical errors prevented, $20M+ biased loans blocked) suggest employees contribute to directly measurable social good. - **Diverse, distributed team**: Team members come from backgrounds including Cerner, SAP, Freshworks, Kenya Power, and others, indicating a mix of enterprise and startup experience. ## Sources 1. [TechCrunch](https://techcrunch.com/2025/10/28/elloe-ai-wants-to-be-the-immune-system-for-ai-check-it-out-at-disrupt-2025/) 2. [LinkedIn](https://linkedin.com/company/elloe-inc) 3. [Elloe AI – Who We Are](https://www.elloe.ai/who-we-are) 4. [Built In – Elloe AI Careers](https://builtin.com/company/elloe-ai) 5. [Gem – Elloe AI Careers](https://jobs.gem.com/elloe-ai) ## Other roles at Elloe AI - [Member of Technical Staff – AGI Governance Researcher (Constitutional Safety)](https://feeny.ai/job/member-of-technical-staff-agi-governance-researcher-constitutional-safety-elloe-e2zwy76efft6) — San Francisco, CA - [Principal Engineer – Distributed Systems (GPU Edge + Inference)](https://feeny.ai/job/principal-engineer-distributed-systems-gpu-edge-inference-elloe-ai-austin-1ebxd0k753ff) — Austin, TX - [Product Manager – Enforcement Infrastructure (LLM + Vault)](https://feeny.ai/job/product-manager-enforcement-infrastructure-llm-vault-elloe-ai-austin-sz8p3dj7n48d) — Austin, TX - [Member of Technical Staff – ML Security Engineer (Red Teaming + Patch Risk Forecasting)](https://feeny.ai/job/member-of-technical-staff-ml-security-engineer-red-teaming-patch-risk-kfqrxs0gn25j) — Austin, TX - [Technical Product Intern (Compliance x LLM Ops)](https://feeny.ai/job/technical-product-intern-compliance-x-llm-ops-elloe-ai-austin-c254wg70xaws) — Austin, TX - [MBA Intern – Strategic Growth & GTM (Founder Office)](https://feeny.ai/job/mba-intern-strategic-growth-gtm-founder-office-elloe-ai-austin-genfb4sn5e6x) — Austin, TX - [Member of Technical Staff – Explainability Engineer (Alignment + SHAP)](https://feeny.ai/job/member-of-technical-staff-explainability-engineer-alignment-shap-elloe-ai-austin-qfm6yx8k24r7) — Austin, TX - [Design Intern (Narrative + Interface Systems)](https://feeny.ai/job/design-intern-narrative-interface-systems-elloe-ai-austin-pa5m274e6mn9) — Austin, TX - [Sales Ops & Revenue Intern (Pilot GTM x Signal Ops)](https://feeny.ai/job/sales-ops-revenue-intern-pilot-gtm-x-signal-ops-elloe-ai-austin-x1b53xdajp4q) — Austin, TX - [Communications Intern (AI Policy + GTM)](https://feeny.ai/job/communications-intern-ai-policy-gtm-elloe-ai-austin-h6r7bdzaqtb8) — Austin, TX