--- title: 'ML Ops Engineer — Agentic AI Lab (Founding Team) at Fabrion' canonical: 'https://feeny.ai/job/ml-ops-engineer-agentic-ai-lab-founding-team-fabrion-san-francisco-na47hpj9nvs7' type: 'job' last_seen: '2026-09-15' --- # ML Ops Engineer — Agentic AI Lab (Founding Team) at Fabrion - **Company:** Fabrion - **Location:** San Francisco, CA - **Employment:** full-time - **Work type:** onsite - **Posted:** 2025-08-11 - **Last confirmed live:** 2026-09-15 - **Apply:** https://jobs.ashbyhq.com/fabrion/277bb01a-80b0-4ee3-b0ef-20287174f597 ## Job description ML Ops Engineer — Agentic AI Lab (Founding Team) Location: San Francisco Bay Area Type: Full-Time Compensation: Competitive salary + meaningful equity (founding tier) Backed by 8VC, we're building a world-class team to tackle one of the industry’s most critical infrastructure problems. ## About the Role Our AI Lab is pioneering the future of intelligent infrastructure through open-source LLMs, agent-native pipelines, retrieval-augmented generation (RAG), and knowledge-graph-grounded models. We’re hiring an ML Ops Engineer to be the glue between ML research and production systems — responsible for automating the model training, deployment, versioning, and observability pipelines that power our agents and AI data fabric. You’ll work across compute orchestration, GPU infrastructure, fine-tuned model lifecycle management, model governance, and security e ## Responsibilities - Build and maintain secure, scalable, and automated pipelines for: - LLM fine-tuning, SFT, LoRA, RLHF, DPO training - RAG embedding pipelines with dynamic updates - Model conversion, quantization, and inference rollout - Manage hybrid compute infrastructure (cloud, on-prem, GPU clusters) for training and inference workloads using Kubernetes, Ray, and Terraform - Containerize models and agents using Docker, with reproducible builds and CI/CD via GitHub Actions or ArgoCD - Implement and enforce model governance: versioning, metadata, lineage, reproducibility, and evaluation capture - Create and manage evaluation and benchmarking frameworks (e.g. OpenLLM-Evals, RAGAS, LangSmith) - Integrate with security and access control layers (OPA, ABAC, Keycloak) to enforce model policies per tenant - Instrument observability for model latency, token usage, performance metrics, error tracing, and drift detection - Support deployment of agentic apps with LangGraph, LangChain, and custom inference backends (e.g. vLLM, TGI, Triton) Desired Experience Model Infrastructure: - 4+ years in MLOps, ML platform engineering, or infra-focused ML roles - Deep familiarity with model lifecycle management tools: MLflow, Weights & Biases, DVC, - HuggingFace Hub - Experience with large model deployments (open-source LLMs preferred): LLaMA, - Mistral, Falcon, Mixtral - Comfortable with tuning libraries (HuggingFace Trainer, DeepSpeed, FSDP, QLoRA) - Familiarity with inference serving: vLLM, TGI, Ray Serve, Triton Inference Server Automation + Infra: - Proficient with Terraform, Helm, K8s, and container orchestration - Experience with CI/CD for ML (e.g. GitHub Actions + model checkpoints) - Managed hybrid workloads across GPU cloud (Lambda, Modal, HuggingFace Inference, - Sagemaker) - Familiar with cost optimization (spot instance scaling, batch prioritization, model sharding) Agent + Data Pipeline Support:● Familiarity with LangChain, LangGraph, LlamaIndex or similar RAG/agent orchestration tools Built embedding pipelines for multi-source documents (PDF, JSON, CSV, HTML) Integrated with vector databases (Weaviate, Qdrant, FAISS, Chroma) Security & Governance: Implemented model-level RBAC, usage tracking, audit trails Integrated with API rate limits, tenant billing, and SLA observability Experience with policy-as-code systems (OPA, Rego) and access layers ## Preferred Stack - LLM Ops: HuggingFace, DeepSpeed, MLflow, Weights & Biases, DVC - Infra: Kubernetes (GKE/EKS), Ray, Terraform, Helm, GitHub Actions, ArgoCD - Serving: vLLM, TGI, Triton, Ray Serve - Pipelines: Prefect, Airflow, Dagster - Monitoring: Prometheus, Grafana, OpenTelemetry, LangSmith - Security: OPA (Rego), Keycloak, Vault - Languages: Python (primary), Bash, optionally Rust or Go for tooling Mindset & Culture Fit - Builder's mindset with startup autonomy: you automate what slows you down - Obsessive about reproducibility, observability, and traceability - Comfortable with a hybrid team of AI researchers, DevOps, and backend engineers - Interested in aligning ML systems to product delivery, not just papers - Bonus: experience with SOC2, HIPAA, or GovCloud-grade model operations ## What We’re Looking For Experience: - 5+ years as a full stack or backend engineer - Experience owning and delivering production systems end-to-end - Prior experience with modern frontend frameworks (React, Next.js) - Familiarity with building APIs, databases, cloud infrastructure, or deployment workflows at scale - Comfortable working in early-stage startups or autonomous roles, prior experience as a founder, founding engineer, or a 0-1 pre-seed startup is a big plus Mindset: - Comfortable with ambiguity, eager to prototype and iterate quickly - Strong sense of ownership — prefers to build systems rather than wait for tickets - Enjoys thinking about architecture, performance, and tradeoffs at every level - Clear communicator and pragmatic team player - Values equity and impact over prestige or hierarchy - Prior startup or founding team experience ## Why This Role Matters Your work will enable models and agents to be trained, evaluated, deployed, and governed at scale — across many tenants, models, and tasks. This is the backbone of a secure, reliable, and scalable AI-native enterprise system. If you dream about using AI to solve some really hard real world problems – we would love to hear from you. ## About Fabrion ## Company Overview - **One-liner**: Fabrion is an AI-native platform and operating system that provides a real-time intelligence layer for industrial manufacturers, transforming complex value chains with agentic AI, knowledge graphs, and data fabrics. - **Entity Type**: Private – Seed stage (Seed round with one investor) - **Headquarters**: San Francisco, California, United States - **Founded**: 2025 - **Founders**: Roy Ng (CEO), Kunal B. (CTO & CPO), Christian V. Jordan, Jake Medwell (Board Member) – experienced operators from AWS, Meta, SAP, and Twilio. ## Core Business - **Primary industries**: Industrial manufacturing, supply chain, enterprise AI - **Target customers**: B2B – Enterprise industrial manufacturers, OEMs, and tier suppliers - **Mission/purpose**: “AI will be the operating system for the next industrial era” – building autonomous, governed AI systems that let teams see, decide, and act on complex manufacturing and supply chain data ## Products & Services - **Fabrion Platform**: An AI-native operating system that ingests, normalizes, and acts on messy, fragmented industrial data. Uses LLMs, knowledge graphs, vector databases, and agentic systems to deliver real-time intelligence. - **Fabrion AI Lab**: A full-stack vertical AI research platform focused on: - Data fabric for AI (metadata-driven cleanse/correlate/stream) - Governance-by-design (observable, reversible, compliant AI) - Context memory management (short/long-term recall for agents) - Industry knowledge graphs (internal data + external signals) - Agent mesh architectures (distributed, goal-conditioned agents) - Hybrid & multi-cloud infrastructure - Custom fine-tuned models (RAG, SLMs, RLHIL, pre-trained enterprise data) ## Market Standing - **Valuation**: Not disclosed - **Key metric**: Seed funding round (amount not publicly stated); 1 investor identified (8VC is prominently featured as a partner) - **Notable investors/partners**: 8VC (lead through “8VC Build” program); partnering with leading OEMs and forward-thinking industrial companies - **Growth signals**: - Workforce of ~7 employees (all founding/early team) with **9 open job postings** across engineering, research, and business development - Web traffic growth: +115.2% monthly (4,264 monthly visits) - LinkedIn followers: 343 (+7.5% monthly) - Founding team has scaled products from zero to hundreds of millions in revenue at previous companies ## Competitive Advantages - **AI-native from day one** – not a thin wrapper on foundation models; built as a full-stack vertical AI research platform - **Deep industry partnerships** with OEMs and industrial companies actively modernizing - **Proven founding team** with experience at AWS, Meta, SAP, and Twilio – domain expertise in manufacturing, logistics, Big Data, SaaS, and enterprise - **Focus on governance and compliance** – every AI action is explainable, reversible, and measurable, a critical differentiator for regulated industrial environments - **Knowledge graph + agent mesh approach** – moves beyond static dashboards to autonomous, semi-autonomous agents ## Strategic Focus - Build autonomous or semi-autonomous AI agents for manufacturing supply chains (self-healing, rebalancing) - Embed compliance and governance as first-class citizens in the AI stack - Advance research in data fabrics, context memory, and custom fine-tuned models for industrial verticals - Expand agentic capabilities from simulation to real-world production environments ## Why Work Here - **Culture**: “High EQ, low ego, no assholes”; owners not renters; first-principles thinking; win with customer outcomes; score points, not yardage - **Work modality**: Hybrid – jobs are listed as “In-Office or Remote” with multiple locations (HQ at Pier 5, The Embarcadero, San Francisco; also remote options) - **Technical stack**: Kubernetes, Keycloak, Python, SQL, Docker, GraphQL, React, TypeScript, Terraform, Snowflake, Neo4j, OpenAI, MLflow, Grafana, Prometheus – cutting-edge AI/ML infrastructure - **Impact**: Solving mission-critical problems in global trade and manufacturing – value chains representing billions of dollars; opportunity to define a new industrial AI paradigm - **Team**: Small founding team with deep experience; opportunity to work across engineering, research, and product from the ground up ## Sources 1. [fabrion.com/careers](https://www.fabrion.com/careers) 2. [builtin.com](https://builtin.com/company/fabrion) 3. [jobs.ashbyhq.com](https://jobs.ashbyhq.com/fabrion) 4. [linkedin.com](https://www.linkedin.com/company/fabrionai) 5. [fabrion.com/ai-lab](https://www.fabrion.com/ai-lab) ## Other roles at Fabrion - [Founding AI Research Lead - Agentic AI Lab](https://feeny.ai/job/founding-ai-research-lead-agentic-ai-lab-fabrion-san-francisco-c251tdtjjvrw) — San Francisco, CA - [Founding Designer](https://feeny.ai/job/founding-designer-fabrion-san-francisco-zww16cz5wbwb) — San Francisco, CA - [Data Partnerships Analyst](https://feeny.ai/job/data-partnerships-analyst-fabrion-san-francisco-2h67mf9ssa3z) — San Francisco, CA - [Business Development Analyst – Auto](https://feeny.ai/job/business-development-analyst-auto-fabrion-san-francisco-4agd1wrm1x3z) — San Francisco, CA - [ML/AI Research Engineer — Agentic AI Lab (Founding Team)](https://feeny.ai/job/ml-ai-research-engineer-agentic-ai-lab-founding-team-fabrion-san-francisco-m7ry3fmyn2rd) — San Francisco, CA - [DevOps Engineer (Founding Team)](https://feeny.ai/job/devops-engineer-founding-team-fabrion-san-francisco-rhgsyhr9s3kx) — San Francisco, CA - [Frontend Engineer (Founding Team)](https://feeny.ai/job/frontend-engineer-founding-team-fabrion-san-francisco-p4nbwzbq1kgq) — San Francisco, CA - [Data Engineer (Founding Team)](https://feeny.ai/job/data-engineer-founding-team-fabrion-san-francisco-33hf28h6rgc7) — San Francisco, CA - [Founding Engineer](https://feeny.ai/job/founding-engineer-fabrion-san-francisco-50vz9w20z840) — San Francisco, CA