--- title: 'Solutions Architect - AI Model Specialist at FriendliAI' canonical: 'https://feeny.ai/job/solutions-architect-ai-model-specialist-friendliai-san-francisco-raxb34mrn257' type: 'job' last_seen: '2026-09-12' --- # Solutions Architect - AI Model Specialist at FriendliAI - **Company:** FriendliAI - **Location:** San Francisco, CA - **Employment:** full-time - **Work type:** hybrid - **Posted:** 2026-03-15 - **Last confirmed live:** 2026-09-12 - **Apply:** https://jobs.ashbyhq.com/friendliai/b131cb30-8386-407e-a984-a75d5f2845da ## Job description ## About the job FriendliAI is seeking a Solution Architect specializing in open-source AI models, AI inference API integration, and agentic systems. You will work closely with our customers to integrate FriendliAI’s inference and agent frameworks into real-world products, enabling them to build and scale AI applications effectively. You will work directly on our customers’ projects, collaborating with their engineering teams to solve challenges in integrating tools, environments, and models with AI agents. This is a hands-on, customer-embedded role. ## Key Responsibilities - Design and implement AI-powered products using FriendliAI’s APIs - Guide customers on selecting, evaluating, and operating AI models across different domains - Integrate and extend open-source frameworks for FriendliAI integration - Build and deploy custom inference endpoints, chat flows, and multi-agent orchestration pipelines - Develop SDKs, example applications, and reference APIs for agentic and generative AI use cases - Provide deep technical guidance on prompt engineering, API composition, and workflow orchestration - Debug and optimize context and memory across long-running agent sessions - Gather customer feedback and translate it into product-level improvements - Lead technical demos, developer workshops, or webinars ## Qualifications - 3+ years of software engineering experience, ideally in backend or API development - Proficient in Python and modern web frameworks (FastAPI, Flask, or similar) - Strong experience deploying LLMs and integrating into generative AI APIs - Familiarity with agentic AI frameworks (LangChain, CrewAI, AutoGen, etc.) - Strong experience in integrating open-source generative AI models into applications - Excellent communication skills and a passion for improving developer experience - Excellent problem-solving and debugging skills in real-world environments ## Preferred Experience - Contributions to open-source AI libraries or projects - Familiarity with multi-agent orchestration, memory systems, RAG, and workflow DAGs - Experience with serverless backends or API gateways ## Benefits - A front-row seat to the generative AI infrastructure revolution - Competitive compensation and benefits package - Daily lunch and dinner provided; unlimited snacks and beverages - Health check-up and top-tier hardware support - Flexible working hours and a highly collaborative environment ## About us FriendliAI is building the next-generation AI inference platform that accelerates the deployment of large language and multimodal models with unmatched performance and efficiency. Our infrastructure powers high-throughput, low-latency workloads for global organizations and integrates directly with Hugging Face, providing instant access to over 600,000 open-source models. We are on a mission to deliver the world’s best platform for AI inference. ## About FriendliAI ## Company Overview - **One-liner**: FriendliAI is The Frontier AI Inference Cloud, providing a highly optimized platform for deploying, scaling, and monitoring large language and multimodal models with unmatched speed and cost efficiency. - **Entity Type**: Private (Seed stage; $26.7M total funding) - **Headquarters**: San Francisco, California, United States (with offices in Seoul, South Korea and Redwood City, CA) - **Founded**: 2021 - **Founders**: Byung-Gon Chun (CEO) and Gyeong-In Yu (CTO) – the researchers who invented the continuous batching technique now standard in AI inference. ## Core Business - **Primary industry**: AI Infrastructure / Cloud Inference Platform (Infrastructure as a Service) - **Target customers**: B2B – enterprises and AI teams deploying generative AI and agent workloads at production scale; focuses on open-weight and custom models. - **Mission**: Democratize access to production-grade AI so teams can focus on building great products. ## Products & Services - **[FriendliAI Inference Cloud](https://friendli.ai/)**: A fully managed platform that instantly deploys 580,000+ Hugging Face models (language, audio, vision) with one click. Supports bring-your-own fine-tuned or proprietary models. Key features: - Custom GPU kernels, smart caching, continuous batching, speculative decoding, and parallel inference. - 2×+ faster inference than standard solutions, up to 3× faster than vLLM. - 50% to 90% cost savings relative to closed model APIs. - 99.99% uptime SLAs, geo-distributed infrastructure, enterprise-grade fault tolerance. - SOC 2 Type II and HIPAA compliant (March 2026). ## Market Standing - **Valuation/Market Cap**: Not disclosed (private company) - **Key Metric**: Total funding $26.7M (two seed rounds: $6.7M in 2021 and $20M led by Capstone Partners in September 2025) - **Notable Investors/Partners**: Capstone Partners (lead investor in latest round), plus two other investors in the first seed round. - **Growth Signals**: - Headcount grew 90% YoY to ~45 employees (as of mid-2026). - Active job postings: 19 positions (up 1800% YoY) – strong hiring momentum. - Appointed Brian Yoo (former Moloco COO) as Chief Business Officer in April 2026 to drive hypergrowth. - Achieved SOC 2 Type II and HIPAA compliance in March 2026, opening regulated markets. - LinkedIn followers grew 257% yearly (8,221 followers). ## Competitive Advantages - **Inventor of continuous batching** – the technique is now industry standard, giving FriendliAI deep technical expertise. - **Purpose-built inference engine** that constantly evolves for state-of-the-art models, delivering 2–3× speed improvements over alternatives like vLLM. - **Massive model library** – 580,000+ Hugging Face models deployable instantly with zero manual optimization. - **Enterprise-grade reliability** – 99.99% SLA, SOC 2 Type II, HIPAA, multi-cloud scaling. - **Cost leadership** – 50–90% savings vs. closed model APIs, maximizing tokens per dollar. ## Strategic Focus - **Hypergrowth** – scaling sales, engineering, and solutions architecture teams globally (US and Korea). - **Compliance and enterprise readiness** – recent SOC 2/HIPAA certification to serve healthcare and regulated industries. - **Multi-cloud and geo-distribution** – expanding infrastructure footprint for global low-latency inference. - **Open-weight and custom model support** – enabling model ownership and flexibility for enterprises. ## Why Work Here - **Culture**: “Boldest innovations come from great teams” – passionate, humble, curious builders. Mission-driven to democratize AI. - **Work policy**: Hybrid (San Francisco and Seoul offices); some roles are in-office (Seoul) or hybrid (SF). Employees engage in a mix of remote and on-site work. - **Engineering focus**: Heavy emphasis on AI inference engine, GPU kernels, Python developer tools, AI agents, backend, and platform security. High technical bar. - **Perks**: Not explicitly listed, but fast-growing startup with significant impact in the AI infrastructure space; opportunity to work on cutting-edge inference optimization. - **Locations**: San Francisco (HQ), Redwood City, CA, and Gangnam-gu, Seoul – global team with cross-cultural collaboration. ## Sources 1. [friendli.ai](https://friendli.ai/) 2. [friendli.ai/careers](https://friendli.ai/careers) 3. [jobs.ashbyhq.com/friendliai](https://jobs.ashbyhq.com/friendliai) 4. [builtin.com/company/friendliai](https://builtin.com/company/friendliai) 5. [linkedin.com/company/friendliai](https://www.linkedin.com/company/friendliai) ## Other roles at FriendliAI - [Director of Product Management](https://feeny.ai/job/director-of-product-management-friendliai-san-francisco-vckbhh439mws) — San Francisco, CA - [Software Engineer - Full Stack](https://feeny.ai/job/software-engineer-full-stack-friendliai-seoul-2kvnfm2hejx8) — Seoul, South Korea - [Software Engineer – Cloud Infrastructure](https://feeny.ai/job/software-engineer-cloud-infrastructure-friendliai-san-francisco-z8fjtpn12kpn) — San Francisco, CA - [Software Engineer - Cloud Infrastructure](https://feeny.ai/job/software-engineer-cloud-infrastructure-friendliai-seoul-rf473fpa9850) — Seoul, South Korea - [Account Executive](https://feeny.ai/job/account-executive-friendliai-san-francisco-z6e1g3bcymfg) — San Francisco, CA - [Software Engineer – Python Developer Tools](https://feeny.ai/job/software-engineer-python-developer-tools-friendliai-seoul-z74wd9kp8fex) — Seoul, South Korea - [Software Engineer – GPU Kernel](https://feeny.ai/job/software-engineer-gpu-kernel-friendliai-seoul-yg1eq3ks7m5a) — Seoul, South Korea - [Software Engineer – AI Inference Engine](https://feeny.ai/job/software-engineer-ai-inference-engine-friendliai-seoul-8pdpe18z4ejk) — Seoul, South Korea - [Customer Success Engineer (contract based)](https://feeny.ai/job/customer-success-engineer-contract-based-friendliai-seoul-9ang4dybr9c9) — Seoul, South Korea - [Software Engineer - Senior Backend](https://feeny.ai/job/software-engineer-senior-backend-friendliai-san-francisco-h5x2ghcnfrg7) — San Francisco, CA