--- title: 'Software Engineer - Cloud Infrastructure at FriendliAI' canonical: 'https://feeny.ai/job/software-engineer-cloud-infrastructure-friendliai-seoul-rf473fpa9850' type: 'job' last_seen: '2026-09-05' --- # Software Engineer - Cloud Infrastructure at FriendliAI - **Company:** FriendliAI - **Location:** Seoul, South Korea - **Employment:** full-time - **Work type:** onsite - **Posted:** 2026-08-07 - **Last confirmed live:** 2026-09-05 - **Apply:** https://jobs.ashbyhq.com/friendliai/9a0d8751-64bb-48e6-8ee7-bd2b92f08377 ## Job description ## ABOUT THE JOB FriendliAI is looking for a Cloud Infrastructure Engineer to own the architecture and evolution of the cluster platform behind our GPU-accelerated AI inference cloud. As a Software Engineer, Cloud Infrastructure, you will design how our clusters are built and connected, extend Kubernetes where its defaults fall short, and own the network path that inference traffic depends on. Inference is an unforgiving workload for Kubernetes. Traffic is bursty and latency-sensitive, GPU capacity is scarce and inelastic, tenants must stay isolated, and multi-node serving depends on the network holding up under sustained load. This is a hands-on architecture role for an engineer who has already run large clusters in production and wants to push them further. ## KEY RESPONSIBILITIES Cluster Architecture - Own the architecture of our multi-cluster, multi-tenant Kubernetes fleet across both managed and self-managed clusters: cluster topology, control plane and etcd lifecycle, and zero-downtime upgrades. - Extend Kubernetes with custom controllers, operators, and CRDs so platform behavior is encoded in software rather than runbooks. - Design GPU scheduling and capacity strategy, including topology-aware placement, node pools, priority and preemption, and quota across tenants. - Build autoscaling that matches inference traffic: queue-driven pod scaling, node autoscaling, scale-to-zero, and cold-start reduction. Networking - Own the Kubernetes network data plane: CNI, IPAM, DNS, ingress, and L4/L7 load balancing. - Design cross-AZ, cross-region, and cross-cluster connectivity, and operate the service mesh for routing, mTLS, and traffic policy. - Debug production network issues (packet loss, conntrack exhaustion, MTU mismatches, DNS latency, load balancer behavior) and drive permanent fixes. Reliability & Collaboration - Define SLOs for platform-critical systems and lead post-incident hardening. - Deliver infrastructure as code with Terraform, Helm, and GitOps. - Partner with the inference engine, platform, SRE, and security teams to turn serving requirements into platform capabilities. ## QUALIFICATIONS - 5+ years designing, building, and operating large-scale Kubernetes infrastructure in production. - Bachelor’s or Master’s degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent. - Proven experience operating large-scale, high-traffic network services in production. - Deep understanding of Kubernetes internals: API server, scheduler, controller loops, kubelet, and etcd. - Strong command of Kubernetes and cloud networking: CNI, kube-proxy/eBPF datapaths, DNS, load balancing, service mesh, and VPC routing. - Proficiency with AWS, Terraform, Helm, and Ansible. - Programming skills in Go or Python, with the ability to build infrastructure tooling and automation. - Strong debugging skills across distributed systems, containers, and the Linux networking stack. - Clear written and verbal communication, including the ability to document architectural decisions for other engineers. ## PREFERRED EXPERIENCE - Large-scale Kubernetes operations in a high-traffic domain such as gaming, e-commerce, or public cloud. - Cilium and eBPF, including kube-proxy replacement or upstream contributions. - Cluster provisioning and lifecycle management with Kubespray or similar Ansible-based tooling. - GPU orchestration: NVIDIA GPU Operator, device plugins, or Dynamic Resource Allocation (DRA). - High-performance networking for distributed workloads: RDMA/RoCE, InfiniBand, EFA, SR-IOV, or NCCL tuning. - Multi-cloud, hybrid-cloud, or bare-metal Kubernetes operations. - Contributions to Kubernetes, Cilium, Istio, or other CNCF projects. ## BENEFITS - Flexible working hours - Daily lunch and dinner provided; unlimited snacks and beverages - Supportive and highly collaborative work environment - Health check-up support and top-tier equipment/hardware support - A front-row seat to the generative AI infrastructure revolution - Competitive compensation, startup equity, health insurance, and other benefits. ## ABOUT FRIENDLIAI FriendliAI is the fastest inference cloud for agents, built to run frontier open-weight models in production at scale. It delivers up to 7x faster output token speed, up to 90% lower inference costs, and 99.99% uptime across the most demanding agent workloads — long-context inference, real-time streaming, and accurate tool calling. We are a small, fast-moving team doing work that matters at one of the most exciting moments in the history of technology. With our world-class inference stack, we are building the platform teams can actually rely on. ## About FriendliAI ## Company Overview - **One-liner**: FriendliAI is The Frontier AI Inference Cloud, providing a highly optimized platform for deploying, scaling, and monitoring large language and multimodal models with unmatched speed and cost efficiency. - **Entity Type**: Private (Seed stage; $26.7M total funding) - **Headquarters**: San Francisco, California, United States (with offices in Seoul, South Korea and Redwood City, CA) - **Founded**: 2021 - **Founders**: Byung-Gon Chun (CEO) and Gyeong-In Yu (CTO) – the researchers who invented the continuous batching technique now standard in AI inference. ## Core Business - **Primary industry**: AI Infrastructure / Cloud Inference Platform (Infrastructure as a Service) - **Target customers**: B2B – enterprises and AI teams deploying generative AI and agent workloads at production scale; focuses on open-weight and custom models. - **Mission**: Democratize access to production-grade AI so teams can focus on building great products. ## Products & Services - **[FriendliAI Inference Cloud](https://friendli.ai/)**: A fully managed platform that instantly deploys 580,000+ Hugging Face models (language, audio, vision) with one click. Supports bring-your-own fine-tuned or proprietary models. Key features: - Custom GPU kernels, smart caching, continuous batching, speculative decoding, and parallel inference. - 2×+ faster inference than standard solutions, up to 3× faster than vLLM. - 50% to 90% cost savings relative to closed model APIs. - 99.99% uptime SLAs, geo-distributed infrastructure, enterprise-grade fault tolerance. - SOC 2 Type II and HIPAA compliant (March 2026). ## Market Standing - **Valuation/Market Cap**: Not disclosed (private company) - **Key Metric**: Total funding $26.7M (two seed rounds: $6.7M in 2021 and $20M led by Capstone Partners in September 2025) - **Notable Investors/Partners**: Capstone Partners (lead investor in latest round), plus two other investors in the first seed round. - **Growth Signals**: - Headcount grew 90% YoY to ~45 employees (as of mid-2026). - Active job postings: 19 positions (up 1800% YoY) – strong hiring momentum. - Appointed Brian Yoo (former Moloco COO) as Chief Business Officer in April 2026 to drive hypergrowth. - Achieved SOC 2 Type II and HIPAA compliance in March 2026, opening regulated markets. - LinkedIn followers grew 257% yearly (8,221 followers). ## Competitive Advantages - **Inventor of continuous batching** – the technique is now industry standard, giving FriendliAI deep technical expertise. - **Purpose-built inference engine** that constantly evolves for state-of-the-art models, delivering 2–3× speed improvements over alternatives like vLLM. - **Massive model library** – 580,000+ Hugging Face models deployable instantly with zero manual optimization. - **Enterprise-grade reliability** – 99.99% SLA, SOC 2 Type II, HIPAA, multi-cloud scaling. - **Cost leadership** – 50–90% savings vs. closed model APIs, maximizing tokens per dollar. ## Strategic Focus - **Hypergrowth** – scaling sales, engineering, and solutions architecture teams globally (US and Korea). - **Compliance and enterprise readiness** – recent SOC 2/HIPAA certification to serve healthcare and regulated industries. - **Multi-cloud and geo-distribution** – expanding infrastructure footprint for global low-latency inference. - **Open-weight and custom model support** – enabling model ownership and flexibility for enterprises. ## Why Work Here - **Culture**: “Boldest innovations come from great teams” – passionate, humble, curious builders. Mission-driven to democratize AI. - **Work policy**: Hybrid (San Francisco and Seoul offices); some roles are in-office (Seoul) or hybrid (SF). Employees engage in a mix of remote and on-site work. - **Engineering focus**: Heavy emphasis on AI inference engine, GPU kernels, Python developer tools, AI agents, backend, and platform security. High technical bar. - **Perks**: Not explicitly listed, but fast-growing startup with significant impact in the AI infrastructure space; opportunity to work on cutting-edge inference optimization. - **Locations**: San Francisco (HQ), Redwood City, CA, and Gangnam-gu, Seoul – global team with cross-cultural collaboration. ## Sources 1. [friendli.ai](https://friendli.ai/) 2. [friendli.ai/careers](https://friendli.ai/careers) 3. [jobs.ashbyhq.com/friendliai](https://jobs.ashbyhq.com/friendliai) 4. [builtin.com/company/friendliai](https://builtin.com/company/friendliai) 5. [linkedin.com/company/friendliai](https://www.linkedin.com/company/friendliai) ## Other roles at FriendliAI - [Director of Product Management](https://feeny.ai/job/director-of-product-management-friendliai-san-francisco-vckbhh439mws) — San Francisco, CA - [Software Engineer - Full Stack](https://feeny.ai/job/software-engineer-full-stack-friendliai-seoul-2kvnfm2hejx8) — Seoul, South Korea - [Software Engineer – Cloud Infrastructure](https://feeny.ai/job/software-engineer-cloud-infrastructure-friendliai-san-francisco-z8fjtpn12kpn) — San Francisco, CA - [Account Executive](https://feeny.ai/job/account-executive-friendliai-san-francisco-z6e1g3bcymfg) — San Francisco, CA - [Software Engineer – Python Developer Tools](https://feeny.ai/job/software-engineer-python-developer-tools-friendliai-seoul-z74wd9kp8fex) — Seoul, South Korea - [Software Engineer – GPU Kernel](https://feeny.ai/job/software-engineer-gpu-kernel-friendliai-seoul-yg1eq3ks7m5a) — Seoul, South Korea - [Software Engineer – AI Inference Engine](https://feeny.ai/job/software-engineer-ai-inference-engine-friendliai-seoul-8pdpe18z4ejk) — Seoul, South Korea - [Customer Success Engineer (contract based)](https://feeny.ai/job/customer-success-engineer-contract-based-friendliai-seoul-9ang4dybr9c9) — Seoul, South Korea - [Software Engineer - Senior Backend](https://feeny.ai/job/software-engineer-senior-backend-friendliai-san-francisco-h5x2ghcnfrg7) — San Francisco, CA - [Software Engineer – AI Agents](https://feeny.ai/job/software-engineer-ai-agents-friendliai-san-francisco-7v0cdvt2ded6) — San Francisco, CA