--- title: 'Inference Engineer at Hyperbolic Labs' canonical: 'https://feeny.ai/job/inference-engineer-hyperbolic-labs-san-francisco-208hxvbmj3jn' type: 'job' last_seen: '2026-09-15' --- # Inference Engineer at Hyperbolic Labs - **Company:** Hyperbolic Labs - **Location:** San Francisco, CA - **Employment:** full-time - **Posted:** 2026-09-09 - **Last confirmed live:** 2026-09-15 - **Apply:** https://jobs.ashbyhq.com/hyperbolic/38121cf9-aa24-44f6-a500-ebcd2d3ca3ef ## Job description ## Who We Are Hyperbolic Labs is on a mission to democratize AI by breaking down the barriers to computing power with our Open-Access AI Cloud. By making better use of idle computing resources across the globe, we offer an innovative GPU marketplace and AI inference service that promise affordability and accessibility for all. As pioneers at the intersection of AI and open-source technology, we believe in an open future where AI innovation is limited only by imagination, not by access to resources. We're looking for forward-thinking individuals who share our passion for making AI universally accessible, secure, and affordable. Join us in building a platform that empowers innovators everywhere to turn their visionary AI projects into reality. ## About the Role We're looking for an Inference Engineer to build inference capabilities on top of Forge, our unified control plane, so customers can consume model tokens without managing GPUs and our NeoCloud partners get a full-stack path to their own token-factory offering. You'll own how models get deployed and served across clusters distributed around the world, on heterogeneous hardware. Deployment comes first: serving models on Forge and our Kubernetes offering, evaluating inference frameworks, and standing up the monitoring, gateways, and endpoints that make a deployment production-ready. From there the work expands into optimization, autoscaling, KV-cache orchestration, and customer inference debugging. This is the primary seat for inference at Hyperbolic \u2014 you'll build it end to end, with real influence over where the scope lands. ## Who You Are - Strong general inference background with a broad, high-level command of the stack rather than a narrow specialty — you can reason about the whole path from request to token - Deep Kubernetes experience, including hands-on ability to operate clusters in production, not just deploy to them - Solid grasp of the concepts that govern inference performance: TTFT, disaggregated inference, speculative decoding, and KV cache and its inner workings - Familiarity with modern inference frameworks and serving engines, and the judgment to evaluate and select among them for a given workload - Working knowledge of NVIDIA Dynamo and how it fits into a distributed serving architecture - Experience setting up monitoring, gateways, and endpoints for production inference services - Proven ability to build a product end to end — you've taken something from nothing to serving real traffic - Strong self-initiative and comfort operating as the primary owner of an area with minimal direction - Generalist instincts: you're willing to pick up adjacent work when it's what the product needs ## Preferred Qualifications - Experience spanning both inference deployment and inference optimization - Hands-on model optimization work — quantization, batching strategies, kernel-level tuning, or similar - Understanding of RDMA and high-performance networking as they apply to distributed serving - Experience deploying inference across heterogeneous accelerators - Background supporting customers directly on inference debugging and performance issues - Experience at a GPU cloud, inference provider, or AI infrastructure company Hyperbolic is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. ## About Hyperbolic Labs ## Company Overview - **One-liner**: Hyperbolic provides an open-access cloud platform offering affordable, on-demand GPU compute and AI inference services for developers, researchers, and AI teams. - **Entity Type**: Private (Series A) - **Headquarters**: San Francisco, California, United States (also listed as Irvine, California, per CB Insights) [LinkedIn], [CB Insights] - **Founded**: 2022 [CB Insights] - **Founders**: Jasper Zhang (CEO) and Yuchen Jin (CTO, departed April 2026) [hyperbolic.ai/about], [LinkedIn] ## Core Business - **Primary industry/industries**: AI Infrastructure, Cloud Computing, GPU-as-a-Service - **Target customers**: B2B – AI startups, ML engineers, academic researchers, enterprise AI teams, and individual developers. - **Mission or purpose statement**: “Build the world's most accessible and comprehensive AI platform, empowering developers with affordable, seamless access to meet all their AI needs in one unified platform.” [hyperbolic.ai/about] ## Products & Services - **On-Demand GPU Clusters**: Rent H100/H200 GPUs in under one minute with no sales calls or quota limits. Pay-as-you-go, scale up/down as needed. [hyperbolic.ai] - **Reserved Clusters**: Guaranteed, isolated capacity for long-term workloads at discounted prepaid pricing. Ideal for 24/7 inference or large training jobs. [hyperbolic.ai] - **Dedicated Endpoints**: Single-tenant GPU instances with private endpoints for high-throughput inference (100K+ tokens/min) and full control. Hourly pricing. [hyperbolic.ai] - **Serverless Inference API**: OpenAI-compatible API to run models like Llama, Qwen, DeepSeek, SDXL, Flux, etc. Swap base URL and key with minimal code changes. [hyperbolic.ai] - **AI Consulting Services**: Engineering support for setup, scaling, sharding, and debugging across training, fine-tuning, and inference. [hyperbolic.ai] ## Market Standing - **Valuation/Market Cap**: Not disclosed - **Total Funding**: $19.73M (Series A: $12M in December 2024; Seed: $7M in July 2024; additional earlier rounds) [LinkedIn], [CB Insights] - **Notable Investors/Partners**: Variant, Polychain Capital, Chapter One, Faction, Bankless Ventures, Blockchain Builders Fund, and 29+ others. [CB Insights] - **Growth Signals**: 200,000+ builders on the platform; claims 3–10x less expensive than inference competitors; presence in 5 countries (US, Singapore, Vietnam, France, South Korea). [hyperbolic.ai], [LinkedIn] - **Key Metric**: Total funding $19.73M; no public revenue data. ## Competitive Advantages - **Cost leadership**: 3–10x cheaper than competing inference providers, with no hidden fees or long-term commitments. [hyperbolic.ai] - **Zero quota limits**: No artificial caps on GPU usage – provision instantly. [hyperbolic.ai] - **Unique model support**: Only platform serving Llama-3.1-405B-Base in BF16 (high-precision) and FP8 (ultra-fast inference). Endorsed by Andrej Karpathy. [hyperbolic.ai] - **Open-access ethos**: Focus on democratizing compute by aggregating underutilized GPUs (consumer-grade and data center), enabling AI development without vendor lock-in. [hyperbolic.ai/about] - **Speed**: Deploy a cluster in under one minute; pay with credit card or crypto. [hyperbolic.ai] ## Strategic Focus - **Scalability & automation**: Building the most automated platform to eliminate sales calls and friction. [hyperbolic.ai/about] - **Open ecosystem**: Partnering with top research labs, institutions, and major AI/ML teams to advance open-source AI and open-access compute. [hyperbolic.ai/about] - **Expanding model variety**: Continuously adding new state-of-the-art models and inference optimizations. [hyperbolic.ai] ## Why Work Here - **Culture**: Emphasizes “open access, collaboration, innovation, and automation.” Founding story rooted in removing barriers for developers and researchers. [hyperbolic.ai/about] - **Team**: Small (17 employees as of mid-2026), with a flat structure and hands-on roles. Engineering-centric (5 technical staff) plus growing GTM and product teams. [LinkedIn] - **Remote/Hybrid**: HQ in San Francisco; employees across 5 countries – likely offers remote flexibility, though policy not explicitly stated. [LinkedIn] - **Notable perks**: Work on cutting-edge AI infrastructure; direct impact on product; collaboration with top AI labs; use of latest GPUs (H100/H200). [hyperbolic.ai/about] - **Caution**: Recent high-profile departures (co-founder/CTO Yuchen Jin, Head of Business Operations, and a founding AI engineer in April 2026) – candidates should investigate stability and leadership continuity. [LinkedIn] ## Sources 1. [Hyperbolic – Official Website](https://www.hyperbolic.ai/) 2. [Hyperbolic – About Page](https://www.hyperbolic.ai/about) 3. [Hyperbolic – LinkedIn Company Page](https://www.linkedin.com/company/hyperbolic-labs) 4. [Hyperbolic Labs – CB Insights Profile](https://www.cbinsights.com/company/hyperbolic-labs) 5. [Hyperbolic – Careers Page (Ashby)](https://jobs.ashbyhq.com/hyperbolic) ## Other roles at Hyperbolic Labs - [Quantitative Researcher](https://feeny.ai/job/quantitative-researcher-hyperbolic-labs-san-francisco-1j2y376nj205) — San Francisco, CA - [Head of Marketing](https://feeny.ai/job/head-of-marketing-hyperbolic-labs-san-francisco-07cmdwq0kxt9) — San Francisco, CA - [GTM (Operations)](https://feeny.ai/job/gtm-operations-hyperbolic-labs-remote-ew8zr8za3vbr) - [Technical Support Engineer](https://feeny.ai/job/technical-support-engineer-hyperbolic-labs-remote-7vnm5z6a127b) - [Capital Markets Lead](https://feeny.ai/job/capital-markets-lead-hyperbolic-labs-san-francisco-985j3avwskz1) — San Francisco, CA - [Supply Specialist](https://feeny.ai/job/supply-specialist-hyperbolic-labs-remote-91ddkx7r0sm1) - [VP of Engineering](https://feeny.ai/job/vp-of-engineering-hyperbolic-labs-san-francisco-fyhpfb9k3fxp) — San Francisco, CA - [Technical Writer](https://feeny.ai/job/technical-writer-hyperbolic-labs-remote-gj2zdmf5d7kf) - [Head of Supply](https://feeny.ai/job/head-of-supply-hyperbolic-labs-san-francisco-qf31th4bdmsj) — San Francisco, CA - [Forward Deployed Infrastructure Engineer - Eastern US](https://feeny.ai/job/forward-deployed-infrastructure-engineer-eastern-us-hyperbolic-labs-remote-nwccwv8jz4bb)