--- title: 'Baseten — company profile' canonical: 'https://feeny.ai/companies/baseten' type: 'company' updated: '2026-07-02' --- # Baseten > AI inference platform for deploying, optimizing, and running machine learning models in production at scale. - **Website:** https://www.baseten.co/ - **Type:** Private (Series F) - **Headquarters:** San Francisco, California, United States - **Founded:** 2019 - **Founders:** Tuhin Srivastava, Amir Haghighat, Phil Howes, Pankaj Gupta - **Valuation:** $13B - **Total raised:** $2B+ - **Latest round:** Series F · $1.5B · June 2026 (valuations of $13B and $11B across two tranches) - **Investors:** Altimeter Capital, Conviction, Spark Capital, Sands Capital, Wellington Management, IVP, Greylock, CapitalG, Bond, South Park Commons, Battery Ventures, Durable Capital Partners, D. E. Shaw Ventures, 01A, Blackbird, Verified Capital - **Business model:** B2B usage-based infrastructure (inference-as-a-service) - **Industries:** Infrastructure - **Open roles:** 61 - **Profile:** https://feeny.ai/companies/baseten ## What they do Baseten runs trained AI models in production for other companies, handling the GPUs, autoscaling, runtime, and performance tuning so engineering teams get a fast, reliable API without operating the infrastructure themselves. It supports open source, custom, and fine tuned models across managed cloud, hybrid, and self hosted deployments. ## Overview **The company betting that inference is the whole game** Baseten has one thesis, and it says it out loud: inference is everything. Not training, not the model itself, but the unglamorous work of running a model in production fast enough, cheaply enough, and reliably enough that a real product can depend on it. That bet has aged well. The company now says its platform handles more than a billion inference requests a day across 87 clusters on 18 cloud environments. Founded in 2019 by Tuhin Srivastava, Amir Haghighat, Phil Howes, and Pankaj Gupta, Baseten spent its early years as a quieter MLOps tool before the generative AI wave turned inference into the bottleneck of the entire industry. It rode that wave hard. In roughly 18 months it raised four rounds and, in mid 2026, closed a $1.5B Series F at a valuation reported as high as $13B. ## What They Do **One platform to put any AI model into production** Baseten takes a trained model and gets it serving live traffic, then keeps it fast and up when demand spikes. You bring an open source, custom, or fine tuned model. Baseten handles the GPUs, the autoscaling, the cold starts, the runtime, and the performance tuning underneath. The pitch to engineers is that they stop babysitting infrastructure. Package a model with the open source Truss framework, push it, and get a production API with 99.99% uptime and horizontal scale. For teams with strict data rules, the same stack runs inside their own cloud or on premises rather than only on Baseten's. ## Problems **Container runtimes were never built for a trillion-parameter model** The problems Baseten chases are the ones that only show up at scale. A model that needs to spin up thousands of replicas hits container images that take minutes to pull. GPU memory constraints that off the shelf runtimes don't understand. Multi tenant isolation the industry never designed for. Its own job posts describe purpose building the container runtime and storage layers for inference, led by containerd maintainers, because patching around a decade of general purpose tooling only goes so far. The payoff it points to is concrete: cold starts cut 2 to 3x, sub 10 second startup for trillion parameter models, and no thundering herd failures when traffic bursts. This is deep systems work, not a wrapper over someone else's cloud. ### Problems addressed - Deploying AI models to production without building and operating GPU infrastructure in-house - Slow cold starts and image-pull latency when scaling to thousands of GPU replicas - Container runtimes and isolation not designed for GPU-heavy, multi-tenant inference - Unreliable, high-latency serving for mission-critical workloads that can't tolerate downtime - High inference cost relative to closed-model APIs - Serving open-source, custom, and fine-tuned models under strict data-governance requirements ## Who It's For **The AI-native companies that cannot afford a slow response** Baseten sells to the teams where latency and downtime are business risk, not an inconvenience. Think AI phone calls, real time transcription, coding assistants, and clinical tools where a stalled model is a lost customer or a missed diagnosis. Its named customers skew exactly that way: Cursor, Notion, Abridge, OpenEvidence, Clay, Gamma, Writer, Hebbia, and Mercor. Inside those companies the buyer is technical. This is a platform for ML engineers and infrastructure teams, not a point and click tool for a non technical manager, and reviewers who wandered in expecting the latter say so plainly. ### Ideal customer profiles - **ML / inference engineers** — Getting a trained model to fast, reliable production serving; Optimizing latency, throughput, and cost per request - **Infrastructure / platform engineers** — Running scalable GPU infrastructure without owning the ops; Cold starts, autoscaling, and multi-cloud capacity management - **AI product teams at AI-native companies** — Shipping AI features that stay up under bursty, high-volume traffic; Controlling inference spend as usage grows - **Enterprises with data-governance requirements** — Running models inside their own cloud or on-prem; Meeting security and compliance constraints while scaling AI ## Products **An inference stack, plus the tools to feed it** The core product is Baseten Cloud, a managed home for models with performance tuning built in. Under it sits the proprietary Inference Stack: custom model runtimes, caching, and custom kernels that Baseten claims deliver higher throughput and lower latency than the defaults. Around that core, the lineup keeps widening. Truss packages models, Chains wires together compound AI systems, Frontier Gateway serves pre optimized versions of the latest open models, and a Training product now closes the loop from fine tuning to serving. The December 2025 acquisition of reinforcement learning startup Parsed pushed it further into post training, so a customer can train, tune, and serve in one place. ## Business Model **Pay for the GPUs you burn, priced below the big labs** Baseten makes money the way a cloud provider does: usage based, you pay for what you run. The bill splits roughly into Model APIs, where you pay per token or per request on shared pre optimized models, and Dedicated Deployments, where you reserve GPUs and autoscale them. The commercial hook is cost. Baseten markets model access at more than 50% below comparable OpenAI rates, and its enterprise motion leans on a forward deployed engineering team that embeds with big accounts. The catch is the flip side of usage pricing: reviewers warn that autoscaling under spiky traffic can make bills hard to predict. ### Plans Baseten prices like a cloud provider: usage based, pay for what you run. It splits roughly into Model APIs, where you pay per token or per request on shared pre optimized models, and Dedicated Deployments, where you reserve and autoscale GPUs. The company markets model access at more than 50% below comparable OpenAI rates, and enterprise deals are custom with a forward deployed engineering team. The tradeoff of usage pricing is that autoscaling under spiky traffic can make bills hard to predict. - **Model APIs** — Usage-based · Developers and teams using shared, pre-optimized models Pay per token or request on the latest open and custom models - Pre-optimized open-source, custom, and fine-tuned models via Frontier Gateway - Pay-per-use pricing - Model access marketed at 50%+ below comparable OpenAI rates - Instant deployment of the latest models (Qwen, DeepSeek, GLM, gpt-oss) - **Dedicated Deployments** — Usage-based (reserved GPUs) · Teams running production workloads at scale Reserved, autoscaling GPU deployments with the full Inference Stack - Dedicated GPU deployments with autoscaling - Fast cold starts and horizontal scale - 99.99% uptime - Access to a range of GPUs (H100, B200) - Single-tenant and self-hosted / hybrid (in-VPC) options - **Enterprise** — Custom · Mission-critical, high-volume AI companies Custom contracts plus hands-on forward deployed engineering - Forward Deployed Engineering team embedded with the account - Self-hosted deployment inside the customer's own cloud / on-prem - Enterprise SLAs and support - Multi-cloud capacity management Good to know: Usage-based billing can lead to unpredictable costs under spiky traffic and autoscaling, per user reviews; Self-hosted / hybrid deployment available for strict data-governance requirements; Enterprise pricing is custom (contact sales) ## Competition **Racing the serverless GPU crowd on how deep the stack goes** Baseten competes with a crowded field of inference and serverless GPU platforms, from Modal and Replicate to Together AI, Fireworks, and Anyscale, plus the hyperscalers' own serving tools. Its argument is depth. Where many rivals sit on top of standard runtimes, Baseten owns the pipeline from the moment a model is pushed to the moment a request returns, which is what lets it fix cold starts and multi tenant isolation at the root. The forward deployed engineering team is the other moat. It is consulting glued to infrastructure, and it makes the biggest accounts sticky. ### Their edge - **Vertical ownership of the inference pipeline** — Baseten controls the stack from model push to response, letting it fix cold starts, image pull latency, and multi-tenant isolation at the root instead of at higher layers. - **Proprietary performance stack** — Custom runtimes, caching, and custom kernels aimed at higher throughput and lower latency than off-the-shelf serving; the team writes its own communication kernels and hardens containerd. - **Forward Deployed Engineering** — A hands-on engineering team embeds with major accounts, making high-value customers like Abridge and Notion sticky. - **Deployment flexibility** — Managed cloud, hybrid, and fully self-hosted in a customer VPC, which wins security-conscious enterprises that rivals' pure-managed offerings can't serve. ### Where they're betting - Dominating inference-as-a-service on the thesis that inference becomes the largest AI market - Owning the full model lifecycle (inference + training + post-training) after the Parsed acquisition - Deep performance research: custom kernels, GPU networking (RDMA), and bleeding-edge Blackwell/Rubin hardware - Reliability and cost for mission-critical, high-volume workloads (voice AI, real-time streaming) ## Proof **A billion requests a day and 20x revenue growth** The traction numbers are the loudest part of the story. Baseten reported roughly 20x year over year revenue growth heading into its Series F, and says the platform now clears more than a billion inference requests daily across 87 clusters on 18 clouds. Investors clearly believed it: the valuation went from about $2.15B in September 2025 to $5B in January 2026 to as high as $13B by mid 2026. The customer side backs it up. Abridge runs clinical notes in production on Baseten, and Cursor, whose coding assistant crossed $2B in ARR, routes inference through it. ## What People Say **Great when it works, expensive when it scales** Developers who like Baseten like it for the reasons it sells itself: fast, dependable model serving, a clean path from a model to a live API, and low ops overhead, helped along by a hands on engineering team. Truss gets specific praise as a clean way to package models. The recurring gripe is money. Usage pricing plus aggressive autoscaling can produce bills that are hard to forecast, especially for dedicated deployments under bursty traffic, and a few users report slow support responses. None of it contradicts the pitch so much as name its price. ## Funding **$1.5B in one round, and four raises in 18 months** Baseten's fundraising cadence reads like a company trying to keep up with its own demand. The June 2026 Series F brought in $1.5B across two tranches at valuations of $13B and $11B, led by Altimeter Capital, Conviction, and Spark Capital, with Sands Capital and Wellington Management as co leads. That followed a $300M Series E at $5B in January 2026 and a $150M round at roughly $2.15B in September 2025. All told, Baseten has raised more than $2B since 2019. The backer list is deep: IVP, Greylock, CapitalG, Bond, South Park Commons, Battery Ventures, Durable Capital, and D. E. Shaw Ventures among them. ## Team & Culture **Built by engineers who lived through deployment hell** Baseten frames itself as an engineer's company, and the founding story is the tell: a team that felt the pain of shipping models firsthand and decided the infrastructure was the real product. The culture the job posts describe is high agency, first principles across the whole stack, and customer obsessed, with talent pulled from Meta, Stripe, Google, NVIDIA, and Databricks. The work is unusually low level for a startup this young. Engineers contribute to containerd upstream, write custom communication kernels, and qualify Blackwell GPU clusters before most of the industry has touched them. Roles run out of San Francisco with an in person lean, and the company is hiring hard across engineering, infrastructure, go to market, and legal. ## Compensation **Wide engineering bands, equity in every offer** Baseten discloses pay on many of its roles, and the spread is wide because the ladder is: individual contributors up through managers of managers all sit in the same San Francisco market. Engineering bands run from roughly $95K at the low end to about $380K at the senior end, with sales roles carrying commission driven upside into the mid $300Ks. Every posting states the same line about equity, so ownership is part of the standard package, not a perk reserved for early hires. All disclosed figures are US dollars and base salary; equity and commission sit on top. ## In the News **A funding-and-acquisition streak through late 2025 and 2026** The headline run is hard to miss: three up rounds and an acquisition in under a year. The $1.5B Series F in June 2026 drew wide coverage for its size and $13B valuation, and the December 2025 acquisition of reinforcement learning startup Parsed signaled a push from pure inference into owning the full model lifecycle. The links below trace the raises, the Parsed deal, and the company's own posts on the systems work underneath the pitch. ### Coverage - [Baseten Raises $1.5 Billion to Power the Next Era of AI Inference](https://www.businesswire.com/news/home/20260622645563/en/Baseten-Raises-%241.5-Billion-to-Power-the-Next-Era-of-AI-Inference) — Business Wire (2026-06-22) - [Baseten Raises $1.5 Billion Series F at Up to $13 Billion Valuation](https://www.citybiz.co/article/863525/baseten-raises-1-5-billion-series-f-at-up-to-13-billion-valuation/) — citybiz (2026-06-22) - [Announcing our Series F](https://www.baseten.co/blog/announcing-our-series-f/) — Baseten Blog (2026-06-22) - [Baseten Raises $300M at a $5B Valuation to Power a Multi-Model Future](https://www.businesswire.com/news/home/20260123035833/en/Baseten-Raises-$300M-at-a-$5B-Valuation-to-Power-a-Multi-Model-Future) — Business Wire (2026-01-23) - [Baseten Acquires Parsed to Enable Companies to Own Their Intelligence](https://www.businesswire.com/news/home/20251210284353/en/Baseten-Acquires-Parsed-to-Enable-Companies-to-Own-Their-Intelligence) — Business Wire (2025-12-10) - [Inference Provider Baseten Acquires Reinforcement Learning Startup Parsed](https://www.theinformation.com/newsletters/ai-agenda/inference-provider-baseten-acquires-reinforcement-learning-startup-parsed) — The Information (2025-12-10) - [Exclusive: Baseten, AI inference unicorn, raises $150 million at $2.15 billion valuation](https://fortune.com/2025/09/05/exclusive-baseten-ai-inference-unicorn-raises-150-million-at-2-15-billion-valuation/) — Fortune (2025-09-05) - [AI startup Baseten raises $75 million following DeepSeek's emergence](https://www.cnbc.com/2025/02/19/ai-inference-startup-baseten-raises-75-million.html) — CNBC (2025-02-19) - [Baseten revenue, valuation & funding](https://sacra.com/c/baseten/) — Sacra (2026) ## Company details - **Mission:** Give teams the tooling, expertise, and hardware to bring performant AI products to market as fast and reliably as possible, on the belief that inference is the largest market AI will create. - **Products:** Baseten Cloud, Baseten Inference Stack, Truss, Baseten Chains, Frontier Gateway, Training on Baseten, Forward Deployed Engineering - **Notable customers:** Cursor, Notion, Abridge, OpenEvidence, Clay, Gamma, Writer, Hebbia, Mercor - **Customer segments:** AI-native companies, Enterprises deploying AI in production, ML engineers and developers, Healthcare / clinical AI, Developer tools and coding assistants - **Buyers / users:** ML engineers, Infrastructure / platform engineers, AI product teams, CTOs and VPs of Engineering at AI-native companies - **Competitors:** Modal, Replicate, Together AI, Fireworks AI, Anyscale, Hyperscaler serving tools (SageMaker, Vertex AI) - **What sets them apart:** Owns the full pipeline from model push to response, fixing cold starts and isolation at the root rather than patching over standard runtimes; Proprietary Inference Stack with custom runtimes, caching, and custom kernels; Forward Deployed Engineering team that embeds with top accounts; Cloud, hybrid, and self-hosted (in-VPC / on-prem) deployment for strict data governance; Founding team with deep systems and applied-AI credibility, hiring containerd maintainers and GPU-networking specialists - **Tech stack:** NVIDIA H100 / B200 GPUs, Kubernetes, Multi-cloud (AWS, GCP), vLLM, TensorRT-LLM, containerd - **Integrations:** vLLM, TensorRT-LLM, Triton, ComfyUI, Hugging Face ## Open roles (61) - [IT Support / Operations Engineer](https://jobs.ashbyhq.com/baseten/c07eb44f-b5fa-4808-90b7-03b265d97836) — New York, NY - [Strategic Finance, GTM](https://jobs.ashbyhq.com/baseten/3677992a-2a97-4a55-98ce-e2719922f263) — San Francisco, CA - [Immigration and Mobility Lead](https://jobs.ashbyhq.com/baseten/9d293c95-898b-4dbe-9479-c8f0e0b35688) — San Francisco, CA - [GRC Manager](https://jobs.ashbyhq.com/baseten/d1b53083-15f9-4f92-9c7d-11792a041f55) — San Francisco, CA - [Field Marketing Manager](https://jobs.ashbyhq.com/baseten/c6e86fec-9412-4d47-b23c-19da03003364) — New York, NY - [People Business Partner, GTM](https://jobs.ashbyhq.com/baseten/ddbd38d0-1914-4bac-9381-75a72f5023ca) — San Francisco, CA - [Account Executive - Enterprise](https://jobs.ashbyhq.com/baseten/75d81d67-465a-4d65-9745-d1c06b0c7aec) — San Francisco, CA - [Product Manager, Enterprise](https://jobs.ashbyhq.com/baseten/342a4c0e-7da3-490e-927f-dffb51f847b2) — San Francisco, CA - [Software Engineer - GPU Fabric Observability](https://jobs.ashbyhq.com/baseten/cf7fa477-0290-4846-9294-f7bc006ae997) — San Francisco, CA - [Brand Designer](https://jobs.ashbyhq.com/baseten/c2f8fe4f-07c1-4d43-ae5c-05f057842e57) — San Francisco, CA - [Base Labs Fellowship](https://jobs.ashbyhq.com/baseten/58d7d8e6-86ee-43a1-baec-3dddcb661d51) — San Francisco, CA - [Post-Training Research Scientist](https://jobs.ashbyhq.com/baseten/7c9d2bb0-ac03-4a3c-86c3-cf720cd314e8) — San Francisco, CA - [Manager, Solutions Architect](https://jobs.ashbyhq.com/baseten/1532aa00-5a93-4967-bb24-6561495d9605) — San Francisco, CA - [Product Manager, Inference Platform](https://jobs.ashbyhq.com/baseten/3027e0bc-731f-4fef-b081-2031590766fd) — San Francisco, CA - [Software Engineer - Voice AI (Inference Runtime)](https://jobs.ashbyhq.com/baseten/6e396eb7-acb3-436a-89ec-05e755c477f2) — San Francisco, CA - [Cloud Platform Engineer](https://jobs.ashbyhq.com/baseten/916e9d62-492a-4562-971f-3e97e3868392) — San Francisco, CA - [Software Engineer - Baseten Inference Stack](https://jobs.ashbyhq.com/baseten/c8701794-bdc1-4932-bffa-a444ce57ed73) — San Francisco, CA - [Account Executive - AI Native: Startups](https://jobs.ashbyhq.com/baseten/1261fe76-f577-4b3d-aac8-df65846ae445) — New York, NY - [Solutions Architect](https://jobs.ashbyhq.com/baseten/c64515f9-a8f7-4633-9340-17cda56b1ef0) — San Francisco, CA - [Account Executive - AI Native: Strategic](https://jobs.ashbyhq.com/baseten/8608bdbc-1b8b-4b69-a433-83864f151375) — New York, NY - [Software Engineer - Infrastructure](https://jobs.ashbyhq.com/baseten/ae64d1d4-7b0a-4be4-8d77-7f5ce63849a7) — San Francisco, CA - [Engineering Manager, Cloud Platform](https://jobs.ashbyhq.com/baseten/0870ed34-7365-4b9f-a50a-481783b8c266) — San Francisco, CA - [Software Engineer - Enterprise Platform](https://jobs.ashbyhq.com/baseten/e2638fb0-15c8-4109-b7ee-70df1325885d) — San Francisco, CA - [Engineering Manager, Forward Deployed Engineering (LLM)](https://jobs.ashbyhq.com/baseten/1b6fc00e-f3fa-440b-9e4e-3013f7a5010e) — San Francisco, CA - [Infrastructure Ops Engineer](https://jobs.ashbyhq.com/baseten/844e2c96-690a-4365-9822-f092ffa03944) — San Francisco, CA - [Data Engineer](https://jobs.ashbyhq.com/baseten/c2ba6c50-d282-4478-98e9-668f94facde8) — San Francisco, CA - [OS / K8s Systems Engineer](https://jobs.ashbyhq.com/baseten/dd166491-a4a2-4fb8-b42a-6a60e4049d15) — San Francisco, CA - [Software Engineer - Billing & Internal Tooling](https://jobs.ashbyhq.com/baseten/d073254f-729c-471e-9a33-520358ead183) — San Francisco, CA - [Ecosystem Partnerships Product Marketing Manager](https://jobs.ashbyhq.com/baseten/a34233b9-d4ed-4a8a-b942-e16df27f8935) — San Francisco, CA - [Revenue Strategy & Operations](https://jobs.ashbyhq.com/baseten/6d32aa11-ac93-4f90-8f62-bdeb79214ee5) — San Francisco, CA - [Forward Deployed Engineer](https://jobs.ashbyhq.com/baseten/84c1801c-1a65-49fb-aaaa-beeafd530e7e) — San Francisco, CA - [Product Manager, Developer Experience](https://jobs.ashbyhq.com/baseten/2d78fdcf-53e1-45d3-a047-2aefb5ad3153) — San Francisco, CA - [Software Engineer - GPU Networking & Distributed Systems](https://jobs.ashbyhq.com/baseten/1f7d7fda-5540-4205-890b-cdbf774f0814) — San Francisco, CA - [Software Engineer - Capacity](https://jobs.ashbyhq.com/baseten/902a7ddb-c21f-4272-aaab-879680697986) — San Francisco, CA - [Software Engineer - Model Performance](https://jobs.ashbyhq.com/baseten/d29e748c-7209-460d-a024-8f77ae0a3d4d) — San Francisco, CA - [Account Executive - Enterprise](https://jobs.ashbyhq.com/baseten/223f44cf-c7b8-416c-b694-e265579aa1c2) — New York, NY - [Post-Training Research Engineer](https://jobs.ashbyhq.com/baseten/68328a94-f785-463b-a3d7-b0dfc11839b5) — San Francisco, CA - [Software Engineer - Internal Platform](https://jobs.ashbyhq.com/baseten/081cb52b-5e88-40a1-8def-1e82c8bc97de) — San Francisco, CA - [Manager, Compensation](https://jobs.ashbyhq.com/baseten/8cac65b1-f67c-4a5c-89b1-1f90ecdd4ddc) — San Francisco, CA - [Strategic Finance Associate / Sr. Associate](https://jobs.ashbyhq.com/baseten/71a011b6-0f17-4c0a-b0ba-38deffce1adb) — San Francisco, CA - [Field Productivity & Enablement Lead](https://jobs.ashbyhq.com/baseten/5a1c6228-3906-4ccc-9988-d9bd67383b9d) — San Francisco, CA - [Manager, Startup Sales](https://jobs.ashbyhq.com/baseten/b1b5ab31-f90c-41f9-b9f2-13408b1c072b) — San Francisco, CA - [Account Executive - AI Native: Startups](https://jobs.ashbyhq.com/baseten/df2f0ebb-a84a-45c2-bc74-dc9589fcc1ae) — San Francisco, CA - [Software Engineer- Model Performance Systems](https://jobs.ashbyhq.com/baseten/75d7beac-0298-40fa-b206-2e0c0c08a64f) — San Francisco, CA - [Software Engineer - GPU Kernels](https://jobs.ashbyhq.com/baseten/ddb5bc98-6116-49a2-802e-1c05398663f1) — San Francisco, CA - [Software Engineer - AI Enablement](https://jobs.ashbyhq.com/baseten/b88a68b7-d2bc-4a30-a79a-3ef292ad7c26) — San Francisco, CA - [Software Engineer - Training Product](https://jobs.ashbyhq.com/baseten/126d54b4-a7bc-4456-bf4d-5d224e4f5d63) — San Francisco, CA - [Software Engineer - Dedicated Inference](https://jobs.ashbyhq.com/baseten/fc6e5f2e-eb2d-4a6c-8a51-8422e8662bde) — San Francisco, CA - [Site Reliability Engineer](https://jobs.ashbyhq.com/baseten/ff008b8e-b38d-4941-b24f-9a48c970e7fb) — San Francisco, CA - [Manager, Strategic Sales](https://jobs.ashbyhq.com/baseten/4e507a15-c9bb-4877-b422-f3eb462a3732) — San Francisco, CA _…and 11 more at https://feeny.ai/companies/baseten/jobs_ --- _Source: https://feeny.ai/companies/baseten · profile updated 2026-07-02_