

Cerebras Systems
CBRS · NASDAQCerebras builds wafer-scale AI chips, the largest ever made, plus the supercomputers and fast inference cloud that run on them.

Overview: The company that bet a whole silicon wafer on one chip, and just went public on it
Most chipmakers cut a wafer into hundreds of small dies. Cerebras keeps the whole wafer as a single chip, the Wafer-Scale Engine, which is why its processor is 56 to 58 times larger than a GPU and, the company says, delivers AI inference up to 15x faster. When Andrew Feldman and four co-founders started it in 2015, it was not clear the thing was even manufacturable.
That wager is now a public one. Cerebras listed on the Nasdaq in May 2026 under the ticker CBRS, popped 68% on day one to a roughly $95 billion valuation, and pulled in about $5.55 billion, the biggest US tech IPO since Uber. Revenue hit $510 million last year on 76% growth, with the company swinging to an $88 million profit. The catch sits in plain sight in the filing: for years a single customer bought almost everything Cerebras made.
What They Do: Wafer-scale silicon, sold three ways
Cerebras designs the largest computer chip ever built and wraps it in systems and cloud services that run AI models fast. The pitch is speed: because everything sits on one piece of silicon, there is no slow chip-to-chip interconnect to fight, so training and inference run in a single latency budget where GPU clusters need many.
You can buy it three ways. Rent it by the token through the inference and training cloud with drop-in OpenAI-compatible APIs, take dedicated private capacity, or install a CS-3 supercomputer in your own data center for full control of models and data.
Problems: When AI is too slow to feel useful
The problem Cerebras keeps pointing at is latency. On GPUs, a coding agent can take 20 to 30 seconds to generate, long enough to break a developer's concentration and force a context switch, and reasoning models that think out loud pay that tax on every step. Multi-step agents stall or time out; voice assistants feel robotic when the pause is a beat too long.
Cerebras reframes speed as a design parameter rather than a nice-to-have. Run a model at around 1,000 tokens per second and you can afford more reasoning, deeper retrieval, and tighter agent loops inside the same wall-clock budget, which is the argument behind its work with coding, search, and drug-discovery customers.
How it Happens
Who It's For: Built for the people who run out of patience with GPU inference
This is a B2B story aimed at three rooms. AI-native product teams who live or die on token speed, from coding agents to real-time voice. Large enterprises and the Global 1000 that want frontier models at production scale. And research and government labs doing physics, genomics, and drug discovery that a wafer-scale system can crunch in a fraction of GPU time.
The common buyer is technical: infrastructure and ML leads choosing where inference runs, and researchers who treat a hundreds-times speedup as the difference between months and years.
Ideal Customer Profiles
- Coding agents that stall on GPU inference
- Agent loops that time out
- Choosing between fast-but-dumb and smart-but-slow models
- High GPU-cloud inference costs
- Scaling frontier models to production latency
- Getting more reasoning inside a fixed latency budget
- Simulations that take years on GPU clusters
- Field-equation and genomic modeling throughput
- Drug-response prediction speed
Products: One giant chip, a supercomputer, and a cloud to rent it
The lineup all traces back to the same piece of silicon. The Wafer-Scale Engine is the chip; the CS-3 is the system built around it; the Condor Galaxy network is what you get when you wire a lot of them together. The cloud and code tiers are how everyone who cannot install a supercomputer still gets to use one.
Business Model: Sell the supercomputer, rent the tokens
Cerebras makes money two ways. It sells and deploys wafer-scale systems to enterprises, labs, and cloud partners, the high-ticket contract side that has driven most of its revenue so far. And it rents inference by the token through a pay-as-you-go cloud, with a free tier to get developers hooked, a self-serve Developer tier starting at $10, and enterprise deals for guaranteed throughput and custom model weights.
The pricing page splits neatly: raw inference API access on one side, a Cerebras Code subscription for developers on the other. Everything above the developer tier is a sales conversation.
Cerebras splits pricing into inference API access and a Cerebras Code subscription. Inference has a free tier for getting started, a self-serve Developer tier that begins at $10 with higher rate limits, and custom Enterprise contracts for guaranteed throughput and custom model weights. Cerebras Code is a flat monthly subscription for developers, sized by how many tokens per day you can send.
Plans
Developers getting started · Fastest way to try Cerebras inference
- Access to all Cerebras-powered models
- Community support via Discord
Power users · Generous rate limits, self-serve
- Everything in Free
- 10x higher rate limits than free tier
- Higher priority processing
- Self-serve payment
Production workloads · Highest throughput and guaranteed uptime
- Highest rate limits for production
- Lowest latency with dedicated queue priority
- Support for custom model weights
- Model fine-tuning and training services
- Dedicated support with response-time SLAs
Indie devs and weekend projects · Fast, high-context completions
- Top open-source model access
- Up to 24 million tokens/day
- Simple agentic workflows
Full-time developers · Heavy coding workflows
- Top open-source model for heavy use
- Up to 120 million tokens/day
- IDE integrations, refactoring, multi-agent systems
Good to know
- Developer-tier model pricing is per million tokens (e.g. Gemma 4 31B at $0.99/M input, $1.49/M output)
- Preview models are for evaluation only, not production
- Also available through partner APIs: AWS Marketplace, OpenRouter, Hugging Face, and Vercel
Competition: Racing Nvidia on the one axis Nvidia does not own outright: speed
Cerebras is not trying to out-volume Nvidia; it is trying to be the fastest place to run inference, full stop. The moat is architectural. Wafer-scale keeps the whole model on one chip with enormous on-chip memory, so it sidesteps the interconnect bottleneck that GPU clusters have to engineer around.
The strategic bet for 2026 is inference at scale: build out data centers across North America and Europe, ride partnerships with OpenAI, AWS, Meta, and the developer ecosystem, and become the default when speed is the constraint. Whether that holds as Nvidia, Groq, and custom silicon all chase the same latency prize is the open question.
Competes with
Their edge
Where they're betting
- Become the #1 provider of high-speed AI inference
- Data-center buildout across North America and Europe
- Deepen cloud partnerships (OpenAI, AWS)
- Broaden model support and developer ecosystem
Proof: The names on the wall, and the numbers behind them
The customer quotes read like a who's who of AI: OpenAI calls Cerebras its dedicated low-latency inference option, Meta runs its Llama API on it, and Cognition, Notion, AlphaSense, GSK, and Mayo Clinic all show up with concrete wins. Cognition reports its SWE-1.6 coding agent runs up to 5x faster on Cerebras than on GPU, hitting around 950 tokens per second.
The research proof is older and harder to argue with. Argonne National Laboratory won a Gordon Bell Special Prize for COVID-19 variant work on a CS-2 cluster, and NETL clocked a CS-2 at nearly 500x faster than its Joule supercomputer on field-equation modeling.
What People Say: Genuinely fast, genuinely dependent on a couple of customers
The praise is consistent and mostly technical: the speed is real, the OpenAI-compatible API makes it easy to drop in, and researchers cite orders-of-magnitude gains on the right workloads. Independent benchmarks back the direction even where they land below the headline 15x.
The recurring worry is not the technology but the customer list. Analysts keep flagging revenue concentration: G42 was most of the business, and even as that fell in 2025, a small set of Abu Dhabi entities and now OpenAI still account for the bulk of sales. As one analyst put it, the concentration rotated, it did not go away. The other knock is that squeezing the full speedup out of the hardware means working in Cerebras's own tooling, not the familiar GPU stack.
Widely seen as genuinely fast and technically impressive, with the loudest concern being customer and geographic concentration rather than the hardware itself.
Cerebras adds a dedicated low-latency inference solution to our platform. That means faster responses, more natural interactions, and a stronger foundation to scale real-time AI to many more people.
- Genuinely fast inference, backed by independent benchmarks
- Drop-in OpenAI-compatible API is easy to adopt
- Orders-of-magnitude gains on scientific and research workloads
- Marquee customer roster (OpenAI, Meta, AWS)
- Heavy revenue concentration in a few customers (G42, UAE entities, now OpenAI)
- Extracting the full speedup often means working in Cerebras's own tooling rather than the familiar GPU stack
- Post-IPO stock volatility as the market debates the inference bet
Funding: $5.55B raised at IPO, and a $95B debut it has been giving back since
Before going public Cerebras had raised around $2.7 billion privately, backed by an unusually operator-heavy cap table: Sam Altman, Ilya Sutskever, Andy Bechtolsheim, Adam D'Angelo, Intel CEO Lip-Bu Tan, and former AMD and Juniper leaders among them. G42 layered on more than $1.4 billion in commitments.
The May 2026 IPO priced at $185, opened at $350, and closed the first day up 68% for roughly a $95 billion valuation, raising about $5.55 billion. The market has since cooled the stock hard, with the cap sliding toward $60 billion within weeks, which is the price of going public on a story the market is still arguing about.
Total raised
Valuation
Latest round
Backers
Outlook: The fastest inference, if the customer list broadens
The technology story is strong and getting stronger: real speed advantages, marquee customers, and a data-center buildout aimed at making Cerebras the default for high-speed inference. If inference demand keeps climbing, wafer-scale is well positioned to catch it.
The risk is the same one the IPO prospectus flagged. A business where a couple of counterparties drive most of the revenue is exposed to any one of them changing course, and OpenAI in particular is a customer with its own shifting chip strategy. The next chapter is whether Cerebras can turn a handful of enormous contracts into a broad, durable book of business.
Team & Culture: A founding bench of chip veterans still running the place
Cerebras is unusual in that most of its five co-founders are still in the building a decade later, with Andrew Feldman as CEO and Sean Lie, Jean-Philippe Fricker, and Michael James in senior architecture roles. The company frames itself as a place for people who want to work on genuinely new hardware, publish and open-source their research, and skip the bureaucracy, and its own pitch to engineers leans on job stability with startup vitality.
The work is deep-systems work: chip design, compilers for a novel dataflow language, ML frameworks, and the infrastructure to run some of the fastest AI machines on earth. Hiring runs across engineering, research, product, and operations, concentrated in Sunnyvale with teams in San Diego, Toronto, and Bangalore.
- Values
- Extraordinary people, breakthrough innovation, global impact, Research-first: engineers can publish and open-source their work, Low-overhead, low-bureaucracy teams, Job stability with startup vitality
- Work policy
- Hybrid/office, primarily Sunnyvale with additional sites in San Diego, Toronto, and Bangalore; some roles open across the US/Canada.
- Hiring
- Hiring across engineering, research, product, and operations, concentrated in Sunnyvale with teams in San Diego, Toronto, and Bangalore, plus roles in the US/Canada and the UAE.
- Backend
- Python, C++, C, Go, Perl, SQL
- Chip Design & Silicon
- RTL, Synopsys, Calibre, ICV, DRC, LVS, Tcl, Computer architecture
- Compilers
- Tungsten dataflow language, Compiler passes, Code generation, Instruction scheduling, Register allocation
- Infrastructure
- Kubernetes, Docker, Jenkins, CI/CD, Ansible, Terraform, GitOps, Jinja2, MAAS, Foreman, Linux, Red Hat, Rocky Linux
- Networking
- BGP, VXLAN, EVPN, RoCEv2, DCQCN, PFC, ECN, TCP/IP, Ethernet, RDMA, PCIe, gNMI, Cisco, Juniper
- Cloud & Observability
- AWS, Prometheus, Grafana
- AI/ML
- PyTorch, TensorFlow, vLLM, ML frameworks
Engineering culture at Cerebras Systems
- Deep-systems work across chip design, compilers, ML frameworks, and infrastructure
- Small specialist teams (e.g. the Tungsten compiler team) working on genuinely novel hardware
- Emphasis on building tools, not just runbooks
Benefits & perks
- Premium medical, dental, and vision coverage
- Life insurance
- 401(k) (US)
- Group RRSP (Canada)
- Generous vacation policy
- Daily catered meals
- Healthy snacks
- Family-friendly events, including the CEO's BBQ
- Equity on most roles
- Bonus eligibility on most roles
Open roles · 93
View all roles →Cerebras Systems is hiring 93 roles across software engineers, operations, product managers, other roles, and more.
Compensation: Six-figure base bands, with equity and bonus layered on
Disclosed pay clusters in engineering, where posted US base salaries run from roughly $120,000 to $300,000 depending on level, with a few new-grad roles landing in the mid-$100,000s. A handful of non-engineering and design roles disclose bands in a similar range.
These are base numbers only. Cerebras notes that actual compensation for most roles adds bonus and equity, and as a newly public company that equity is now liquid stock rather than paper, which materially changes the math for candidates.
Most roles add bonus and equity on top of base; as a newly public company (Nasdaq: CBRS) that equity is now liquid stock rather than private paper.
Security & Legal: Cerebras Systems Inc., out of Sunnyvale
The registered entity is Cerebras Systems Inc., operating from 1237 E. Arques Ave, Sunnyvale, CA 94085, per its own privacy and terms pages effective August 2024. The company publishes a CCPA disclosure and a subprocessors reference, and states it does not retain the inputs and outputs of its training, inference, and chatbot services.
One nuance worth knowing: for API use, ownership of prompts and outputs is governed by the terms of whichever third-party model you run, not by Cerebras alone.
Legal entity
Registered address
Data residency
Certifications
Data practices
In the News: An IPO, a $20B OpenAI deal, and an AWS tie-up in one stretch
The past year has been loud. Cerebras pulled off the biggest US tech IPO since Uber, signed a multi-year deal with OpenAI to deploy 750 megawatts of wafer-scale systems reported at more than $20 billion, and struck an AWS partnership pairing its CS-3 systems with AWS Trainium. The links below are the source trail.
Cerebras pops 68% in Nasdaq debut, pushing the AI chipmaker's market cap to $95 billion
OpenAI Partners with Cerebras to bring high-speed inference to the mainstream
Cerebras Systems (CBRS) Lands $20 Billion OpenAI And AWS Partnerships After IPO
Cerebras and AWS collaborate to deliver fastest AI inference
Case Study: Cognition x Cerebras, the dawn of real-time coding agents
Cerebras Raises $5.55 Billion in AI Chip IPO, But 86% Revenue Dependence on UAE Entities Unresolved
Cerebras Systems priced its initial public offering
More in Artificial Intelligence
Other companies hiring in the same space.

OpenAI (710 jobs)
Builds frontier AI models and ships them as consumer, developer, and enterprise products — ChatGPT, the API platform, and Codex.

Harvey (331 jobs)
Domain-specific AI for legal and professional services that automates research, drafting, contract analysis, and due diligence.

Applied Intuition (262 jobs)
Applied Intuition builds the software and digital infrastructure that brings physical AI (autonomous driving and robotics) to every moving machine, from cars and trucks to drones and defense platforms.

Legora (229 jobs)
Legora builds a collaborative, agentic AI workspace that helps lawyers review, research, draft, and advise faster.

Sierra (175 jobs)
Enterprise AI platform for building branded customer-service agents that resolve conversations across chat, voice, and messaging.

ElevenLabs (174 jobs)
AI research and product company building foundational audio models for voice synthesis, conversational agents, and creative media generation.

Mistral (151 jobs)
A French AI lab building open and frontier-grade large language models, with the full developer and enterprise stack around them.

SKELAR (134 jobs)
Ukrainian venture builder that co-founds and scales global consumer tech companies, backing each with capital, a shared operating platform, and a network of operators.

Cohere (128 jobs)
Enterprise AI company building secure, privately deployable foundation models and an agentic workspace (North) for regulated businesses.