

Etched
Transformer-specific ASICs and rack-scale systems that hardwire the transformer into silicon for faster, cheaper AI inference.

Overview: The startup that bet the whole company on one AI architecture
Etched made a wager most chip companies would never dare: that the transformer, the architecture behind every major AI model, would win so completely that you could burn it straight into silicon and never look back. Its Sohu chip does exactly that. Instead of the flexible general-purpose logic a GPU carries, Sohu hardwires transformer inference into the metal, trading versatility for raw speed and efficiency.
The bet nearly killed the company first. Founders Gavin Uberti and Chris Zhu pitched a detailed memo on chip specialization in 2023 and watched every major investor pass, running month to month on fumes. Then the AI market consolidated around transformers exactly as they predicted. By mid 2026 Etched had A0 silicon back from TSMC, more than $1 billion in booked customer contracts, and an $800 million war chest. If transformers stay on top, this is one of the most valuable hardware companies in the world. If the field moves on, Sohu is an expensive paperweight, and Etched has said as much.
What They Do: Frontier inference clusters, co-designed from the transistor to the token
Etched does not just sell a chip. It sells the whole rack. The company co-designs chips, racks, software, and even manufacturing methods so frontier models can run inference with better throughput, latency, cost, and power than general-purpose hardware, across both the prefill and decode halves of a request.
The pitch is vertical integration taken to an extreme. Math block designers sit next to inference engineers, thermal experts next to supply-chain managers, all aimed at one goal the company repeats like a mantra: get to gigawatt scale as fast as possible. Its systems target the hardest workloads out there, many-trillion-parameter mixture-of-experts models, long context, and agentic runs, the exact places where GPUs choke.
Problems: Why GPUs waste most of their silicon on transformer inference
The core problem Etched attacks is waste. On a general-purpose GPU, transformer inference typically uses only 30 to 40 percent of peak FLOPs, because the chip carries logic for workloads it isn't running and throttles its clock as power and heat climb. Etched calls out two specific bottlenecks and claims to have engineered around both.
Low Voltage Inference runs the chip's math blocks at under half the voltage of typical AI chips, which the company says lets it sustain 80 percent or more of peak FLOPs on trillion-parameter sparse models without thermal throttling. Cluster Scale Memory tackles the other wall, memory latency, with a proprietary low-latency interconnect and an HBM/SRAM hybrid that aims to hit SRAM-class decode speeds without giving up capacity. Both are claims, not yet independently benchmarked, but they name the real pain precisely.
How it Happens
Who It's For: Built for the labs and hyperscalers running inference at industrial scale
Etched is not selling to hobbyists or small teams. Its buyers are the frontier AI labs, cloud providers, and hyperscalers running transformer inference at a scale where a few percent of efficiency translates into millions of dollars and megawatts. The company says it made co-design decisions hand in hand with leading AI companies and cloud providers, and ran terabytes of production traffic through its simulator before committing to silicon.
That focus is also the constraint. A turnkey inference cluster that ships by the rack and books contracts in the hundreds of millions is a product for a very short list of customers, the ones already spending at data-center scale on AI.
Ideal Customer Profiles
- Cost and power ceiling of GPU inference at scale
- Sourcing turnkey inference capacity fast
- Serving many-trillion-parameter MoEs and long context economically
- Latency on decode-heavy agentic workloads
Products: One chip, one system, one architecture on purpose
Everything Etched builds points at a single idea: run transformers, run them fast, skip everything else. The Sohu chip is the wager made physical, a transformer-only ASIC fabbed on TSMC's N4P process. Around it the company wraps a full rack-scale system, custom PCBs, cold plates, interconnects, and power delivery, plus the compiler and runtime software to drive it.
The throughput claims are eye-watering. Etched says an eight-chip Sohu server can push more than 500,000 tokens per second on Llama 70B, against roughly 23,000 on eight H100s. No third party has verified those numbers yet, but the roadmap is public: first racks shipping in summer 2026, with performance updates promised alongside.
Business Model: Sell the rack, book the contract, ramp the factory
Etched makes money the old-fashioned hardware way: it sells complete inference systems under large forward contracts, not chips off a shelf. Customers order full clusters, and the company has already booked more than $1 billion in those orders, which it is now racing to fulfill as its first racks ship.
There is no public pricing and no self-serve tier, which fits the buyer. This is enterprise capital equipment sold direct, with a Taiwan factory, a data center, and an in-house prototyping lab all stood up to turn that order book into shipped silicon.
enterprise
Competition: Taking on NVIDIA by giving up everything NVIDIA does well
Etched is a direct shot at NVIDIA's grip on AI compute, but it competes by inverting NVIDIA's strategy. Where a GPU is a Swiss Army knife, Sohu is a single blade, useless for anything but transformers and, Etched argues, far better at that one thing. Its rack-scale system is pitched as a turnkey alternative to NVIDIA's DGX and HGX.
The edge and the risk are the same fact. A transformer-only chip can devote all its transistors to the workload that matters today, but if the industry pivots to a new architecture, the company estimates it would take about three years to respond. Etched is betting the leadership pedigree, engineers from NVIDIA, Google's TPU program, Broadcom, and TSMC, plus a head start on transformer-specific design, buys enough time for the bet to pay off.
Competes with
Their edge
Where they're betting
- Ship first racks and ramp to volume against $1B in contracts
- Reach gigawatt-scale deployment as fast as possible
- Deepen the TSMC manufacturing partnership
- Prove SOTA throughput/latency/power in real deployments
Proof: What Etched can actually point to today
The concrete milestones are real even where the benchmarks aren't. A0 silicon came back from TSMC's N4P process earlier in 2026, the company is validating its first rack-scale product with customers, and it has booked over $1 billion in contracts against those systems. To run engineering around the clock, it opened a Taiwan factory and built a data center, test house, and prototyping lab inside its San Jose office.
The caution worth stating plainly: the headline performance figures come from Etched, not from independent testing, and its first racks are only shipping now. The order book and the working silicon are the strongest proof on the table.
What People Say: A bold bet that observers admire and distrust in equal measure
Etched draws a rare mix of excitement and open skepticism. On Hacker News and in the ML community, people credit the clarity of the thesis and the sheer speed advantage on paper, the idea that stripping out general-purpose overhead could free 2 to 3x more useful compute from the same transistors.
The doubts are just as consistent. Critics note that none of the performance claims have been independently verified, that early materials leaned on renders rather than shipped chips, and that betting the company on transformers is a single point of failure if the architecture ever falls out of fashion. It is admired as a gutsy call and distrusted as an unproven one, often in the same breath.
Widely seen as one of the boldest bets in AI hardware, admired for the thesis and distrusted for the unverified claims and transformer-only risk.
- Clear, gutsy thesis that AI will consolidate on transformers
- Large theoretical speed and efficiency advantage over GPUs
- Elite engineering and investor bench
- Performance claims not independently verified
- Early materials leaned on renders rather than shipped chips
- Single point of failure: betting the company entirely on transformers
Funding: $800M raised, a $5B valuation, and a near-miss that came first
Etched has raised $800 million in total, and reached a $5 billion post-money valuation on a $500 million round led by Stripes that closed in late 2025 and surfaced publicly in January 2026, with Peter Thiel and Ribbit Capital among the participants. The company also touts a strategic investment from VentureTech Alliance, tied to its TSMC relationship.
What makes the number striking is where it started. In 2023 the founders could not get a single major investor to bite. The cap table now reads like a who's who of AI, angels and backers including Andrej Karpathy, Geoffrey Hinton, Fei-Fei Li, Arthur Mensch, and Stanley Druckenmiller, plus quant firms Jane Street, Hudson River Trading, and Two Sigma.
Total raised
Valuation
Latest round
Backers
Outlook: Everything now rides on shipping racks that hit the numbers
Etched has done the hard part of a hardware startup, working silicon, a real order book, and a factory to build against it. What it hasn't done is prove the performance claims in the open or ship at volume, and both come due now that first racks are leaving the dock in summer 2026.
The next year is the whole game. Independent benchmarks that back the token-per-second numbers would validate the entire thesis. A transformer-specific chip that lands as promised turns $1 billion in contracts into a durable business. The one risk it cannot engineer away is the one it chose on purpose: if the models the world runs stop being transformers, none of the rest matters.
Team & Culture: Two Harvard dropouts and a bench of silicon veterans, all in-office
Etched was founded in 2022 by Gavin Uberti and Chris Zhu, math students and Thiel Fellows who walked away from Harvard, with Robert Wachen as co-founder and president. Around them sits an unusually senior bench: a CTO who ran Cypress, a VP of Platform who spent 22 years at NVIDIA building HGX and DGX, a VP of Software who built Google's TPU software team across five generations, and engineers pulled from NVIDIA, Google, Broadcom, SK Hynix, and TSMC.
The culture is deliberate and demanding. Every technical hire is expected to work across engineering and research with no boundary between them, and the team is fully in person in San Jose at Santana Row, with a second hub in Taipei next to its manufacturing partners. The company's own values are blunt: own outcomes end to end with no training wheels, and treat production, not the prototype, as the real product.
- Values
- End-to-end ownership with no training wheels, Production is the product, No boundary between engineering and research; technical staff work across disciplines, Fully in-person, Extreme vertical integration and co-design
- Work policy
- Fully in-person in San Jose (Santana Row), with an office in Taipei, Taiwan; a small number of roles are remote.
- Hiring
- Hiring aggressively across silicon (ASIC design, verification, architecture), systems and infrastructure software, security/IT, data center, and operations. Mostly on-site in San Jose (Santana Row), with roles in Taipei, Taiwan.
- Silicon / ASIC
- SystemVerilog, Verilog, UVM, RTL, Digital Design, Synopsys, Cadence, Verilator, Design Verification
- Systems / Firmware
- C, C++, Rust, PCIe, I2C, SPI, Ethernet, DRAM, Kernel development, eBPF, Linux internals
- Software
- Python, Go, Rust, Bash, Git, TensorRT-LLM, vLLM
- Infrastructure
- Linux, Kubernetes, Docker, SLURM, Terraform, Ansible, Puppet, Bazel, AWS, GCP, CI/CD
- Observability & Security
- Prometheus, Grafana, VictoriaMetrics, Okta, FreeIPA, Fortinet, Arista, EDR/XDR
Engineering culture at Etched
- No boundary between engineering and research; expected to contribute to both
- Treat infrastructure and data center work as an engineering discipline, from first principles
- High ownership: independently scope and drive projects to completion
ASIC / Silicon culture at Etched
- Co-design across chip, package, PCB, and system
- Same rigor, testing, and design discipline applied to infra as to the ASIC itself
- Feed characterization data directly back into RTL and physical design decisions
Benefits & perks
- Medical, dental, and vision packages with generous premium coverage
- $500 per month credit for waiving medical benefits
- Various wellness benefits covering fitness, mental health, and more
- Housing subsidy of $2,000 per month for those living within walking distance of the office
- Relocation support for those moving to San Jose (Santana Row)
- Daily lunch and dinner in the office
- Competitive compensation with generous equity/option packages
- Unlimited compute budget subject to ROI justification
Open roles · 107
View all roles →Etched is hiring 107 roles across software engineers, operations, other roles, designers, and more.
Compensation: Bay Area silicon pay, equity-heavy, with perks aimed at the office
Disclosed pay clusters where you'd expect for a well-funded Bay Area hardware startup. Across the many engineering roles that list a band, base salaries generally run from roughly $130k to $275k, with a handful of specialized and leadership roles reaching to about $300k. Comp is quoted in US dollars, matching Etched's San Jose center of gravity.
Equity is a real part of the package, with roles citing generous equity or option grants on top of base. The benefits lean toward keeping an in-person team fed and close to the office: full medical, dental, and vision, a $2,000 per month housing subsidy for people within walking distance, relocation help to Santana Row, and daily lunch and dinner in the office.
Roles cite generous equity or option grants on top of base; comp is described as competitive for the AI hardware space and equity-heavy.
Security & Legal: A Delaware chipmaker guarding irreplaceable design IP
The registered entity is Etched, Inc., a Delaware corporation, operating from 3155 Olsen Dr, San Jose, CA 95117. Its published terms lean enterprise-standard and cautious, with a mandatory arbitration clause, a class-action and jury-trial waiver, and website access granted only for internal evaluation.
Security is treated as existential rather than boilerplate, which fits a company whose crown jewels are chip design files. Job postings describe zero-trust network architecture, segmentation that isolates ASIC development from the rest of the network, and DLP controls built to keep design IP from leaking, with SOC 2 and ISO 27001 named as the frameworks it works toward. Etched does not publish a subprocessors page, so its vendor list isn't public.
Legal entity
Registered address
Data residency
Certifications
Data practices
In the News: From stealth to a $5B headline in six months
Etched spent years heads-down and then arrived loudly. The story that carried in mid 2026 was the same set of facts told two ways: a company emerging from stealth with working silicon and a billion-dollar order book, and an NVIDIA challenger reaching a $5 billion valuation. The $500 million Stripes round was reported earlier in the year and set up the moment.
Nvidia competitor Etched hits $5B valuation, $1B in sales for AI chip
Etched Emerges From Stealth With Working Chip, $800M Raised, and Over $1B in Customer Contracts
Etched.ai raises $500m for a $5bn valuation – report
AI Chip Startup Etched Raises $500 Million to Take on Nvidia
Sohu: The First Transformer ASIC | Hacker News
More in Artificial Intelligence
Other companies hiring in the same space.

OpenAI (710 jobs)
Builds frontier AI models and ships them as consumer, developer, and enterprise products — ChatGPT, the API platform, and Codex.

Harvey (331 jobs)
Domain-specific AI for legal and professional services that automates research, drafting, contract analysis, and due diligence.

Applied Intuition (262 jobs)
Applied Intuition builds the software and digital infrastructure that brings physical AI (autonomous driving and robotics) to every moving machine, from cars and trucks to drones and defense platforms.

Legora (229 jobs)
Legora builds a collaborative, agentic AI workspace that helps lawyers review, research, draft, and advise faster.

Sierra (175 jobs)
Enterprise AI platform for building branded customer-service agents that resolve conversations across chat, voice, and messaging.

ElevenLabs (174 jobs)
AI research and product company building foundational audio models for voice synthesis, conversational agents, and creative media generation.

Mistral (151 jobs)
A French AI lab building open and frontier-grade large language models, with the full developer and enterprise stack around them.

SKELAR (134 jobs)
Ukrainian venture builder that co-founds and scales global consumer tech companies, backing each with capital, a shared operating platform, and a network of operators.

Cohere (128 jobs)
Enterprise AI company building secure, privately deployable foundation models and an agentic workspace (North) for regulated businesses.
Backed by Ribbit Capital
Companies that share an investor.

Crusoe (363 jobs)
Energy-first AI infrastructure company that sources power, builds hyperscale AI data centers, and runs a GPU cloud purpose-built for AI workloads.

Fluidstack (196 jobs)
Builds and operates gigawatt-scale AI data centers and fully managed GPU clusters for frontier AI labs, governments, and enterprises.

Checkout.com (175 jobs)
Global enterprise payments platform that helps large merchants accept, move, protect, and optimize money through one API.

monday.com (172 jobs)
A no-code work platform, now built around AI agents, that lets any team shape its own projects, CRM, dev, and IT workflows in one place.

Base Power Company (162 jobs)
Base Power installs home batteries and sells below-market energy in Texas, earning its margin from the grid instead of your bill.

Decagon (115 jobs)
Autonomous AI agents that resolve enterprise customer support end to end across chat, voice, email, and SMS.

Plaid (107 jobs)
Plaid provides the data network and APIs that let apps securely connect to users' bank accounts.

Mach Industries (91 jobs)
Builds low-cost autonomous aircraft, precision munitions, and high-altitude platforms for the U.S. military and its allies.

Cognition (75 jobs)
Applied AI lab behind Devin, the autonomous AI software engineer, and the Windsurf IDE.