Etched website
Etched

Etched

Transformer-specific ASICs and rack-scale systems that hardwire the transformer into silicon for faster, cheaper AI inference.

Careers(107)
Etched website preview

Overview: The startup that bet the whole company on one AI architecture

Etched made a wager most chip companies would never dare: that the transformer, the architecture behind every major AI model, would win so completely that you could burn it straight into silicon and never look back. Its Sohu chip does exactly that. Instead of the flexible general-purpose logic a GPU carries, Sohu hardwires transformer inference into the metal, trading versatility for raw speed and efficiency.

The bet nearly killed the company first. Founders Gavin Uberti and Chris Zhu pitched a detailed memo on chip specialization in 2023 and watched every major investor pass, running month to month on fumes. Then the AI market consolidated around transformers exactly as they predicted. By mid 2026 Etched had A0 silicon back from TSMC, more than $1 billion in booked customer contracts, and an $800 million war chest. If transformers stay on top, this is one of the most valuable hardware companies in the world. If the field moves on, Sohu is an expensive paperweight, and Etched has said as much.

What They Do: Frontier inference clusters, co-designed from the transistor to the token

Etched does not just sell a chip. It sells the whole rack. The company co-designs chips, racks, software, and even manufacturing methods so frontier models can run inference with better throughput, latency, cost, and power than general-purpose hardware, across both the prefill and decode halves of a request.

The pitch is vertical integration taken to an extreme. Math block designers sit next to inference engineers, thermal experts next to supply-chain managers, all aimed at one goal the company repeats like a mantra: get to gigawatt scale as fast as possible. Its systems target the hardest workloads out there, many-trillion-parameter mixture-of-experts models, long context, and agentic runs, the exact places where GPUs choke.

Problems: Why GPUs waste most of their silicon on transformer inference

The core problem Etched attacks is waste. On a general-purpose GPU, transformer inference typically uses only 30 to 40 percent of peak FLOPs, because the chip carries logic for workloads it isn't running and throttles its clock as power and heat climb. Etched calls out two specific bottlenecks and claims to have engineered around both.

Low Voltage Inference runs the chip's math blocks at under half the voltage of typical AI chips, which the company says lets it sustain 80 percent or more of peak FLOPs on trillion-parameter sparse models without thermal throttling. Cluster Scale Memory tackles the other wall, memory latency, with a proprietary low-latency interconnect and an HBM/SRAM hybrid that aims to hit SRAM-class decode speeds without giving up capacity. Both are claims, not yet independently benchmarked, but they name the real pain precisely.

How it Happens

GPUs use only ~30-40% of peak FLOPs on transformer inference due to general-purpose overhead
AI chips throttle clock speed as FLOPs utilization and power rise, cutting sustained throughput
HBM-based chips can't reach SRAM-level decode latency; SRAM-only chips sacrifice FLOPs density and capacity
Runaway cost and power of large-scale transformer inference (MoEs, long context, agentic workloads)

Who It's For: Built for the labs and hyperscalers running inference at industrial scale

Etched is not selling to hobbyists or small teams. Its buyers are the frontier AI labs, cloud providers, and hyperscalers running transformer inference at a scale where a few percent of efficiency translates into millions of dollars and megawatts. The company says it made co-design decisions hand in hand with leading AI companies and cloud providers, and ran terabytes of production traffic through its simulator before committing to silicon.

That focus is also the constraint. A turnkey inference cluster that ships by the rack and books contracts in the hundreds of millions is a product for a very short list of customers, the ones already spending at data-center scale on AI.

Ideal Customer Profiles

AI infrastructure leaders
  • Cost and power ceiling of GPU inference at scale
  • Sourcing turnkey inference capacity fast
Frontier AI labs
  • Serving many-trillion-parameter MoEs and long context economically
  • Latency on decode-heavy agentic workloads

Products: One chip, one system, one architecture on purpose

Everything Etched builds points at a single idea: run transformers, run them fast, skip everything else. The Sohu chip is the wager made physical, a transformer-only ASIC fabbed on TSMC's N4P process. Around it the company wraps a full rack-scale system, custom PCBs, cold plates, interconnects, and power delivery, plus the compiler and runtime software to drive it.

The throughput claims are eye-watering. Etched says an eight-chip Sohu server can push more than 500,000 tokens per second on Llama 70B, against roughly 23,000 on eight H100s. No third party has verified those numbers yet, but the roadmap is public: first racks shipping in summer 2026, with performance updates promised alongside.

Sohu
A transformer-only ASIC fabbed on TSMC's N4P process, hardwiring transformer inference into silicon. Etched claims an eight-chip Sohu server exceeds 500,000 tokens/sec on Llama 70B, versus roughly 23,000 on eight H100s (company figures, not independently verified).
Frontier Inference Clusters
Full rack-scale systems that pair Sohu with custom PCBs, cold plates, interconnects, and power delivery. Built for many-trillion-parameter MoEs, long context, and agentic workloads across both prefill and decode.
Low Voltage Inference (LVI)
An architecture that runs the chip's math blocks at under half the voltage of typical AI chips, which the company says sustains 80%+ of peak FLOPs on trillion-parameter sparse MoEs without thermal throttling.
Cluster Scale Memory (CSM)
A low-latency shared memory pool across the scale-up domain using a proprietary interconnect and an HBM/SRAM hybrid design, aimed at SRAM-class decode speeds without sacrificing memory capacity.

Business Model: Sell the rack, book the contract, ramp the factory

Etched makes money the old-fashioned hardware way: it sells complete inference systems under large forward contracts, not chips off a shelf. Customers order full clusters, and the company has already booked more than $1 billion in those orders, which it is now racing to fulfill as its first racks ship.

There is no public pricing and no self-serve tier, which fits the buyer. This is enterprise capital equipment sold direct, with a Taiwan factory, a data center, and an in-house prototyping lab all stood up to turn that order book into shipped silicon.

Pricing

enterprise

Competition: Taking on NVIDIA by giving up everything NVIDIA does well

Etched is a direct shot at NVIDIA's grip on AI compute, but it competes by inverting NVIDIA's strategy. Where a GPU is a Swiss Army knife, Sohu is a single blade, useless for anything but transformers and, Etched argues, far better at that one thing. Its rack-scale system is pitched as a turnkey alternative to NVIDIA's DGX and HGX.

The edge and the risk are the same fact. A transformer-only chip can devote all its transistors to the workload that matters today, but if the industry pivots to a new architecture, the company estimates it would take about three years to respond. Etched is betting the leadership pedigree, engineers from NVIDIA, Google's TPU program, Broadcom, and TSMC, plus a head start on transformer-specific design, buys enough time for the bet to pay off.

Competes with

NVIDIA (DGX / HGX, H100, B200)GroqCerebrasSambaNovaAMD (Instinct)Google TPU

Their edge

Architecture-specialized silicon
By hardwiring transformers into the chip, Sohu can devote all its transistors to the workload that matters today, versus a GPU that wastes silicon on flexibility it isn't using for inference.
Turnkey rack, not just a chip
Etched ships complete co-designed inference clusters as a direct alternative to NVIDIA's DGX/HGX, capturing system-level value.
Deep silicon pedigree
Engineers and leaders from NVIDIA, Google's TPU program, Broadcom, SK Hynix, and TSMC, including people who built HGX/DGX and the TPU software stack.

Where they're betting

  • Ship first racks and ramp to volume against $1B in contracts
  • Reach gigawatt-scale deployment as fast as possible
  • Deepen the TSMC manufacturing partnership
  • Prove SOTA throughput/latency/power in real deployments

Proof: What Etched can actually point to today

The concrete milestones are real even where the benchmarks aren't. A0 silicon came back from TSMC's N4P process earlier in 2026, the company is validating its first rack-scale product with customers, and it has booked over $1 billion in contracts against those systems. To run engineering around the clock, it opened a Taiwan factory and built a data center, test house, and prototyping lab inside its San Jose office.

The caution worth stating plainly: the headline performance figures come from Etched, not from independent testing, and its first racks are only shipping now. The order book and the working silicon are the strongest proof on the table.

A0 silicon returned from TSMC N4P (2026)
Over $1B in booked customer contracts
$800M
raised across four financings
$5B
post-money valuation
Opened a Taiwan factory plus an in-house data center, test house, and prototyping lab in San Jose
First rack-scale product shipping summer 2026

What People Say: A bold bet that observers admire and distrust in equal measure

Etched draws a rare mix of excitement and open skepticism. On Hacker News and in the ML community, people credit the clarity of the thesis and the sheer speed advantage on paper, the idea that stripping out general-purpose overhead could free 2 to 3x more useful compute from the same transistors.

The doubts are just as consistent. Critics note that none of the performance claims have been independently verified, that early materials leaned on renders rather than shipped chips, and that betting the company on transformers is a single point of failure if the architecture ever falls out of fashion. It is admired as a gutsy call and distrusted as an unproven one, often in the same breath.

Widely seen as one of the boldest bets in AI hardware, admired for the thesis and distrusted for the unverified claims and transformer-only risk.

Loved
  • Clear, gutsy thesis that AI will consolidate on transformers
  • Large theoretical speed and efficiency advantage over GPUs
  • Elite engineering and investor bench
Gripes
  • Performance claims not independently verified
  • Early materials leaned on renders rather than shipped chips
  • Single point of failure: betting the company entirely on transformers

Funding: $800M raised, a $5B valuation, and a near-miss that came first

Etched has raised $800 million in total, and reached a $5 billion post-money valuation on a $500 million round led by Stripes that closed in late 2025 and surfaced publicly in January 2026, with Peter Thiel and Ribbit Capital among the participants. The company also touts a strategic investment from VentureTech Alliance, tied to its TSMC relationship.

What makes the number striking is where it started. In 2023 the founders could not get a single major investor to bite. The cap table now reads like a who's who of AI, angels and backers including Andrej Karpathy, Geoffrey Hinton, Fei-Fei Li, Arthur Mensch, and Stanley Druckenmiller, plus quant firms Jane Street, Hudson River Trading, and Two Sigma.

Total raised

$800M

Valuation

$5B

Latest round

$500M · led by Stripes · Dec 2025 (reported Jan 2026)

Backers

StripesPeter ThielRibbit CapitalVentureTech AllianceJane StreetHudson River TradingTwo SigmaAndrej KarpathyGeoffrey HintonFei-Fei LiArthur MenschScott WuStanley Druckenmiller

Outlook: Everything now rides on shipping racks that hit the numbers

Etched has done the hard part of a hardware startup, working silicon, a real order book, and a factory to build against it. What it hasn't done is prove the performance claims in the open or ship at volume, and both come due now that first racks are leaving the dock in summer 2026.

The next year is the whole game. Independent benchmarks that back the token-per-second numbers would validate the entire thesis. A transformer-specific chip that lands as promised turns $1 billion in contracts into a durable business. The one risk it cannot engineer away is the one it chose on purpose: if the models the world runs stop being transformers, none of the rest matters.

Team & Culture: Two Harvard dropouts and a bench of silicon veterans, all in-office

Etched was founded in 2022 by Gavin Uberti and Chris Zhu, math students and Thiel Fellows who walked away from Harvard, with Robert Wachen as co-founder and president. Around them sits an unusually senior bench: a CTO who ran Cypress, a VP of Platform who spent 22 years at NVIDIA building HGX and DGX, a VP of Software who built Google's TPU software team across five generations, and engineers pulled from NVIDIA, Google, Broadcom, SK Hynix, and TSMC.

The culture is deliberate and demanding. Every technical hire is expected to work across engineering and research with no boundary between them, and the team is fully in person in San Jose at Santana Row, with a second hub in Taipei next to its manufacturing partners. The company's own values are blunt: own outcomes end to end with no training wheels, and treat production, not the prototype, as the real product.

Values
End-to-end ownership with no training wheels, Production is the product, No boundary between engineering and research; technical staff work across disciplines, Fully in-person, Extreme vertical integration and co-design
Work policy
Fully in-person in San Jose (Santana Row), with an office in Taipei, Taiwan; a small number of roles are remote.
Hiring
Hiring aggressively across silicon (ASIC design, verification, architecture), systems and infrastructure software, security/IT, data center, and operations. Mostly on-site in San Jose (Santana Row), with roles in Taipei, Taiwan.
Silicon / ASIC
SystemVerilog, Verilog, UVM, RTL, Digital Design, Synopsys, Cadence, Verilator, Design Verification
Systems / Firmware
C, C++, Rust, PCIe, I2C, SPI, Ethernet, DRAM, Kernel development, eBPF, Linux internals
Software
Python, Go, Rust, Bash, Git, TensorRT-LLM, vLLM
Infrastructure
Linux, Kubernetes, Docker, SLURM, Terraform, Ansible, Puppet, Bazel, AWS, GCP, CI/CD
Observability & Security
Prometheus, Grafana, VictoriaMetrics, Okta, FreeIPA, Fortinet, Arista, EDR/XDR

Engineering culture at Etched

  • No boundary between engineering and research; expected to contribute to both
  • Treat infrastructure and data center work as an engineering discipline, from first principles
  • High ownership: independently scope and drive projects to completion

ASIC / Silicon culture at Etched

  • Co-design across chip, package, PCB, and system
  • Same rigor, testing, and design discipline applied to infra as to the ASIC itself
  • Feed characterization data directly back into RTL and physical design decisions

Benefits & perks

Health & wellness
  • Medical, dental, and vision packages with generous premium coverage
  • $500 per month credit for waiving medical benefits
  • Various wellness benefits covering fitness, mental health, and more
Living & relocation
  • Housing subsidy of $2,000 per month for those living within walking distance of the office
  • Relocation support for those moving to San Jose (Santana Row)
  • Daily lunch and dinner in the office
Compensation & work
  • Competitive compensation with generous equity/option packages
  • Unlimited compute budget subject to ROI justification

Compensation: Bay Area silicon pay, equity-heavy, with perks aimed at the office

Disclosed pay clusters where you'd expect for a well-funded Bay Area hardware startup. Across the many engineering roles that list a band, base salaries generally run from roughly $130k to $275k, with a handful of specialized and leadership roles reaching to about $300k. Comp is quoted in US dollars, matching Etched's San Jose center of gravity.

Equity is a real part of the package, with roles citing generous equity or option grants on top of base. The benefits lean toward keeping an in-person team fed and close to the office: full medical, dental, and vision, a $2,000 per month housing subsidy for people within walking distance, relocation help to Santana Row, and daily lunch and dinner in the office.

Engineering
$130,000$275,000 · yearly
based on many disclosed roles
Operations
$125,000$175,000 · yearly
based on a few disclosed roles
G&A
$100,000$220,000 · yearly
based on a few disclosed roles
Sales
$100,000$220,000 · yearly
based on a few disclosed roles

Roles cite generous equity or option grants on top of base; comp is described as competitive for the AI hardware space and equity-heavy.

In the News: From stealth to a $5B headline in six months

Etched spent years heads-down and then arrived loudly. The story that carried in mid 2026 was the same set of facts told two ways: a company emerging from stealth with working silicon and a billion-dollar order book, and an NVIDIA challenger reaching a $5 billion valuation. The $500 million Stripes round was reported earlier in the year and set up the moment.

More in Artificial Intelligence

Other companies hiring in the same space.

OpenAI

OpenAI (710 jobs)

710 jobs

Builds frontier AI models and ships them as consumer, developer, and enterprise products — ChatGPT, the API platform, and Codex.

Harvey

Harvey (331 jobs)

331 jobs

Domain-specific AI for legal and professional services that automates research, drafting, contract analysis, and due diligence.

Applied Intuition

Applied Intuition (262 jobs)

262 jobs

Applied Intuition builds the software and digital infrastructure that brings physical AI (autonomous driving and robotics) to every moving machine, from cars and trucks to drones and defense platforms.

Legora

Legora (229 jobs)

229 jobs

Legora builds a collaborative, agentic AI workspace that helps lawyers review, research, draft, and advise faster.

Sierra

Sierra (175 jobs)

175 jobs

Enterprise AI platform for building branded customer-service agents that resolve conversations across chat, voice, and messaging.

ElevenLabs

ElevenLabs (174 jobs)

174 jobs

AI research and product company building foundational audio models for voice synthesis, conversational agents, and creative media generation.

Mistral

Mistral (151 jobs)

151 jobs

A French AI lab building open and frontier-grade large language models, with the full developer and enterprise stack around them.

SKELAR

SKELAR (134 jobs)

134 jobs

Ukrainian venture builder that co-founds and scales global consumer tech companies, backing each with capital, a shared operating platform, and a network of operators.

Cohere

Cohere (128 jobs)

128 jobs

Enterprise AI company building secure, privately deployable foundation models and an agentic workspace (North) for regulated businesses.

Backed by Ribbit Capital

Companies that share an investor.