--- title: 'Head of Platform Product Reliability at Etched' canonical: 'https://feeny.ai/job/head-of-platform-product-reliability-etched-san-jose-dsg885yvdty5' type: 'job' last_seen: '2026-09-07' --- # Head of Platform Product Reliability at Etched - **Company:** [Etched](https://feeny.ai/companies/etched) - **Location:** San Jose, CA - **Employment:** full-time - **Work type:** onsite - **Posted:** 2026-05-12 - **Last confirmed live:** 2026-09-07 - **Apply:** https://jobs.ashbyhq.com/etched/fbade92c-43c8-4e8d-931a-be9c5ec27b5b ## Job description ## About Etched Etched is building hardware for frontier intelligence. We co-design chips, racks, software, and manufacturing to deliver best-in-class throughput and latency across both prefill and decode workloads. Our first products are heavily focused on inference. Backed by hundreds of millions from top-tier investors and staffed by leading engineers, Etched is redefining the infrastructure layer for the fastest growing industry in history. Job Summary We are seeking a highly technical and execution-focused Head of Platform Product Reliability to lead reliability engineering across Etched's server, rack, and datacenter platform products. This role owns system-level product reliability from architecture through fleet deployment. You will define reliability strategy, qualification methodologies, accelerated stress testing programs, failure analysis processes, and long-term reliability standards for complex AI infrastructure systems. This team focuses specifically on product reliability engineering for platform hardware and deployed systems — ensuring every Etched product ships with the reliability profile that enterprise and hyperscale customers demand. You will work cross-functionally with Platform Engineering, Mechanical Engineering, Thermal, Firmware, Manufacturing, Supply Chain, Datacenter Operations, and Program teams to ensure Etched products achieve exceptional reliability at scale. ## Key Responsibilities - Define and own the end-to-end reliability strategy for AI servers, accelerator platforms, rack systems, and datacenter infrastructure, from design requirements through field deployment - Establish reliability requirements, qualification standards, and validation methodologies that scale across product generations - Build and institutionalize reliability engineering processes spanning the full product lifecycle: - EVT / DVT / PVT qualification gates and exit criteria - Accelerated life testing (ALT) and accelerated stress testing (AST) - Environmental testing: temperature, humidity, altitude, contamination - HALT / HASS programs for design margin and production screening - Vibration, shock, and transportation stress testing - Power cycling, thermal cycling, and long-duration soak testing - Lead root-cause investigations for reliability failures surfaced during development, manufacturing, and field deployment, driving corrective actions across hardware, firmware, thermal, and mechanical domains - Develop comprehensive system reliability models including MTBF projections, FIT rate analysis, Weibull lifetime modeling, component derating methodologies, and reliability growth tracking - Ensure reliability is considered early, partnering with Platform Engineering architects and design leads so reliability requirements shape decisions before they become expensive to change - Work closely with ODMs, JDMs, contract manufacturers, and component suppliers to validate and enforce long-term platform reliability commitments - Build fleet reliability infrastructure: telemetry analysis pipelines, field feedback loops, and monitoring frameworks that give Etched visibility into deployed system health at scale - Drive reliability signoff criteria and lead product release readiness reviews across engineering and program teams - Build and lead a high-performing product reliability engineering organization — hiring, developing, and retaining technical talent as the company scales You may be a good fit if you have (Must-have qualifications) - BS, MS, or PhD in Electrical Engineering, Mechanical Engineering, Reliability Engineering, or a related technical field - 10+ years of reliability engineering experience in hardware-centric organizations, with meaningful time spent on complex systems rather than component-level work - Experience leading reliability programs for one or more of: - AI accelerator or GPU-class compute systems - Hyperscale or cloud server infrastructure - Networking platforms, storage systems, or rack-scale infrastructure - Deep understanding of system-level failure mechanisms — including thermal, power delivery, mechanical, and connector/interconnect failure modes — and how design decisions affect long-term field reliability - Hands-on experience with FMEA, Weibull analysis, HALT/HASS, qualification planning, failure analysis methodologies, and reliability statistics and modeling - A track record of driving cross-functional root-cause investigations in fast-moving hardware organizations where schedule pressure is real and accountability is high - Strong technical judgment — capable of making defensible tradeoffs between reliability targets, cost, schedule, and performance without losing sight of customer expectations - Excellent communication skills and the credibility to influence design decisions with engineering leads, program managers, and executive stakeholders Strong candidates may also have experience with (Nice-to-have qualifications) - Experience with liquid-cooled systems, high-density power delivery, or thermal management for high-power AI infrastructure - Direct experience supporting hyperscale or cloud datacenter deployments at scale, including customer-facing reliability commitments and SLA management - Demonstrated experience building a reliability organization from early-stage — establishing processes, tooling, and team norms in environments without established infrastructure - Familiarity with fleet telemetry systems, large-scale field reliability analytics, and data-driven approaches to proactive reliability management - Experience working closely with ODM or JDM partners in Taiwan or broader Asia, including NPI support and on-site qualification engagement - Background in high-speed digital systems, GPU compute platforms, or accelerator-based architectures — with an understanding of how these affect system-level reliability behavior ## Benefits - Medical, dental, and vision packages with generous premium coverage - $500 per month credit for waiving medical benefits - Housing subsidy of $2k per month for those living within walking distance of the office - Relocation support for those moving to San Jose (Santana Row) - Various wellness benefits covering fitness, mental health, and more - Daily lunch and dinner in our office - Unlimited compute budget subject to ROI justification ## How we’re different Etched believes in the Bitter Lesson http://www.incompleteideas.net/IncIdeas/BitterLesson.html. We are the first inference-focused frontier AI system. Our addressable market is the entirety of inference, unlike many of our competitors. We are a fully in-person team in San Jose (Santana Row), and greatly value engineering skills. We do not have boundaries between engineering and research, and we expect all of our technical staff to contribute to both and work across disciplines as needed. ## About Etched ## Company Overview - **One-liner**: Etched designs and manufactures specialized ASICs (application-specific integrated circuits) that are hardwired to run transformer-based AI models, offering an order-of-magnitude improvement in inference cost and energy efficiency over general-purpose GPUs. - **Entity Type**: Private (Series B / later stage; total raised $800M as of June 2026) - **Headquarters**: San Jose, California, USA (also has offices in Cupertino, CA; Sacramento, CA; and small presence in Taiwan, Canada, UK, Bulgaria) - **Founded**: 2022 - **Founders**: Gavin Uberti (CEO), Robert Wachen (President & COO), Chris Zhu (Co-Founder) ## Core Business - **Primary industry**: AI hardware / semiconductor design / computer hardware manufacturing - **Target customers**: Frontier AI labs, hyperscalers, and enterprises running large-scale transformer inference workloads (B2B, enterprise) - **Mission / purpose**: Build hardware for superintelligence – specifically, the world’s most powerful servers for transformer inference. ## Products & Services - **Sohu Chip (ASIC)**: A transformer-specific ASIC manufactured by TSMC. It burns the transformer architecture directly into silicon, eliminating the overhead of general-purpose GPUs for inference. - **Frontier Inference Clusters**: Full-stack systems that bundle the Sohu chip with custom-designed racks, networking, and software. Customers place orders for complete systems, not just chips. Etched claims they deliver more tokens per dollar and per watt for dense models, sparse MoEs, diffusion models, and more. - **Software Stack**: Includes compiler and runtime optimizations built by a team with deep experience (former Apache TVM developer, ex-Google TPU software lead). ## Market Standing - **Valuation**: $5 billion post-money valuation (as of December 2025 / disclosed June 2026) - **Key Metric**: $1 billion in contract orders (booked) for its systems; total funding of $800M (including an unannounced $500M round closed Dec 2025 led by Stripes) - **Notable Investors**: Stripes (lead), VentureTech Alliance, Jane Street, Hudson River Trading, Two Sigma, Ribbit Capital. Angel investors include Andrej Karpathy, Geoffrey Hinton, Fei-Fei Li, Arthur Mensch, Scott Wu, Stanley Druckenmiller, Peter Thiel. Earlier backers: Positive Sum, Primary Venture Partners. - **Growth Signals**: - Headcount: 258 employees (as of mid-2026), up 141% year-over-year - 100+ active job postings across engineering, operations, and business functions - $500M raise in late 2025; first chip manufactured by TSMC in 2025 and now being tested with customers - Talent sources include Google (21), Apple (17), NVIDIA (16), Tesla (11), Intel (10), AMD (5) ## Competitive Advantages - **Transformer-specific architecture**: Unlike GPUs (which are general-purpose), Etched’s chip is hardwired only for transformer operations, enabling drastically lower cost and latency for today’s dominant AI model architecture. - **Full-stack system play**: By selling entire inference clusters (chip + rack + software), Etched captures more value and delivers a turnkey solution competitive with NVIDIA’s DGX/HGX systems. - **Deeply experienced leadership**: - CEO Gavin Uberti (Harvard math researcher, AI compiler expert, Apache TVM contributor) - President Robert Wachen (co-founded Prod AI incubator; Thiel Fellow) - VP of Platform (ex-NVIDIA 22 years, built HGX/DGX systems) - VP of Software (ex-Google DeepMind, lead TPU v1–v5 software) - Chief Architect Saptadeep Pal (former Auradine co-founder, chiplet expert) - VP of Hardware Engineering Ajat Hukkoo (ex-Cypress CTO, shipped >$1B revenue products) ## Strategic Focus - **Scale manufacturing and customer deliveries**: With the first chip successfully fabricated, the company is now testing production systems and aims to fulfill $1B in pre-orders. - **Expand engineering talent**: Aggressively hiring across silicon design, validation, infrastructure, optics, and software to support volume ramp. - **Maintain technological lead**: Continue refining the ASIC for future transformer variants and MoE mixtures while building out a partner ecosystem. ## Why Work Here - **Mission-driven, high-impact engineering**: “Building the hardware for superintelligence” – the company is solving one of the hardest problems in compute (inference efficiency) with a focused team. - **Top-tier team and learning environment**: Colleagues from Google, NVIDIA, Apple, Tesla, Broadcom, and Intel; exposure to full-stack chip-to-system design. - **In-office culture**: All roles are in-office (Cupertino and San Jose, CA). The company emphasizes on-site collaboration for hardware development. Notable perks are not heavily advertised, but compensation is competitive for the AI hardware space (equity-heavy). - **Growth trajectory**: 141% headcount growth YoY, hundreds of open roles, and a $5B valuation signal strong backing and career acceleration potential. - **Intellectual challenge**: Work on chip simulation, RTL design, system validation, or software infrastructure for cutting-edge AI inference. ## Sources 1. [etched.com/careers](https://www.etched.com/careers) 2. [builtin.com/company/etched](https://builtin.com/company/etched) 3. [linkedin.com/company/etched-ai](https://www.linkedin.com/company/etched-ai) 4. [cbinsights.com/company/etched](https://www.cbinsights.com/company/etched) 5. [techcrunch.com/2026/06/30/nvidia-competitor-etched-hits-5b-valuation-1b-in-sales-for-ai-chip/](https://techcrunch.com/2026/06/30/nvidia-competitor-etched-hits-5b-valuation-1b-in-sales-for-ai-chip/) ## Other roles at Etched - [EE Hardware System Engineer (TW)](https://feeny.ai/job/ee-hardware-system-engineer-tw-etched-taipei-fgncym9kzy6w) — Taipei, Taiwan - [Silicon Validation Engineer, Software](https://feeny.ai/job/silicon-validation-engineer-software-etched-san-jose-qhp3k9xqy5vr) — San Jose, CA - [RMA & Repair Lead](https://feeny.ai/job/rma-repair-lead-etched-san-jose-q0tmchh0yyrj) — San Jose, CA - [Advanced Packaging SI/PI Engineer](https://feeny.ai/job/advanced-packaging-si-pi-engineer-etched-san-jose-47he7k0hp901) — San Jose, CA - [Substrate IC Package Layout Design Engineer](https://feeny.ai/job/substrate-ic-package-layout-design-engineer-etched-san-jose-g36wmxc8ww9r) — San Jose, CA - [ECAD Librarian](https://feeny.ai/job/ecad-librarian-etched-san-jose-0h5kvj2cxzk4) — San Jose, CA - [Inventory Warehouse Coordinator](https://feeny.ai/job/inventory-warehouse-coordinator-etched-san-jose-49b0bs49ce57) — San Jose, CA - [Logistics Analyst (Taiwan)](https://feeny.ai/job/logistics-analyst-taiwan-etched-taipei-5xzc094cqgzn) — Taipei, Taiwan - [Communications](https://feeny.ai/job/communications-etched-san-jose-3f4bdr97hysp) — San Jose, CA - [Production Finance](https://feeny.ai/job/production-finance-etched-san-jose-qfypbv9eh0ws) — San Jose, CA