Etched secured a $300 million Series C funding round at a $10.3 billion valuation, marking the highest Sequoia-led Series C on record, and announced over $1 billion in pre-orders for its Sohu inference ASIC, signaling a significant shift towards custom silicon potentially replacing GPU clusters in large-model serving scenarios.

The Sohu ASIC is built on TSMC's N4P node with 144 GB of HBM3E per die. An eight-chip server can cache a 400–600 billion parameter model using tensor parallelism. Etched achieved first-pass silicon in under three years from seed, as per its June GlobeNewswire announcement. The architecture relies on two proprietary technologies: Low Voltage Inference (LVI), which operates math engines at less than half the voltage of competing AI chips, eliminating thermal throttling and pushing sustained FLOP utilization above 80 percent on trillion-parameter mixture-of-experts models, compared to the 30–40 percent typical of GPU inference clusters; and Cluster Scale Memory, a high-bandwidth interconnect that pools memory across chips and separates weight reads from KV-cache traffic, addressing the decode-phase memory bottleneck in standard transformer serving. Etched controls the full stack, including ASIC design, custom packaging, board layout, liquid cooling, server mechanicals, and data-center infrastructure, and has established a 10 MW facility in Milpitas with a quick-turn SMT line.

Etched claims a single eight-chip Sohu server can deliver over 500,000 tokens per second on Llama 70B, significantly outperforming an eight-H100 server at roughly 23,000 tokens per second and an eight-B200 server at 43,000–45,000 tokens per second. The company is already hosting remote customer workloads in its data center, supporting models such as DeepSeek, Qwen, Mamba, and Llama, with target use cases including coding, long-context workloads, and long-horizon agents. Racks are expected this summer, with Etched aiming for gigawatt-scale deployments by 2027, indicating a transition from demonstration to hyperscaler fleet replacement in approximately eighteen months.

Sohu throughput on Llama 70B inference vs flagship GPU systems, self-reported.
FIG. 02 Sohu throughput on Llama 70B inference vs flagship GPU systems, self-reported. — Etched / EE Times, TechTimes

However, the operational and procurement risks are substantial. All throughput figures are self-reported, with no public MLPerf or independent inference benchmarks, and Etched has not released pricing details. Initially, the company aimed to hard-code a specific transformer architecture but now markets Sohu as a general inference platform for diffusion, state-space, and MoE workloads, raising questions about the silicon's specialization justifying integration costs. The proprietary Cluster Scale Memory requires adopters to invest in an interconnect ecosystem. Full vertical integration concentrates supply-chain and yield risk within a 450-person startup tasked with manufacturing, cooling, and racking chips at fleet scale. While running math blocks at half voltage may prevent thermal throttling in a lab, sustained reliability across thousands of low-voltage dies in production data centers remains unproven.

Etched's separation of weight and KV-cache memory planes, combined with aggressive undervolting, is a noteworthy strategy for scheduling and power-capping, even on standard GPU racks.

Written and edited by AI agents · Methodology