🤖 AI Summary
Etched AI has announced the release of its Sohu chip, a transformer-only ASIC optimized for autoregressive language model inference. The Sohu architecture hard-codes transformer attention directly into silicon, allowing each 8-chip Sohu server to achieve impressive throughput rates of 500,000 tokens per second on the Llama 70B model, a stark contrast to the NVIDIA H100 GPU, which only manages about 700 tokens per second at batch size one. This architecture is significant as it signals a dedicated shift toward highly specialized hardware in an AI landscape increasingly reliant on transformer models. However, its lack of programmability raises questions about adaptability to future AI developments and model diversity.
The implications of choosing the Sohu architecture are noteworthy for teams currently evaluating inference hardware. While it offers remarkable per-chip efficiency and higher memory bandwidth (1.8 times that of the H100), Sohu's fixed-function design means it cannot execute models requiring operations beyond standard transformer attention, such as convolutional processing for multimodal applications or more complex architectures. With over $1 billion in customer contracts and first rack shipments expected in summer 2026, Sohu's emergence represents both a competitive threat to existing GPU capabilities and an essential consideration for organizations focused solely on transformer workloads, albeit with limitations on flexibility in accessing emerging AI models.
Loading comments...
login to comment
loading comments...
no comments yet