🤖 AI Summary
At Hot Chips 2026, d-Matrix unveiled its Raptor 3D-DRAM accelerator, designed to tackle the increasing demands of generative inference workloads in artificial intelligence. With model weights and the Key-Value (KV) cache growing substantially—potentially requiring up to 935 GB of storage for a single model—traditional memory solutions face both capacity and bandwidth limitations. d-Matrix addresses these challenges by stacking compute directly atop DRAM dies, achieving impressive bandwidth capabilities (up to 100 TB/s) and significantly lowering energy consumption, projected at around 0.37 pJ/bit, which is considerably more efficient than existing HBM4 architectures.
The Raptor architecture not only enhances performance during the decode phase—where most inference time is spent—but also simplifies the memory architecture by utilizing face-to-face stacking techniques with TSMC's N4 logic die for high yield and cost efficiency. By efficiently interleaving ECC bits and implementing stream blocking techniques, d-Matrix ensures utilization of bandwidth while maintaining thermal and power reliability. The Raptor's ability to serve up to 1,000 tokens per second per user for massive 3-trillion-parameter models signifies a potential breakthrough in optimizing AI inference systems, prompting a reevaluation of how capacity and bandwidth requirements are balanced in future accelerators.
Loading comments...
login to comment
loading comments...
no comments yet