🤖 AI Summary
AMD’s Strix Halo (Ryzen AI MAX) combines a 16-core Zen 5 CPU with a 20-WGP RDNA 3.5 GPU and pairs 256-bit LPDDR5X-8000 (theoretical 256 GB/s) with a 32 MB “Infinity Cache” (MALL). The tester used the chip’s accessible Infinity Fabric performance counters to compare traffic seen at Coherent Stations (CS) versus Unified Memory Controllers (UMCs), using CS-only traffic as a proxy for cache hits. Practical limitations — only eight counters, 1s sampling, CPU-side noise and rare cross-CCX probes — were acknowledged, but the approach uniquely exposes memory-side cache behavior on a mobile AMD part where discrete GPUs’ tools stop at L2.
Results show the 32 MB Infinity Cache meaningfully reduces DRAM pressure: in peak intervals it captured roughly 73% of Infinity Fabric traffic and kept most workloads comfortably below the LPDDR5X theoretical limit. Still, some benchmarks (e.g., 3DMark Time Spy Extreme, Ungine Valley) push close to bandwidth limits and would be DRAM-bound without the cache. Hit rates fall as resolution increases, and very high-demand scenarios would need much larger cache or far more DRAM bandwidth (estimates suggest >335 GB/s would be needed, or GDDR6-like 448 GB/s as in PS5). Overall, Infinity Cache proves an effective bandwidth-amplifier for mobile iGPU designs, but designers must balance cache capacity vs. raw DRAM bandwidth for extreme workloads.
Loading comments...
login to comment
loading comments...
no comments yet