Performance Measurements on four systems
With help of AI, I created a new “roofline” measurement that tries to write microbenchmarks to run system parameters. I have compared four CPU systems below. In some cases there might be issues with the microbenchmarks created that enable prefetching/optimization but still a useful overall comparison
| Metric | Ryzen 9950 + RX 7600 XT | Ryzen AI 9 HP 370 + 890M | Ryzen AI 395 + Radeon 8060S | ARM CP 8180 |
|---|---|---|---|---|
| Cores | 16 | 12 | 16 | 12 |
| Threads | 32 | 24 | 32 | 12 |
| L1d | 16x 48k = 768k | 12x 48k = 576k | 16x 48k = 768k | 4x 32k, 8x 64k = 640k |
| L1i | 16x 32k = 512k | 12x 32k = 384k | 16x 32k = 512k | 4x 32k, 8x 64k = 640k |
| L2 | 16x 1024k = 16M | 12x 1024k = 12M | 16 x 1024k = 16M | 8x 512k = 4M |
| L3 | 2x 32M = 64M | 16M + 8M = 24M | 2 x 32M = 64M | 1x 12M = 12M |
| Memory | 128 GB | 96 GB | 128 GB | 64 GB |
| Best CPU Core | Zen5 5756 MHz | Zen 5 5158 MHz | Zen5 5188 MHz | Cortex-A720 2600 MHz |
| FP64 FLOPS | 45.40 GFLOPS | 41.00 GFLOPS | 41.02 GFLOPS | 10.34 GFLOPS |
| FP32 FLOPS | 45.67 GFLOPS | 40.96 GFLOPS | 40.99 GFLOPS | 20.73 GFLOPS |
| FP16 FLOPS | 6.73 GFLOPS | 6.12 GFLOPS | 6.04 GFLOPS | 10.36 GFLOPS |
| L1 Triad Bandwidth | 223.33 GB/s | 207.39 GB/s | 202.87 GB/s | 52.30 GB/s |
| L2 Triad Bandwidth | 153.51 GB/s | 131.58 GB/s | 129.74 GB/s | 54.36 GB/s |
| L3 Triad Bandwidth | 66.12 GB/s | 73.21 GB/s | 69.30 GB/s | 46.84 GB/s |
| DRAM Bandwidth | 34.98 GB/s | 44.74 GB/s | 39.81 GB/s | 21.60 GB/s |
| L1 Latency | 0.698 ns | 0.781 ns | 0.832 ns | 1.544 ns |
| L2 Latency | 8.461 ns | 4.140 ns | 6.937 ns | 3.731 ns |
| L3 Latency | 11.971 ns | 11.543 ns | 13.326 ns | 6.143 ns |
| DRAM Latency | 81.349 ns | 100.16 ns | 81.067 ns | 156.623 ns |
| Core-Core Latency hyperthread | 79.86 ns | 26.23 ns | 106.63 ns | n/a |
| Core-Core Latency Perf Cores | 20.35 ns | 35.42 ns | 32.19 ns | 157.59 ns |
| Core-Core Latency Cross-CCD | 50.01 ns | 198.06 ns | 70.96 ns | 166.59 ns |
| Best GPU Core | RX 7600 XT – HIP | 890M – HIP | 8060S – HIP | n/a |
| FP16 FLOPS | 6465.80 GFLOPS | 4307.66 GFLOPS | 11290.44 GFLOPS | n/a |
| FP32 FLOPS | 5258.37 GFLOPS | 2792.52 GFLOPS | 7302.15 GFLOPS | n/a |
| FP64 FLOPS | 303.55 GFLOPS | 184.57 GFLOPS | 460.54 GFLOPS | n/a |
| Mem Triad | 303.55 GFLOPS | 79.69 GB/s | 231.22 GB/s | n/a |
| H2D Copy (Pinned) | 14.17 GB/s | 39.26 GB/s | 82.82 GB/s | n/a |
| D2H Copy (Pinned) | 14.29 GB/s | 38.10 GB/s | 64.62 GB/s | n/a |
| Kernel Launch Latency (sync) | 17.10 us | 6.96 us | 7.12 us | n/a |
| Kernel Launch Latency (async) | 2.60 us | 1.51 us | 1.73 us | n/a |
| Event Sync Latency | 15.07 us | 4.99 us | 5.03 us | n/a |

Comments
Performance Measurements on four systems — No Comments
HTML tags allowed in your comment: <a href="" title=""> <abbr title=""> <acronym title=""> <b> <blockquote cite=""> <cite> <code> <del datetime=""> <em> <i> <q cite=""> <s> <strike> <strong>