With my new topology aware scripts I went back to my Ryzen AI 9 HX 370 system to see how things mapped.
Below is output of the topology script
================================================================================
CPU CORE & CACHE HIERARCHY MAP (X86_64)
================================================================================
Total Cores: 24
Core Details Table:
Core Core Model Stepping Max Freq L1I/L1D Cache L2 Cache
----------------------------------------------------------------------------------------
cpu0 AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 5158 MHz 32K / 48K 1024K
cpu1 AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 5158 MHz 32K / 48K 1024K
cpu2 AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 5158 MHz 32K / 48K 1024K
cpu3 AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 5158 MHz 32K / 48K 1024K
cpu4 AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 3289 MHz 32K / 48K 1024K
cpu5 AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 3289 MHz 32K / 48K 1024K
cpu6 AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 3289 MHz 32K / 48K 1024K
cpu7 AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 3289 MHz 32K / 48K 1024K
cpu8 AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 3289 MHz 32K / 48K 1024K
cpu9 AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 3289 MHz 32K / 48K 1024K
cpu10 AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 3289 MHz 32K / 48K 1024K
cpu11 AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 3289 MHz 32K / 48K 1024K
cpu12 AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 5158 MHz 32K / 48K 1024K
cpu13 AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 5158 MHz 32K / 48K 1024K
cpu14 AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 5158 MHz 32K / 48K 1024K
cpu15 AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 5158 MHz 32K / 48K 1024K
cpu16 AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 3289 MHz 32K / 48K 1024K
cpu17 AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 3289 MHz 32K / 48K 1024K
cpu18 AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 3289 MHz 32K / 48K 1024K
cpu19 AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 3289 MHz 32K / 48K 1024K
cpu20 AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 3289 MHz 32K / 48K 1024K
cpu21 AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 3289 MHz 32K / 48K 1024K
cpu22 AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 3289 MHz 32K / 48K 1024K
cpu23 AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 3289 MHz 32K / 48K 1024K
System-wide Cache Capacity Summary:
- L1 Data Cache: 576 KiB (Total across 12x 48K)
- L1 Instruction Cache: 384 KiB (Total across 12x 32K)
- L2 Unified Cache: 12.0 MiB (Total across 12x 1024K)
- L3 Unified Cache: 24.0 MiB (Total across 1x 16384K, 1x 8192K)
================================================================================
TOPOLOGY PLUMBING TREE
================================================================================
L3 Unified Cache [16384K]
├── L2 Unified Cache [1024K]
│ ├── L1 Data Cache [48K]
│ │ ├── Core 15: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 5.16 GHz
│ │ └── Core 3: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 5.16 GHz
│ └── L1 Instruction Cache [32K]
├── L2 Unified Cache [1024K]
│ ├── L1 Data Cache [48K]
│ │ ├── Core 14: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 5.16 GHz
│ │ └── Core 2: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 5.16 GHz
│ └── L1 Instruction Cache [32K]
├── L2 Unified Cache [1024K]
│ ├── L1 Data Cache [48K]
│ │ ├── Core 13: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 5.16 GHz
│ │ └── Core 1: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 5.16 GHz
│ └── L1 Instruction Cache [32K]
└── L2 Unified Cache [1024K]
├── L1 Data Cache [48K]
│ ├── Core 12: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 5.16 GHz
│ └── Core 0: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 5.16 GHz
└── L1 Instruction Cache [32K]
L3 Unified Cache [8192K]
├── L2 Unified Cache [1024K]
│ ├── L1 Data Cache [48K]
│ │ ├── Core 23: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 3.29 GHz
│ │ └── Core 11: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 3.29 GHz
│ └── L1 Instruction Cache [32K]
├── L2 Unified Cache [1024K]
│ ├── L1 Data Cache [48K]
│ │ ├── Core 22: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 3.29 GHz
│ │ └── Core 10: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 3.29 GHz
│ └── L1 Instruction Cache [32K]
├── L2 Unified Cache [1024K]
│ ├── L1 Data Cache [48K]
│ │ ├── Core 21: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 3.29 GHz
│ │ └── Core 9: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 3.29 GHz
│ └── L1 Instruction Cache [32K]
├── L2 Unified Cache [1024K]
│ ├── L1 Data Cache [48K]
│ │ ├── Core 20: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 3.29 GHz
│ │ └── Core 8: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 3.29 GHz
│ └── L1 Instruction Cache [32K]
├── L2 Unified Cache [1024K]
│ ├── L1 Data Cache [48K]
│ │ ├── Core 19: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 3.29 GHz
│ │ └── Core 7: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 3.29 GHz
│ └── L1 Instruction Cache [32K]
├── L2 Unified Cache [1024K]
│ ├── L1 Data Cache [48K]
│ │ ├── Core 18: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 3.29 GHz
│ │ └── Core 6: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 3.29 GHz
│ └── L1 Instruction Cache [32K]
├── L2 Unified Cache [1024K]
│ ├── L1 Data Cache [48K]
│ │ ├── Core 17: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 3.29 GHz
│ │ └── Core 5: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 3.29 GHz
│ └── L1 Instruction Cache [32K]
└── L2 Unified Cache [1024K]
├── L1 Data Cache [48K]
│ ├── Core 16: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 3.29 GHz
│ └── Core 4: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 3.29 GHz
└── L1 Instruction Cache [32K]
One thing that surprised me was previously likwid-topology had given a different mapping of L3 topology. Apparently, there are two core complexes. One has 4 Zen5 cores (8 threads) and 16MB of L3 and the other has 8 Zen5 compact cores (16 threads) and only 8MB of L3. That is consistent with specs from AMD so a spot where likwid-topology isn’t quite complete. I can probably update the topology report above to reflect hyperthreading but otherwise useful additional information.
Similar to previous ARM experiment, I tried a stream sweep to find best configuration
triad_MBps opt_level threads strategy domain cpus run
74213.600000 O2 2 spread_l3 rr 0,4 1
72749.400000 Ofast 2 spread_l3 rr 0,4 1
72687.900000 O3 2 spread_l3 rr 0,4 1
71942.400000 O2 3 spread_l3 rr 0,4,1 1
71505.000000 O2 4 spread_l3 rr 0,4,1,5 1
71500.900000 Ofast 3 spread_l3 rr 0,4,1 1
71464.400000 O3 3 spread_l3 rr 0,4,1 1
71460.800000 Ofast 4 spread_l3 rr 0,4,1,5 1
71403.000000 O3 4 spread_l3 rr 0,4,1,5 1
71041.700000 O2 2 local_l3 0 0,1 1
69966.800000 O2 2 local_l3 1 4,5 1
69621.300000 O2 3 local_l3 0 0,1,2 1
69486.700000 O3 2 local_l3 0 0,1 1
69432.500000 Ofast 2 local_l3 0 0,1 1
69296.300000 O3 3 local_l3 0 0,1,2 1
69168.200000 Ofast 3 local_l3 0 0,1,2 1
69074.700000 O2 3 local_l3 1 4,5,6 1
68989.500000 O2 4 local_l3 0 0,1,2,3 1
68979.600000 O2 4 local_l3 1 4,5,6,7 1
68943.700000 Ofast 3 local_l3 1 4,5,6 1
In this case, the fastest triad comes from using one thread from the Zen5 cores and one from the Zen5c cores, each with their own L3.
coremark_average scenario domain threads cpus run
641556.737750 all_logical all 24 0,1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23 1
465524.800224 thread0_only all 12 0,1,2,3,4,5,6,7,8,9,10,11 1
385663.999579 capability_group 1 16 4,5,6,7,8,9,10,11,16,17,18,19,20,21,22,23 1
381649.327648 l3_complex 1 16 4,5,6,7,8,9,10,11,16,17,18,19,20,21,22,23 1
307810.051852 l3_complex 0 8 0,1,2,3,12,13,14,15 1
305829.027379 capability_group 0 8 0,1,2,3,12,13,14,15 1
Looking at coremark across the different types of cores isn’t surprising.
- Running on all cores gives greatest throughput.
- Running only one thread per core (hyperthreading) gets about 72.5% of the performance.
- The “capability group” (Zen5 vs Zen5c) and “l3 complex” (those with the 16MB L3 and those with the 8 MB L3) are the same sets so results are the same. The performance of 8 ZenC cores (16 threads) still more than performance of 4 Zen cores (8 threads).

