↓
 

Performance analysis, tools and experiments

An eclectic collection

  • Overview
  • Blog
  • Workloads
    • cpu2017
      • 500.perlbench_r
      • 502.gcc_r
      • 503.bwaves_r
      • 505.mcf_r
      • 507.cactuBSSN_r
      • 508.namd_r
      • 510.parest_r
      • 511.povray_r
      • 519.lbm_r
      • 520.omnetpp_r
      • 521.wrf_r
      • 523.xalancbmk_r
      • 525.x264_r
      • 526.blender_r
      • 527.cam4_r
      • 531.deepsjeng_r
      • 538.imagick_r
      • 541.leela_r
      • 544.nab_r
      • 548.exchange2_r
      • 549.fotonik3d_r
      • 554.roms_r
      • 557.xz_r
    • geekbench
    • lmbench
    • passmark
    • pbbs
    • phoronix
      • ai-benchmark
      • aircrack-ng
      • amg
      • aobench
      • aom-av1
      • apache
      • apache-iotdb
      • appleseed
      • arrayfire
      • askap
      • asmfish
      • astcenc
      • avifenc
      • basis
      • blake2
      • blogbench
      • blender
      • blosc
      • bork
      • botan
      • brl-cad
      • build-apache
      • build-clash
      • build-eigen
      • build-erlang
      • build-ffmpeg
      • build-gcc
      • build-gdb
      • build-gem5
      • build-godot
      • build-imagemagick
      • build-linux-kernel
      • build-llvm
      • build-mesa
      • build-mplayer
      • build-nodejs
      • build-php
      • build-python
      • build-wasmer
      • build2
      • bullet
      • byte
      • cachebench
      • cassandra
      • clickhouse
      • clomp
      • cloverleaf
      • cockroach
      • compilebench
      • compress-7zip
      • compress-gzip
      • compress-lz4
      • compress-pbzip2
      • compress-rar
      • compress-xz
      • compress-zstd
      • core-latency
      • coremark
      • cp2k
      • cpp-perf-bench
      • cpuminer-opt
      • crafty
      • c-ray
      • cryptopp
      • cryptsetup
      • ctx-clock
      • cython-bench
      • dacapobench
      • daphne
      • darktable
      • dav1d
      • dbench
      • deepsparse
      • deepspeech
      • dolfyn
      • draco
      • dragonflydb
      • duckdb
      • easywave
      • ebizzy
      • embree
      • encode-flac
      • encode-mp3
      • encode-opus
      • encode-wavpack
      • espeak
      • etcpak
      • faiss
      • fast-cli
      • ffmpeg
      • ffte
      • fftw
      • fhourstones
      • financebench
      • furmark
      • gcrypt
      • gegl
      • gimp
      • git
      • glibc-bench
      • gmpbench
      • gnupg
      • gnuradio
      • go-benchmark
      • gpaw
      • graph500
      • graphics-magick
      • gromacs
      • hackbench
      • hadoop
      • heffte
      • helsing
      • himeno
      • hmmer
      • hpcg
      • incompact3d
      • indigobench
      • inkscape
      • ipc-benchmark
      • java-jmh
      • java-scimark2
      • john-the-ripper
      • jpegxl
      • jpegxl-decode
      • kvazaar
      • kripke
      • lammps
      • lczero
      • libraw
      • libreoffice
      • libxsmm
      • liquid-dsp
      • llama.cpp
      • llamafile
      • lulesh
      • lzbench
      • mbw
      • memcached
      • minibude
      • minife
      • mnn
      • mpcbench
      • m-queens
      • mrbayes
      • mutex
      • namd
      • mt-dgemm
      • ncnn
      • neat
      • nettle
      • nginx
      • ngspice
      • node-octane
      • node-web-tooling
      • npb
      • n-queens
      • numpy
      • nwchem
      • oidn
      • onednn
      • octave-benchmark
      • onnx
      • opencv
      • openfoam
      • openjpeg
      • openssl
      • openradioss
      • openscad
      • openvino
      • openvkl
      • ospray
      • ospray-studio
      • palabos
      • parboil
      • pennant
      • perl-benchmark
      • pgbench
      • phpbench
      • pjsip
      • polybench-c
      • polyhedron
      • povray
      • primesieve
      • pybench
      • pyhpc
      • pyperformance
      • pytorch
      • quadray
      • qe
      • qmcpack
      • quantlib
      • quicksilver
      • ramspeed
      • rav1e
      • rawtherapee
      • rbenchmark
      • redis
      • renaissance
      • rnnoise
      • rocksdb
      • rodinia
      • rsvg
      • schbench
      • scikit-learn
      • scimark2
      • scylladb
      • securemark
      • selenium
      • simdjson
      • smallpt
      • smhasher
      • spark
      • spark-tpcds
      • speedb
      • specfem3d
      • sqlite
      • srsran
      • stargate
      • stockfish
      • stream
      • stress-ng
      • svt-av1
      • svt-hevc
      • svt-vp9
      • sudokut
      • synthmark
      • sysbench
      • tensorflow
      • tensorflow-lite
      • tesseract
      • tjbench
      • tnn
      • toybrot
      • tscp
      • ttsiod-renderer
      • tungsten
      • uvg266
      • vkpeak
      • vpxenc
      • v-ray
      • vvenc
      • webp
      • webp2
      • whisper.cpp
      • whisperfile
      • wireguard
      • x264
      • x265
      • xmrig
      • xnnpack
      • y-cruncher
      • z3
    • stream
  • Tools
    • Compilers
    • likwid
    • perf
    • trace-cmd and kernelshark
    • wspy
  • Experiments
    • Histograms
    • clustering
    • Adding summary statistics for all benchmarks
  • Home
  • Blog
  • Workloads
    • cpu2017
      • 500.perlbench_r
      • 502.gcc_r
      • 503.bwaves_r
      • 505.mcf_r
      • 507.cactuBSSN_r
      • 508.namd_r
      • 510.parest_r
      • 511.povray_r
      • 519.lbm_r
      • 520.omnetpp_r
      • 521.wrf_r
      • 523.xalancbmk_r
      • 525.x264_r
      • 526.blender_r
      • 527.cam4_r
      • 531.deepsjeng_r
      • 538.imagick_r
      • 541.leela_r
      • 544.nab_r
      • 548.exchange2_r
      • 549.fotonik3d_r
      • 554.roms_r
      • 557.xz_r
    • geekbench
    • lmbench
    • passmark
    • pbbs
    • phoronix
      • ai-benchmark
      • aircrack-ng
      • amg
      • aobench
      • aom-av1
      • apache
      • apache-iotdb
      • appleseed
      • arrayfire
      • askap
      • asmfish
      • astcenc
      • avifenc
      • b
      • basis
      • blake2
      • blender
      • blogbench
      • blosc
      • bork
      • botan
      • brl-cad
      • build-apache
      • build-clash
      • build-eigen
      • build-erlang
      • build-ffmpeg
      • build-gcc
      • build-gdb
      • build-gem5
      • build-godot
      • build-imagemagick
      • build-linux-kernel
      • build-llvm
      • build-mesa
      • build-mplayer
      • build-nodejs
      • build-php
      • build-python
      • build-wasmer
      • build2
      • bullet
      • byte
      • c-ray
      • cachebench
      • cassandra
      • clickhouse
      • clomp
      • cloverleaf
      • cockroach
      • compilebench
      • compress-7zip
      • compress-gzip
      • compress-lz4
      • compress-pbzip2
      • compress-rar
      • compress-xz
      • compress-zstd
      • core-latency
      • coremark
      • cp2k
      • cpp-perf-bench
      • cpuminer-opt
      • crafty
      • cryptopp
      • cryptsetup
      • ctx-clock
      • cython-bench
      • dacapobench
      • daphne
      • darktable
      • dav1d
      • dbench
      • deepsparse
      • deepspeech
      • dolfyn
      • draco
      • dragonflydb
      • duckdb
      • easywave
      • ebizzy
      • embree
      • encode-flac
      • encode-mp3
      • encode-opus
      • encode-wavpack
      • espeak
      • etcpak
      • faiss
      • fast-cli
      • ffmpeg
      • ffte
      • fftw
      • fhourstones
      • financebench
      • furmark
      • gcrypt
      • gegl
      • gimp
      • git
      • glibc-bench
      • gmpbench
      • gnupg
      • gnuradio
      • go-benchmark
      • gpaw
      • graph500
      • graphics-magick
      • gromacs
      • hackbench
      • hadoop
      • heffte
      • helsing
      • himeno
      • hmmer
      • hpcg
      • incompact3d
      • indigobench
      • inkscape
      • ipc-benchmark
      • java-jmh
      • java-scimark2
      • john-the-ripper
      • jpegxl
      • jpegxl-decode
      • kripke
      • kvazaar
      • lammps
      • lczero
      • libraw
      • libreoffice
      • libxsmm
      • liquid-dsp
      • llama.cpp
      • llamafile
      • lulesh
      • lzbench
      • m-queens
      • mbw
      • memcached
      • minibude
      • minife
      • mnn
      • mpcbench
      • mrbayes
      • mt-dgemm
      • mutex
      • n-queens
      • namd
      • ncnn
      • neat
      • nettle
      • nginx
      • ngspice
      • node-octane
      • node-web-tooling
      • npb
      • numpy
      • nwchem
      • octave-benchmark
      • oidn
      • onednn
      • onnx
      • opencv
      • openfoam
      • openjpeg
      • openradioss
      • openscad
      • openssl
      • openvino
      • openvkl
      • ospray
      • ospray-studio
      • palabos
      • parboil
      • pennant
      • perl-benchmark
      • pgbench
      • phpbench
      • pjsip
      • polybench-c
      • polyhedron
      • povray
      • primesieve
      • pybench
      • pyhpc
      • pyperformance
      • pytorch
      • qe
      • qmcpack
      • quadray
      • quantlib
      • quicksilver
      • ramspeed
      • rav1e
      • rawtherapee
      • rays1bench
      • rbenchmark
      • redis
      • renaissance
      • rnnoise
      • rocksdb
      • rodinia
      • rsvg
      • schbench
      • scikit-learn
      • scimark2
      • scylladb
      • securemark
      • selenium
      • simdjson
      • smallpt
      • smhasher
      • spark
      • spark-tpcds
      • specfem3d
      • speedb
      • sqlite
      • srsran
      • stargate
      • stockfish
      • stream
      • stress-ng
      • sudokut
      • svt-av1
      • svt-hevc
      • svt-vp9
      • synthmark
      • sysbench
      • tensorflow
      • tensorflow-lite
      • tesseract
      • tjbench
      • tnn
      • toybrot
      • tscp
      • ttsiod-renderer
      • tungsten
      • uvg266
      • v-ray
      • vkpeak
      • vpxenc
      • vvenc
      • webp
      • webp2
      • whisper.cpp
      • whisperfile
      • wireguard
      • x264
      • x265
      • xmrig
      • xnnpack
      • y-cruncher
      • z3
    • stream
  • Tools
    • Compilers
    • likwid
    • perf
    • trace-cmd and kernelshark
    • wspy
  • Experiments
Home→Published 2026

Yearly Archives: 2026

Performance Measurements on four systems

Performance analysis, tools and experiments Posted on July 12, 2026 by mevJuly 12, 2026

With help of AI, I created a new “roofline” measurement that tries to write microbenchmarks to run system parameters. I have compared four CPU systems below. In some cases there might be issues with the microbenchmarks created that enable prefetching/optimization but still a useful overall comparison

MetricRyzen 9950 + RX 7600 XTRyzen AI 9 HP 370 + 890MRyzen AI 395 + Radeon 8060SARM CP 8180
Cores16121612
Threads32243212
L1d16x 48k = 768k12x 48k = 576k16x 48k = 768k4x 32k, 8x 64k = 640k
L1i16x 32k = 512k12x 32k = 384k16x 32k = 512k4x 32k, 8x 64k = 640k
L216x 1024k = 16M12x 1024k = 12M16 x 1024k = 16M8x 512k = 4M
L32x 32M = 64M16M + 8M = 24M2 x 32M = 64M1x 12M = 12M
Memory128 GB96 GB128 GB64 GB
Best CPU CoreZen5 5756 MHzZen 5 5158 MHzZen5 5188 MHzCortex-A720 2600 MHz
FP64 FLOPS45.40 GFLOPS41.00 GFLOPS41.02 GFLOPS10.34 GFLOPS
FP32 FLOPS45.67 GFLOPS40.96 GFLOPS40.99 GFLOPS20.73 GFLOPS
FP16 FLOPS6.73 GFLOPS6.12 GFLOPS6.04 GFLOPS10.36 GFLOPS
L1 Triad Bandwidth223.33 GB/s207.39 GB/s202.87 GB/s52.30 GB/s
L2 Triad Bandwidth153.51 GB/s131.58 GB/s129.74 GB/s54.36 GB/s
L3 Triad Bandwidth66.12 GB/s73.21 GB/s69.30 GB/s46.84 GB/s
DRAM Bandwidth34.98 GB/s44.74 GB/s39.81 GB/s21.60 GB/s
L1 Latency0.698 ns0.781 ns0.832 ns1.544 ns
L2 Latency8.461 ns4.140 ns6.937 ns3.731 ns
L3 Latency11.971 ns11.543 ns13.326 ns6.143 ns
DRAM Latency81.349 ns100.16 ns81.067 ns156.623 ns
Core-Core Latency hyperthread79.86 ns26.23 ns106.63 nsn/a
Core-Core Latency Perf Cores20.35 ns35.42 ns32.19 ns157.59 ns
Core-Core Latency Cross-CCD50.01 ns198.06 ns70.96 ns166.59 ns
Best GPU CoreRX 7600 XT – HIP890M – HIP8060S – HIPn/a
FP16 FLOPS6465.80 GFLOPS4307.66 GFLOPS11290.44 GFLOPSn/a
FP32 FLOPS5258.37 GFLOPS2792.52 GFLOPS7302.15 GFLOPSn/a
FP64 FLOPS303.55 GFLOPS184.57 GFLOPS460.54 GFLOPSn/a
Mem Triad303.55 GFLOPS79.69 GB/s231.22 GB/sn/a
H2D Copy (Pinned)14.17 GB/s39.26 GB/s82.82 GB/sn/a
D2H Copy (Pinned)14.29 GB/s38.10 GB/s64.62 GB/sn/a
Kernel Launch Latency (sync)17.10 us6.96 us7.12 usn/a
Kernel Launch Latency (async)2.60 us1.51 us1.73 usn/a
Event Sync Latency15.07 us4.99 us5.03 usn/a
Posted in hardware | Tagged Aarch64, benchmarks, Ryzen AI 9 HX 370, Zen5 | Leave a reply

New scripts with Ryzen AI 9 HX 370

Performance analysis, tools and experiments Posted on July 12, 2026 by mevJuly 12, 2026

With my new topology aware scripts I went back to my Ryzen AI 9 HX 370 system to see how things mapped.

Below is output of the topology script

================================================================================
CPU CORE & CACHE HIERARCHY MAP (X86_64)
================================================================================
Total Cores: 24

Core Details Table:
Core   Core Model                   Stepping   Max Freq     L1I/L1D Cache    L2 Cache  
----------------------------------------------------------------------------------------
cpu0   AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 5158 MHz     32K / 48K        1024K     
cpu1   AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 5158 MHz     32K / 48K        1024K     
cpu2   AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 5158 MHz     32K / 48K        1024K     
cpu3   AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 5158 MHz     32K / 48K        1024K     
cpu4   AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 3289 MHz     32K / 48K        1024K     
cpu5   AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 3289 MHz     32K / 48K        1024K     
cpu6   AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 3289 MHz     32K / 48K        1024K     
cpu7   AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 3289 MHz     32K / 48K        1024K     
cpu8   AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 3289 MHz     32K / 48K        1024K     
cpu9   AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 3289 MHz     32K / 48K        1024K     
cpu10  AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 3289 MHz     32K / 48K        1024K     
cpu11  AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 3289 MHz     32K / 48K        1024K     
cpu12  AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 5158 MHz     32K / 48K        1024K     
cpu13  AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 5158 MHz     32K / 48K        1024K     
cpu14  AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 5158 MHz     32K / 48K        1024K     
cpu15  AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 5158 MHz     32K / 48K        1024K     
cpu16  AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 3289 MHz     32K / 48K        1024K     
cpu17  AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 3289 MHz     32K / 48K        1024K     
cpu18  AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 3289 MHz     32K / 48K        1024K     
cpu19  AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 3289 MHz     32K / 48K        1024K     
cpu20  AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 3289 MHz     32K / 48K        1024K     
cpu21  AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 3289 MHz     32K / 48K        1024K     
cpu22  AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 3289 MHz     32K / 48K        1024K     
cpu23  AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 3289 MHz     32K / 48K        1024K     

System-wide Cache Capacity Summary:
 - L1 Data        Cache: 576 KiB      (Total across 12x 48K)
 - L1 Instruction Cache: 384 KiB      (Total across 12x 32K)
 - L2 Unified     Cache: 12.0 MiB     (Total across 12x 1024K)
 - L3 Unified     Cache: 24.0 MiB     (Total across 1x 16384K, 1x 8192K)

================================================================================
TOPOLOGY PLUMBING TREE
================================================================================
L3 Unified Cache [16384K]
├── L2 Unified Cache [1024K]
│   ├── L1 Data Cache [48K]
│   │   ├── Core 15: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 5.16 GHz
│   │   └── Core 3: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 5.16 GHz
│   └── L1 Instruction Cache [32K]
├── L2 Unified Cache [1024K]
│   ├── L1 Data Cache [48K]
│   │   ├── Core 14: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 5.16 GHz
│   │   └── Core 2: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 5.16 GHz
│   └── L1 Instruction Cache [32K]
├── L2 Unified Cache [1024K]
│   ├── L1 Data Cache [48K]
│   │   ├── Core 13: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 5.16 GHz
│   │   └── Core 1: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 5.16 GHz
│   └── L1 Instruction Cache [32K]
└── L2 Unified Cache [1024K]
    ├── L1 Data Cache [48K]
    │   ├── Core 12: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 5.16 GHz
    │   └── Core 0: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 5.16 GHz
    └── L1 Instruction Cache [32K]
L3 Unified Cache [8192K]
├── L2 Unified Cache [1024K]
│   ├── L1 Data Cache [48K]
│   │   ├── Core 23: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 3.29 GHz
│   │   └── Core 11: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 3.29 GHz
│   └── L1 Instruction Cache [32K]
├── L2 Unified Cache [1024K]
│   ├── L1 Data Cache [48K]
│   │   ├── Core 22: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 3.29 GHz
│   │   └── Core 10: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 3.29 GHz
│   └── L1 Instruction Cache [32K]
├── L2 Unified Cache [1024K]
│   ├── L1 Data Cache [48K]
│   │   ├── Core 21: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 3.29 GHz
│   │   └── Core 9: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 3.29 GHz
│   └── L1 Instruction Cache [32K]
├── L2 Unified Cache [1024K]
│   ├── L1 Data Cache [48K]
│   │   ├── Core 20: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 3.29 GHz
│   │   └── Core 8: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 3.29 GHz
│   └── L1 Instruction Cache [32K]
├── L2 Unified Cache [1024K]
│   ├── L1 Data Cache [48K]
│   │   ├── Core 19: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 3.29 GHz
│   │   └── Core 7: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 3.29 GHz
│   └── L1 Instruction Cache [32K]
├── L2 Unified Cache [1024K]
│   ├── L1 Data Cache [48K]
│   │   ├── Core 18: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 3.29 GHz
│   │   └── Core 6: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 3.29 GHz
│   └── L1 Instruction Cache [32K]
├── L2 Unified Cache [1024K]
│   ├── L1 Data Cache [48K]
│   │   ├── Core 17: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 3.29 GHz
│   │   └── Core 5: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 3.29 GHz
│   └── L1 Instruction Cache [32K]
└── L2 Unified Cache [1024K]
    ├── L1 Data Cache [48K]
    │   ├── Core 16: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 3.29 GHz
    │   └── Core 4: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 3.29 GHz
    └── L1 Instruction Cache [32K]

One thing that surprised me was previously likwid-topology had given a different mapping of L3 topology. Apparently, there are two core complexes. One has 4 Zen5 cores (8 threads) and 16MB of L3 and the other has 8 Zen5 compact cores (16 threads) and only 8MB of L3. That is consistent with specs from AMD so a spot where likwid-topology isn’t quite complete. I can probably update the topology report above to reflect hyperthreading but otherwise useful additional information.

Similar to previous ARM experiment, I tried a stream sweep to find best configuration

triad_MBps	opt_level	threads	strategy	domain	cpus	run
74213.600000	O2	2	spread_l3	rr	0,4	1
72749.400000	Ofast	2	spread_l3	rr	0,4	1
72687.900000	O3	2	spread_l3	rr	0,4	1
71942.400000	O2	3	spread_l3	rr	0,4,1	1
71505.000000	O2	4	spread_l3	rr	0,4,1,5	1
71500.900000	Ofast	3	spread_l3	rr	0,4,1	1
71464.400000	O3	3	spread_l3	rr	0,4,1	1
71460.800000	Ofast	4	spread_l3	rr	0,4,1,5	1
71403.000000	O3	4	spread_l3	rr	0,4,1,5	1
71041.700000	O2	2	local_l3	0	0,1	1
69966.800000	O2	2	local_l3	1	4,5	1
69621.300000	O2	3	local_l3	0	0,1,2	1
69486.700000	O3	2	local_l3	0	0,1	1
69432.500000	Ofast	2	local_l3	0	0,1	1
69296.300000	O3	3	local_l3	0	0,1,2	1
69168.200000	Ofast	3	local_l3	0	0,1,2	1
69074.700000	O2	3	local_l3	1	4,5,6	1
68989.500000	O2	4	local_l3	0	0,1,2,3	1
68979.600000	O2	4	local_l3	1	4,5,6,7	1
68943.700000	Ofast	3	local_l3	1	4,5,6	1

In this case, the fastest triad comes from using one thread from the Zen5 cores and one from the Zen5c cores, each with their own L3.

coremark_average	scenario	domain	threads	cpus	run
641556.737750	all_logical	all	24	0,1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23	1
465524.800224	thread0_only	all	12	0,1,2,3,4,5,6,7,8,9,10,11	1
385663.999579	capability_group	1	16	4,5,6,7,8,9,10,11,16,17,18,19,20,21,22,23	1
381649.327648	l3_complex	1	16	4,5,6,7,8,9,10,11,16,17,18,19,20,21,22,23	1
307810.051852	l3_complex	0	8	0,1,2,3,12,13,14,15	1
305829.027379	capability_group	0	8	0,1,2,3,12,13,14,15	1

Looking at coremark across the different types of cores isn’t surprising.

  • Running on all cores gives greatest throughput.
  • Running only one thread per core (hyperthreading) gets about 72.5% of the performance.
  • The “capability group” (Zen5 vs Zen5c) and “l3 complex” (those with the 16MB L3 and those with the 8 MB L3) are the same sets so results are the same. The performance of 8 ZenC cores (16 threads) still more than performance of 4 Zen cores (8 threads).

Posted in experiment, hardware | Tagged coremark, Ryzen AI 9 HX 370, stream | Leave a reply

ARM box, inventory script

Performance analysis, tools and experiments Posted on July 12, 2026 by mevJuly 12, 2026

I have a Minisforum MS-R1 system that will let me experiment further with Aarch 64 cores on Linux. To help me with this discovery, I extended several scripts to inventory and examine the system so will also show those scripts outputs on this system and others to calibrate.

The first script generates the topology of both cores and cache on the system from /sys entries:

================================================================================
CPU CORE & CACHE HIERARCHY MAP (AARCH64)
================================================================================
Total Cores: 12

Core Details Table:
Core   Core Model                   Stepping   Max Freq     L1I/L1D Cache    L2 Cache  
----------------------------------------------------------------------------------------
cpu0   ARM Cortex-A720              r0p1       2600 MHz     64K / 64K        512K      
cpu1   ARM Cortex-A720              r0p1       2600 MHz     64K / 64K        512K      
cpu2   ARM Cortex-A520              r0p1       1800 MHz     32K / 32K        N/A       
cpu3   ARM Cortex-A520              r0p1       1800 MHz     32K / 32K        N/A       
cpu4   ARM Cortex-A520              r0p1       1800 MHz     32K / 32K        N/A       
cpu5   ARM Cortex-A520              r0p1       1800 MHz     32K / 32K        N/A       
cpu6   ARM Cortex-A720              r0p1       2300 MHz     64K / 64K        512K      
cpu7   ARM Cortex-A720              r0p1       2300 MHz     64K / 64K        512K      
cpu8   ARM Cortex-A720              r0p1       2200 MHz     64K / 64K        512K      
cpu9   ARM Cortex-A720              r0p1       2200 MHz     64K / 64K        512K      
cpu10  ARM Cortex-A720              r0p1       2500 MHz     64K / 64K        512K      
cpu11  ARM Cortex-A720              r0p1       2500 MHz     64K / 64K        512K      

System-wide Cache Capacity Summary:
 - L1 Data        Cache: 640 KiB      (Total across 4x 32K, 8x 64K)
 - L1 Instruction Cache: 640 KiB      (Total across 4x 32K, 8x 64K)
 - L2 Unified     Cache: 4.0 MiB      (Total across 8x 512K, 4x unreported)
 - L3 Unified     Cache: 12.0 MiB     (Total across 1x 12288K)

================================================================================
TOPOLOGY PLUMBING TREE
================================================================================
L3 Unified Cache [12288K]
├── Core 11: ARM Cortex-A720 (r0p1) @ 2.50 GHz
│   └─ Private Caches: L1 Data (64K), L1 Instruction (64K), L2 Unified (512K)
├── Core 10: ARM Cortex-A720 (r0p1) @ 2.50 GHz
│   └─ Private Caches: L1 Data (64K), L1 Instruction (64K), L2 Unified (512K)
├── Core 9: ARM Cortex-A720 (r0p1) @ 2.20 GHz
│   └─ Private Caches: L1 Data (64K), L1 Instruction (64K), L2 Unified (512K)
├── Core 8: ARM Cortex-A720 (r0p1) @ 2.20 GHz
│   └─ Private Caches: L1 Data (64K), L1 Instruction (64K), L2 Unified (512K)
├── Core 7: ARM Cortex-A720 (r0p1) @ 2.30 GHz
│   └─ Private Caches: L1 Data (64K), L1 Instruction (64K), L2 Unified (512K)
├── Core 6: ARM Cortex-A720 (r0p1) @ 2.30 GHz
│   └─ Private Caches: L1 Data (64K), L1 Instruction (64K), L2 Unified (512K)
├── Core 5: ARM Cortex-A520 (r0p1) @ 1.80 GHz
│   └─ Private Caches: L1 Data (32K), L1 Instruction (32K)
├── Core 4: ARM Cortex-A520 (r0p1) @ 1.80 GHz
│   └─ Private Caches: L1 Data (32K), L1 Instruction (32K)
├── Core 3: ARM Cortex-A520 (r0p1) @ 1.80 GHz
│   └─ Private Caches: L1 Data (32K), L1 Instruction (32K)
├── Core 2: ARM Cortex-A520 (r0p1) @ 1.80 GHz
│   └─ Private Caches: L1 Data (32K), L1 Instruction (32K)
├── Core 1: ARM Cortex-A720 (r0p1) @ 2.60 GHz
│   └─ Private Caches: L1 Data (64K), L1 Instruction (64K), L2 Unified (512K)
└── Core 0: ARM Cortex-A720 (r0p1) @ 2.60 GHz
    └─ Private Caches: L1 Data (64K), L1 Instruction (64K), L2 Unified (512K)

This shows this SOC has a mix of both Cortex A720 cores and Cortex A520 cores and even the Cortex A720 cores have different maximum frequencies reflecting possible different between “high” and “mid” cores.

The first thing we do on this new system is run a sweep of stream pinned to different cores. I updated a script to consider the core capability, the optimization level and the number of threads to create the following measurements:

triad_MBps	opt_level	threads	strategy	domain	cpus	run
39170.300000	O2	4	local_l3	0	0,1,10,11	1
38441.200000	O2	2	local_l3	0	0,1	1
37068.000000	O2	2	local_l3_cc	0	0,10	1
36531.400000	O2	3	local_l3	0	0,1,10	1
35612.500000	O2	2	local_l3_cc	0	0,6	1
34266.600000	O2	2	local_l3_cc	0	0,11	1
33733.900000	O2	2	local_l3_cc	0	0,7	1
26004.700000	O2	1	local_l3	0	0	1

The highest stream triad comes from running four threads and pinning to the four highest frequency Cortex 720 cores. Just using the two fastest cores comes close.

The next thing we do is try coremark for each individual core

cpu	run	status	coremark_average	log_file
0	1	OK	25512.373511	/home/mev/source/perf/results/coremark_each_core/coremark.cpu0.run1.txt
1	1	OK	25511.108133	/home/mev/source/perf/results/coremark_each_core/coremark.cpu1.run1.txt
2	1	OK	8182.294456	/home/mev/source/perf/results/coremark_each_core/coremark.cpu2.run1.txt
3	1	OK	8168.957901	/home/mev/source/perf/results/coremark_each_core/coremark.cpu3.run1.txt
4	1	OK	8172.910173	/home/mev/source/perf/results/coremark_each_core/coremark.cpu4.run1.txt
5	1	OK	8176.058957	/home/mev/source/perf/results/coremark_each_core/coremark.cpu5.run1.txt
6	1	OK	22462.563490	/home/mev/source/perf/results/coremark_each_core/coremark.cpu6.run1.txt
7	1	OK	22442.589501	/home/mev/source/perf/results/coremark_each_core/coremark.cpu7.run1.txt
8	1	OK	21470.071023	/home/mev/source/perf/results/coremark_each_core/coremark.cpu8.run1.txt
9	1	OK	21462.437238	/home/mev/source/perf/results/coremark_each_core/coremark.cpu9.run1.txt
10	1	OK	24509.139361	/home/mev/source/perf/results/coremark_each_core/coremark.cpu10.run1.txt
11	1	OK	24504.245930	/home/mev/source/perf/results/coremark_each_core/coremark.cpu11.run1.txt

This shows us results that are consistent with the cores and their maximum frequencies. It surprises me how much quicker a Cortex 720 is than a Cortex 520, so some strategy that pins to the 8 Cortex 720 cores might make sense in some situations.

coremark_average	scenario	domain	threads	cpus	run
172518.777163	thread0_only	all	12	0,1,2,3,4,5,6,7,8,9,10,11	1
168110.086124	all_logical	all	12	0,1,2,3,4,5,6,7,8,9,10,11	1
51064.558269	capability_group	0	2	0,1	1
49082.209232	capability_group	1	2	10,11	1
44976.394529	capability_group	2	2	6,7	1
43020.862299	capability_group	3	2	8,9	1
33342.065373	capability_group	4	4	2,3,4,5	1

I also tried some aggregation of running on all cores, on all thread 0 (hyperthread) and for each pair of cores that have the same capabilities. The first two numbers are close because they are the same run. The pairs of Cortex cores by themselves have coremark scores according to frequencies. The four Cortex 520 cores by themselves are slower than any pair of Cortex 720 cores.

As a follow on post, I will also document what these same scripts have shown with a Ryzen AI 9 370 also show.

Posted in experiment, hardware, Tools | Tagged Aarch64, coremark, stream | Leave a reply

Meta

  • Log in
  • Entries feed
  • Comments feed
  • WordPress.org

Archives

  • July 2026
  • November 2024
  • October 2024
  • September 2024
  • July 2024
  • June 2024
  • March 2024
  • February 2024
  • January 2024
  • December 2023
  • February 2023

Tags

7840HS Aarch64 bad data benchmarks cachyos cluster compiler coremark cpu2017 data fabric getrusage gnuplot i5-13500H icache ipc kernel l3 metrics namd opcache perf performance counters perf_event_open phoronix Ryzen AI 9 HX 370 Ryzen AI 365 scaling stream threshold topdown tree virtualization website wsl Zen5

Recent Posts

  • Performance Measurements on four systems
  • New scripts with Ryzen AI 9 HX 370
  • ARM box, inventory script
  • Virtualization comparisons
  • Updating to a new kernel and graphics driver
©2026 - Performance analysis, tools and experiments - Weaver Xtreme Theme
↑