↓
 

Performance analysis, tools and experiments

An eclectic collection

  • Overview
  • Blog
  • Workloads
    • cpu2017
      • 500.perlbench_r
      • 502.gcc_r
      • 503.bwaves_r
      • 505.mcf_r
      • 507.cactuBSSN_r
      • 508.namd_r
      • 510.parest_r
      • 511.povray_r
      • 519.lbm_r
      • 520.omnetpp_r
      • 521.wrf_r
      • 523.xalancbmk_r
      • 525.x264_r
      • 526.blender_r
      • 527.cam4_r
      • 531.deepsjeng_r
      • 538.imagick_r
      • 541.leela_r
      • 544.nab_r
      • 548.exchange2_r
      • 549.fotonik3d_r
      • 554.roms_r
      • 557.xz_r
    • geekbench
    • lmbench
    • passmark
    • pbbs
    • phoronix
      • ai-benchmark
      • aircrack-ng
      • amg
      • aobench
      • aom-av1
      • apache
      • apache-iotdb
      • appleseed
      • arrayfire
      • askap
      • asmfish
      • astcenc
      • avifenc
      • basis
      • blake2
      • blogbench
      • blender
      • blosc
      • bork
      • botan
      • brl-cad
      • build-apache
      • build-clash
      • build-eigen
      • build-erlang
      • build-ffmpeg
      • build-gcc
      • build-gdb
      • build-gem5
      • build-godot
      • build-imagemagick
      • build-linux-kernel
      • build-llvm
      • build-mesa
      • build-mplayer
      • build-nodejs
      • build-php
      • build-python
      • build-wasmer
      • build2
      • bullet
      • byte
      • cachebench
      • cassandra
      • clickhouse
      • clomp
      • cloverleaf
      • cockroach
      • compilebench
      • compress-7zip
      • compress-gzip
      • compress-lz4
      • compress-pbzip2
      • compress-rar
      • compress-xz
      • compress-zstd
      • core-latency
      • coremark
      • cp2k
      • cpp-perf-bench
      • cpuminer-opt
      • crafty
      • c-ray
      • cryptopp
      • cryptsetup
      • ctx-clock
      • cython-bench
      • dacapobench
      • daphne
      • darktable
      • dav1d
      • dbench
      • deepsparse
      • deepspeech
      • dolfyn
      • draco
      • dragonflydb
      • duckdb
      • easywave
      • ebizzy
      • embree
      • encode-flac
      • encode-mp3
      • encode-opus
      • encode-wavpack
      • espeak
      • etcpak
      • faiss
      • fast-cli
      • ffmpeg
      • ffte
      • fftw
      • fhourstones
      • financebench
      • furmark
      • gcrypt
      • gegl
      • gimp
      • git
      • glibc-bench
      • gmpbench
      • gnupg
      • gnuradio
      • go-benchmark
      • gpaw
      • graph500
      • graphics-magick
      • gromacs
      • hackbench
      • hadoop
      • heffte
      • helsing
      • himeno
      • hmmer
      • hpcg
      • incompact3d
      • indigobench
      • inkscape
      • ipc-benchmark
      • java-jmh
      • java-scimark2
      • john-the-ripper
      • jpegxl
      • jpegxl-decode
      • kvazaar
      • kripke
      • lammps
      • lczero
      • libraw
      • libreoffice
      • libxsmm
      • liquid-dsp
      • llama.cpp
      • llamafile
      • lulesh
      • lzbench
      • mbw
      • memcached
      • minibude
      • minife
      • mnn
      • mpcbench
      • m-queens
      • mrbayes
      • mutex
      • namd
      • mt-dgemm
      • ncnn
      • neat
      • nettle
      • nginx
      • ngspice
      • node-octane
      • node-web-tooling
      • npb
      • n-queens
      • numpy
      • nwchem
      • oidn
      • onednn
      • octave-benchmark
      • onnx
      • opencv
      • openfoam
      • openjpeg
      • openssl
      • openradioss
      • openscad
      • openvino
      • openvkl
      • ospray
      • ospray-studio
      • palabos
      • parboil
      • pennant
      • perl-benchmark
      • pgbench
      • phpbench
      • pjsip
      • polybench-c
      • polyhedron
      • povray
      • primesieve
      • pybench
      • pyhpc
      • pyperformance
      • pytorch
      • quadray
      • qe
      • qmcpack
      • quantlib
      • quicksilver
      • ramspeed
      • rav1e
      • rawtherapee
      • rbenchmark
      • redis
      • renaissance
      • rnnoise
      • rocksdb
      • rodinia
      • rsvg
      • schbench
      • scikit-learn
      • scimark2
      • scylladb
      • securemark
      • selenium
      • simdjson
      • smallpt
      • smhasher
      • spark
      • spark-tpcds
      • speedb
      • specfem3d
      • sqlite
      • srsran
      • stargate
      • stockfish
      • stream
      • stress-ng
      • svt-av1
      • svt-hevc
      • svt-vp9
      • sudokut
      • synthmark
      • sysbench
      • tensorflow
      • tensorflow-lite
      • tesseract
      • tjbench
      • tnn
      • toybrot
      • tscp
      • ttsiod-renderer
      • tungsten
      • uvg266
      • vkpeak
      • vpxenc
      • v-ray
      • vvenc
      • webp
      • webp2
      • whisper.cpp
      • whisperfile
      • wireguard
      • x264
      • x265
      • xmrig
      • xnnpack
      • y-cruncher
      • z3
    • stream
  • Tools
    • Compilers
    • likwid
    • perf
    • trace-cmd and kernelshark
    • wspy
  • Experiments
    • Histograms
    • clustering
    • Adding summary statistics for all benchmarks
  • Home
  • Blog
  • Workloads
    • cpu2017
      • 500.perlbench_r
      • 502.gcc_r
      • 503.bwaves_r
      • 505.mcf_r
      • 507.cactuBSSN_r
      • 508.namd_r
      • 510.parest_r
      • 511.povray_r
      • 519.lbm_r
      • 520.omnetpp_r
      • 521.wrf_r
      • 523.xalancbmk_r
      • 525.x264_r
      • 526.blender_r
      • 527.cam4_r
      • 531.deepsjeng_r
      • 538.imagick_r
      • 541.leela_r
      • 544.nab_r
      • 548.exchange2_r
      • 549.fotonik3d_r
      • 554.roms_r
      • 557.xz_r
    • geekbench
    • lmbench
    • passmark
    • pbbs
    • phoronix
      • ai-benchmark
      • aircrack-ng
      • amg
      • aobench
      • aom-av1
      • apache
      • apache-iotdb
      • appleseed
      • arrayfire
      • askap
      • asmfish
      • astcenc
      • avifenc
      • b
      • basis
      • blake2
      • blender
      • blogbench
      • blosc
      • bork
      • botan
      • brl-cad
      • build-apache
      • build-clash
      • build-eigen
      • build-erlang
      • build-ffmpeg
      • build-gcc
      • build-gdb
      • build-gem5
      • build-godot
      • build-imagemagick
      • build-linux-kernel
      • build-llvm
      • build-mesa
      • build-mplayer
      • build-nodejs
      • build-php
      • build-python
      • build-wasmer
      • build2
      • bullet
      • byte
      • c-ray
      • cachebench
      • cassandra
      • clickhouse
      • clomp
      • cloverleaf
      • cockroach
      • compilebench
      • compress-7zip
      • compress-gzip
      • compress-lz4
      • compress-pbzip2
      • compress-rar
      • compress-xz
      • compress-zstd
      • core-latency
      • coremark
      • cp2k
      • cpp-perf-bench
      • cpuminer-opt
      • crafty
      • cryptopp
      • cryptsetup
      • ctx-clock
      • cython-bench
      • dacapobench
      • daphne
      • darktable
      • dav1d
      • dbench
      • deepsparse
      • deepspeech
      • dolfyn
      • draco
      • dragonflydb
      • duckdb
      • easywave
      • ebizzy
      • embree
      • encode-flac
      • encode-mp3
      • encode-opus
      • encode-wavpack
      • espeak
      • etcpak
      • faiss
      • fast-cli
      • ffmpeg
      • ffte
      • fftw
      • fhourstones
      • financebench
      • furmark
      • gcrypt
      • gegl
      • gimp
      • git
      • glibc-bench
      • gmpbench
      • gnupg
      • gnuradio
      • go-benchmark
      • gpaw
      • graph500
      • graphics-magick
      • gromacs
      • hackbench
      • hadoop
      • heffte
      • helsing
      • himeno
      • hmmer
      • hpcg
      • incompact3d
      • indigobench
      • inkscape
      • ipc-benchmark
      • java-jmh
      • java-scimark2
      • john-the-ripper
      • jpegxl
      • jpegxl-decode
      • kripke
      • kvazaar
      • lammps
      • lczero
      • libraw
      • libreoffice
      • libxsmm
      • liquid-dsp
      • llama.cpp
      • llamafile
      • lulesh
      • lzbench
      • m-queens
      • mbw
      • memcached
      • minibude
      • minife
      • mnn
      • mpcbench
      • mrbayes
      • mt-dgemm
      • mutex
      • n-queens
      • namd
      • ncnn
      • neat
      • nettle
      • nginx
      • ngspice
      • node-octane
      • node-web-tooling
      • npb
      • numpy
      • nwchem
      • octave-benchmark
      • oidn
      • onednn
      • onnx
      • opencv
      • openfoam
      • openjpeg
      • openradioss
      • openscad
      • openssl
      • openvino
      • openvkl
      • ospray
      • ospray-studio
      • palabos
      • parboil
      • pennant
      • perl-benchmark
      • pgbench
      • phpbench
      • pjsip
      • polybench-c
      • polyhedron
      • povray
      • primesieve
      • pybench
      • pyhpc
      • pyperformance
      • pytorch
      • qe
      • qmcpack
      • quadray
      • quantlib
      • quicksilver
      • ramspeed
      • rav1e
      • rawtherapee
      • rays1bench
      • rbenchmark
      • redis
      • renaissance
      • rnnoise
      • rocksdb
      • rodinia
      • rsvg
      • schbench
      • scikit-learn
      • scimark2
      • scylladb
      • securemark
      • selenium
      • simdjson
      • smallpt
      • smhasher
      • spark
      • spark-tpcds
      • specfem3d
      • speedb
      • sqlite
      • srsran
      • stargate
      • stockfish
      • stream
      • stress-ng
      • sudokut
      • svt-av1
      • svt-hevc
      • svt-vp9
      • synthmark
      • sysbench
      • tensorflow
      • tensorflow-lite
      • tesseract
      • tjbench
      • tnn
      • toybrot
      • tscp
      • ttsiod-renderer
      • tungsten
      • uvg266
      • v-ray
      • vkpeak
      • vpxenc
      • vvenc
      • webp
      • webp2
      • whisper.cpp
      • whisperfile
      • wireguard
      • x264
      • x265
      • xmrig
      • xnnpack
      • y-cruncher
      • z3
    • stream
  • Tools
    • Compilers
    • likwid
    • perf
    • trace-cmd and kernelshark
    • wspy
  • Experiments
Home→Tags stream

Tag Archives: stream

New scripts with Ryzen AI 9 HX 370

Performance analysis, tools and experiments Posted on July 12, 2026 by mevJuly 12, 2026

With my new topology aware scripts I went back to my Ryzen AI 9 HX 370 system to see how things mapped.

Below is output of the topology script

================================================================================
CPU CORE & CACHE HIERARCHY MAP (X86_64)
================================================================================
Total Cores: 24

Core Details Table:
Core   Core Model                   Stepping   Max Freq     L1I/L1D Cache    L2 Cache  
----------------------------------------------------------------------------------------
cpu0   AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 5158 MHz     32K / 48K        1024K     
cpu1   AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 5158 MHz     32K / 48K        1024K     
cpu2   AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 5158 MHz     32K / 48K        1024K     
cpu3   AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 5158 MHz     32K / 48K        1024K     
cpu4   AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 3289 MHz     32K / 48K        1024K     
cpu5   AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 3289 MHz     32K / 48K        1024K     
cpu6   AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 3289 MHz     32K / 48K        1024K     
cpu7   AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 3289 MHz     32K / 48K        1024K     
cpu8   AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 3289 MHz     32K / 48K        1024K     
cpu9   AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 3289 MHz     32K / 48K        1024K     
cpu10  AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 3289 MHz     32K / 48K        1024K     
cpu11  AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 3289 MHz     32K / 48K        1024K     
cpu12  AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 5158 MHz     32K / 48K        1024K     
cpu13  AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 5158 MHz     32K / 48K        1024K     
cpu14  AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 5158 MHz     32K / 48K        1024K     
cpu15  AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 5158 MHz     32K / 48K        1024K     
cpu16  AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 3289 MHz     32K / 48K        1024K     
cpu17  AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 3289 MHz     32K / 48K        1024K     
cpu18  AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 3289 MHz     32K / 48K        1024K     
cpu19  AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 3289 MHz     32K / 48K        1024K     
cpu20  AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 3289 MHz     32K / 48K        1024K     
cpu21  AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 3289 MHz     32K / 48K        1024K     
cpu22  AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 3289 MHz     32K / 48K        1024K     
cpu23  AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 3289 MHz     32K / 48K        1024K     

System-wide Cache Capacity Summary:
 - L1 Data        Cache: 576 KiB      (Total across 12x 48K)
 - L1 Instruction Cache: 384 KiB      (Total across 12x 32K)
 - L2 Unified     Cache: 12.0 MiB     (Total across 12x 1024K)
 - L3 Unified     Cache: 24.0 MiB     (Total across 1x 16384K, 1x 8192K)

================================================================================
TOPOLOGY PLUMBING TREE
================================================================================
L3 Unified Cache [16384K]
├── L2 Unified Cache [1024K]
│   ├── L1 Data Cache [48K]
│   │   ├── Core 15: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 5.16 GHz
│   │   └── Core 3: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 5.16 GHz
│   └── L1 Instruction Cache [32K]
├── L2 Unified Cache [1024K]
│   ├── L1 Data Cache [48K]
│   │   ├── Core 14: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 5.16 GHz
│   │   └── Core 2: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 5.16 GHz
│   └── L1 Instruction Cache [32K]
├── L2 Unified Cache [1024K]
│   ├── L1 Data Cache [48K]
│   │   ├── Core 13: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 5.16 GHz
│   │   └── Core 1: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 5.16 GHz
│   └── L1 Instruction Cache [32K]
└── L2 Unified Cache [1024K]
    ├── L1 Data Cache [48K]
    │   ├── Core 12: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 5.16 GHz
    │   └── Core 0: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 5.16 GHz
    └── L1 Instruction Cache [32K]
L3 Unified Cache [8192K]
├── L2 Unified Cache [1024K]
│   ├── L1 Data Cache [48K]
│   │   ├── Core 23: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 3.29 GHz
│   │   └── Core 11: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 3.29 GHz
│   └── L1 Instruction Cache [32K]
├── L2 Unified Cache [1024K]
│   ├── L1 Data Cache [48K]
│   │   ├── Core 22: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 3.29 GHz
│   │   └── Core 10: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 3.29 GHz
│   └── L1 Instruction Cache [32K]
├── L2 Unified Cache [1024K]
│   ├── L1 Data Cache [48K]
│   │   ├── Core 21: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 3.29 GHz
│   │   └── Core 9: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 3.29 GHz
│   └── L1 Instruction Cache [32K]
├── L2 Unified Cache [1024K]
│   ├── L1 Data Cache [48K]
│   │   ├── Core 20: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 3.29 GHz
│   │   └── Core 8: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 3.29 GHz
│   └── L1 Instruction Cache [32K]
├── L2 Unified Cache [1024K]
│   ├── L1 Data Cache [48K]
│   │   ├── Core 19: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 3.29 GHz
│   │   └── Core 7: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 3.29 GHz
│   └── L1 Instruction Cache [32K]
├── L2 Unified Cache [1024K]
│   ├── L1 Data Cache [48K]
│   │   ├── Core 18: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 3.29 GHz
│   │   └── Core 6: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 3.29 GHz
│   └── L1 Instruction Cache [32K]
├── L2 Unified Cache [1024K]
│   ├── L1 Data Cache [48K]
│   │   ├── Core 17: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 3.29 GHz
│   │   └── Core 5: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 3.29 GHz
│   └── L1 Instruction Cache [32K]
└── L2 Unified Cache [1024K]
    ├── L1 Data Cache [48K]
    │   ├── Core 16: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 3.29 GHz
    │   └── Core 4: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 3.29 GHz
    └── L1 Instruction Cache [32K]

One thing that surprised me was previously likwid-topology had given a different mapping of L3 topology. Apparently, there are two core complexes. One has 4 Zen5 cores (8 threads) and 16MB of L3 and the other has 8 Zen5 compact cores (16 threads) and only 8MB of L3. That is consistent with specs from AMD so a spot where likwid-topology isn’t quite complete. I can probably update the topology report above to reflect hyperthreading but otherwise useful additional information.

Similar to previous ARM experiment, I tried a stream sweep to find best configuration

triad_MBps	opt_level	threads	strategy	domain	cpus	run
74213.600000	O2	2	spread_l3	rr	0,4	1
72749.400000	Ofast	2	spread_l3	rr	0,4	1
72687.900000	O3	2	spread_l3	rr	0,4	1
71942.400000	O2	3	spread_l3	rr	0,4,1	1
71505.000000	O2	4	spread_l3	rr	0,4,1,5	1
71500.900000	Ofast	3	spread_l3	rr	0,4,1	1
71464.400000	O3	3	spread_l3	rr	0,4,1	1
71460.800000	Ofast	4	spread_l3	rr	0,4,1,5	1
71403.000000	O3	4	spread_l3	rr	0,4,1,5	1
71041.700000	O2	2	local_l3	0	0,1	1
69966.800000	O2	2	local_l3	1	4,5	1
69621.300000	O2	3	local_l3	0	0,1,2	1
69486.700000	O3	2	local_l3	0	0,1	1
69432.500000	Ofast	2	local_l3	0	0,1	1
69296.300000	O3	3	local_l3	0	0,1,2	1
69168.200000	Ofast	3	local_l3	0	0,1,2	1
69074.700000	O2	3	local_l3	1	4,5,6	1
68989.500000	O2	4	local_l3	0	0,1,2,3	1
68979.600000	O2	4	local_l3	1	4,5,6,7	1
68943.700000	Ofast	3	local_l3	1	4,5,6	1

In this case, the fastest triad comes from using one thread from the Zen5 cores and one from the Zen5c cores, each with their own L3.

coremark_average	scenario	domain	threads	cpus	run
641556.737750	all_logical	all	24	0,1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23	1
465524.800224	thread0_only	all	12	0,1,2,3,4,5,6,7,8,9,10,11	1
385663.999579	capability_group	1	16	4,5,6,7,8,9,10,11,16,17,18,19,20,21,22,23	1
381649.327648	l3_complex	1	16	4,5,6,7,8,9,10,11,16,17,18,19,20,21,22,23	1
307810.051852	l3_complex	0	8	0,1,2,3,12,13,14,15	1
305829.027379	capability_group	0	8	0,1,2,3,12,13,14,15	1

Looking at coremark across the different types of cores isn’t surprising.

  • Running on all cores gives greatest throughput.
  • Running only one thread per core (hyperthreading) gets about 72.5% of the performance.
  • The “capability group” (Zen5 vs Zen5c) and “l3 complex” (those with the 16MB L3 and those with the 8 MB L3) are the same sets so results are the same. The performance of 8 ZenC cores (16 threads) still more than performance of 4 Zen cores (8 threads).

Posted in experiment, hardware | Tagged coremark, Ryzen AI 9 HX 370, stream | Leave a reply

ARM box, inventory script

Performance analysis, tools and experiments Posted on July 12, 2026 by mevJuly 12, 2026

I have a Minisforum MS-R1 system that will let me experiment further with Aarch 64 cores on Linux. To help me with this discovery, I extended several scripts to inventory and examine the system so will also show those scripts outputs on this system and others to calibrate.

The first script generates the topology of both cores and cache on the system from /sys entries:

================================================================================
CPU CORE & CACHE HIERARCHY MAP (AARCH64)
================================================================================
Total Cores: 12

Core Details Table:
Core   Core Model                   Stepping   Max Freq     L1I/L1D Cache    L2 Cache  
----------------------------------------------------------------------------------------
cpu0   ARM Cortex-A720              r0p1       2600 MHz     64K / 64K        512K      
cpu1   ARM Cortex-A720              r0p1       2600 MHz     64K / 64K        512K      
cpu2   ARM Cortex-A520              r0p1       1800 MHz     32K / 32K        N/A       
cpu3   ARM Cortex-A520              r0p1       1800 MHz     32K / 32K        N/A       
cpu4   ARM Cortex-A520              r0p1       1800 MHz     32K / 32K        N/A       
cpu5   ARM Cortex-A520              r0p1       1800 MHz     32K / 32K        N/A       
cpu6   ARM Cortex-A720              r0p1       2300 MHz     64K / 64K        512K      
cpu7   ARM Cortex-A720              r0p1       2300 MHz     64K / 64K        512K      
cpu8   ARM Cortex-A720              r0p1       2200 MHz     64K / 64K        512K      
cpu9   ARM Cortex-A720              r0p1       2200 MHz     64K / 64K        512K      
cpu10  ARM Cortex-A720              r0p1       2500 MHz     64K / 64K        512K      
cpu11  ARM Cortex-A720              r0p1       2500 MHz     64K / 64K        512K      

System-wide Cache Capacity Summary:
 - L1 Data        Cache: 640 KiB      (Total across 4x 32K, 8x 64K)
 - L1 Instruction Cache: 640 KiB      (Total across 4x 32K, 8x 64K)
 - L2 Unified     Cache: 4.0 MiB      (Total across 8x 512K, 4x unreported)
 - L3 Unified     Cache: 12.0 MiB     (Total across 1x 12288K)

================================================================================
TOPOLOGY PLUMBING TREE
================================================================================
L3 Unified Cache [12288K]
├── Core 11: ARM Cortex-A720 (r0p1) @ 2.50 GHz
│   └─ Private Caches: L1 Data (64K), L1 Instruction (64K), L2 Unified (512K)
├── Core 10: ARM Cortex-A720 (r0p1) @ 2.50 GHz
│   └─ Private Caches: L1 Data (64K), L1 Instruction (64K), L2 Unified (512K)
├── Core 9: ARM Cortex-A720 (r0p1) @ 2.20 GHz
│   └─ Private Caches: L1 Data (64K), L1 Instruction (64K), L2 Unified (512K)
├── Core 8: ARM Cortex-A720 (r0p1) @ 2.20 GHz
│   └─ Private Caches: L1 Data (64K), L1 Instruction (64K), L2 Unified (512K)
├── Core 7: ARM Cortex-A720 (r0p1) @ 2.30 GHz
│   └─ Private Caches: L1 Data (64K), L1 Instruction (64K), L2 Unified (512K)
├── Core 6: ARM Cortex-A720 (r0p1) @ 2.30 GHz
│   └─ Private Caches: L1 Data (64K), L1 Instruction (64K), L2 Unified (512K)
├── Core 5: ARM Cortex-A520 (r0p1) @ 1.80 GHz
│   └─ Private Caches: L1 Data (32K), L1 Instruction (32K)
├── Core 4: ARM Cortex-A520 (r0p1) @ 1.80 GHz
│   └─ Private Caches: L1 Data (32K), L1 Instruction (32K)
├── Core 3: ARM Cortex-A520 (r0p1) @ 1.80 GHz
│   └─ Private Caches: L1 Data (32K), L1 Instruction (32K)
├── Core 2: ARM Cortex-A520 (r0p1) @ 1.80 GHz
│   └─ Private Caches: L1 Data (32K), L1 Instruction (32K)
├── Core 1: ARM Cortex-A720 (r0p1) @ 2.60 GHz
│   └─ Private Caches: L1 Data (64K), L1 Instruction (64K), L2 Unified (512K)
└── Core 0: ARM Cortex-A720 (r0p1) @ 2.60 GHz
    └─ Private Caches: L1 Data (64K), L1 Instruction (64K), L2 Unified (512K)

This shows this SOC has a mix of both Cortex A720 cores and Cortex A520 cores and even the Cortex A720 cores have different maximum frequencies reflecting possible different between “high” and “mid” cores.

The first thing we do on this new system is run a sweep of stream pinned to different cores. I updated a script to consider the core capability, the optimization level and the number of threads to create the following measurements:

triad_MBps	opt_level	threads	strategy	domain	cpus	run
39170.300000	O2	4	local_l3	0	0,1,10,11	1
38441.200000	O2	2	local_l3	0	0,1	1
37068.000000	O2	2	local_l3_cc	0	0,10	1
36531.400000	O2	3	local_l3	0	0,1,10	1
35612.500000	O2	2	local_l3_cc	0	0,6	1
34266.600000	O2	2	local_l3_cc	0	0,11	1
33733.900000	O2	2	local_l3_cc	0	0,7	1
26004.700000	O2	1	local_l3	0	0	1

The highest stream triad comes from running four threads and pinning to the four highest frequency Cortex 720 cores. Just using the two fastest cores comes close.

The next thing we do is try coremark for each individual core

cpu	run	status	coremark_average	log_file
0	1	OK	25512.373511	/home/mev/source/perf/results/coremark_each_core/coremark.cpu0.run1.txt
1	1	OK	25511.108133	/home/mev/source/perf/results/coremark_each_core/coremark.cpu1.run1.txt
2	1	OK	8182.294456	/home/mev/source/perf/results/coremark_each_core/coremark.cpu2.run1.txt
3	1	OK	8168.957901	/home/mev/source/perf/results/coremark_each_core/coremark.cpu3.run1.txt
4	1	OK	8172.910173	/home/mev/source/perf/results/coremark_each_core/coremark.cpu4.run1.txt
5	1	OK	8176.058957	/home/mev/source/perf/results/coremark_each_core/coremark.cpu5.run1.txt
6	1	OK	22462.563490	/home/mev/source/perf/results/coremark_each_core/coremark.cpu6.run1.txt
7	1	OK	22442.589501	/home/mev/source/perf/results/coremark_each_core/coremark.cpu7.run1.txt
8	1	OK	21470.071023	/home/mev/source/perf/results/coremark_each_core/coremark.cpu8.run1.txt
9	1	OK	21462.437238	/home/mev/source/perf/results/coremark_each_core/coremark.cpu9.run1.txt
10	1	OK	24509.139361	/home/mev/source/perf/results/coremark_each_core/coremark.cpu10.run1.txt
11	1	OK	24504.245930	/home/mev/source/perf/results/coremark_each_core/coremark.cpu11.run1.txt

This shows us results that are consistent with the cores and their maximum frequencies. It surprises me how much quicker a Cortex 720 is than a Cortex 520, so some strategy that pins to the 8 Cortex 720 cores might make sense in some situations.

coremark_average	scenario	domain	threads	cpus	run
172518.777163	thread0_only	all	12	0,1,2,3,4,5,6,7,8,9,10,11	1
168110.086124	all_logical	all	12	0,1,2,3,4,5,6,7,8,9,10,11	1
51064.558269	capability_group	0	2	0,1	1
49082.209232	capability_group	1	2	10,11	1
44976.394529	capability_group	2	2	6,7	1
43020.862299	capability_group	3	2	8,9	1
33342.065373	capability_group	4	4	2,3,4,5	1

I also tried some aggregation of running on all cores, on all thread 0 (hyperthread) and for each pair of cores that have the same capabilities. The first two numbers are close because they are the same run. The pairs of Cortex cores by themselves have coremark scores according to frequencies. The four Cortex 520 cores by themselves are slower than any pair of Cortex 720 cores.

As a follow on post, I will also document what these same scripts have shown with a Ryzen AI 9 370 also show.

Posted in experiment, hardware, Tools | Tagged Aarch64, coremark, stream | Leave a reply

New Ryzen AI 9 HX 370 machine

Performance analysis, tools and experiments Posted on October 8, 2024 by mevOctober 10, 2024

I have a new AMD performance machine for experiments. The processor is a Ryzen AI 9 HX 370 in a Beelink SER9 mini-PC.

Following are some of the major parameters.in comparison with my Ryzen 7840HS comparison machine.

ItemRyzen 7840HSRyzen AI 9 HX 370Notes
ArchitectureZen4Zen 5
Cores812
(4x Zen 5 and 8x Zen 5c)
Threads1624
Base Clock3.8 GHz2.0 GHz, 2.0 GHz
Boost Clock5.1 GHz5.1 GHz, 3.3 GHz
TDP35-45W15-54WSet by vendor
Memory32 GB (2 x 16 GiB)

DDR5 – 5600

2 Memory Channels
32 GB (4x 8 GiB)

DDR5 – 7500

2 Memory Channels
Check BIOS for actual speed
StreamCopy: 71400 MB/s
Scale: 70300 MB/s
Add: 73600 MB/s
Triad: 73000 MB/s
Copy: 86725 MB/s
Scale: 86626 MS/s
Add: 88192 MB/s
Triad: 87655 MB/s
Measured
CacheL1 – 32kB, 8 way, 4 clocks

L2 – 1 MB, 8-way, 14 clocks

L3 – 16MB, 24 way, 47 clocks
L1 – 32kB

L2 – 1 MB

L3 – 24 MB
Agner Fog architecture document and likwid-topology
lmbenchL1 – 0.8 ns
L2 – 3 ns
L3 – 8 ns
L1 – 0.8 ns
L2 – 3ns
L3 – 8 ns
Measured in Nanoseconds
GraphicsRadeon 780M

12 cores

2700 MHz
Radeon 890M

16 cores

2900 MHz
Phoronix streamAverage: 40604 MB/sAverage 44500 MB/s
Phoronix coremarkAverage 464076 Iterations/secondAverage 563477 Iterations/second+21%

Following are the results from likwid-topology. This is a hybrid core with four Zen5 cores and eight Zen5c cores. I believe the first four cores are Zen5 and the remaining eight are Zen5c.

--------------------------------------------------------------------------------
CPU name:	AMD Ryzen AI 9 HX 370 w/ Radeon 890M           
CPU type:	nil
CPU stepping:	0
********************************************************************************
Hardware Thread Topology
********************************************************************************
Sockets:		1
Cores per socket:	12
Threads per core:	2
--------------------------------------------------------------------------------
HWThread        Thread        Core        Die        Socket        Available
0               0             0           0          0             *                
1               0             1           0          0             *                
2               0             2           0          0             *                
3               0             3           0          0             *                
4               0             4           0          0             *                
5               0             5           0          0             *                
6               0             6           0          0             *                
7               0             7           0          0             *                
8               0             8           0          0             *                
9               0             9           0          0             *                
10              0             10          0          0             *                
11              0             11          0          0             *                
12              1             0           0          0             *                
13              1             1           0          0             *                
14              1             2           0          0             *                
15              1             3           0          0             *                
16              1             4           0          0             *                
17              1             5           0          0             *                
18              1             6           0          0             *                
19              1             7           0          0             *                
20              1             8           0          0             *                
21              1             9           0          0             *                
22              1             10          0          0             *                
23              1             11          0          0             *                
--------------------------------------------------------------------------------
Socket 0:		( 0 12 1 13 2 14 3 15 4 16 5 17 6 18 7 19 8 20 9 21 10 22 11 23 )
--------------------------------------------------------------------------------
********************************************************************************
Cache Topology
********************************************************************************
Level:			1
Size:			48 kB
Cache groups:		( 0 12 ) ( 1 13 ) ( 2 14 ) ( 3 15 ) ( 4 16 ) ( 5 17 ) ( 6 18 ) ( 7 19 ) ( 8 20 ) ( 9 21 ) ( 10 22 ) ( 11 23 )
--------------------------------------------------------------------------------
Level:			2
Size:			1 MB
Cache groups:		( 0 12 ) ( 1 13 ) ( 2 14 ) ( 3 15 ) ( 4 16 ) ( 5 17 ) ( 6 18 ) ( 7 19 ) ( 8 20 ) ( 9 21 ) ( 10 22 ) ( 11 23 )
--------------------------------------------------------------------------------
Level:			3
Size:			16 MB
Cache groups:		( 0 12 1 13 2 14 3 15 ) ( 4 16 5 17 6 18 7 19 ) ( 8 20 9 21 10 22 11 23 )
--------------------------------------------------------------------------------
********************************************************************************
NUMA Topology
********************************************************************************
NUMA domains:		1
--------------------------------------------------------------------------------
Domain:			0
Processors:		( 0 12 1 13 2 14 3 15 4 16 5 17 6 18 7 19 8 20 9 21 10 22 11 23 )
Distances:		10
Free memory:		22667.5 MB
Total memory:		27574.2 MB
--------------------------------------------------------------------------------

The L3 cache amount may be incorrect as specifications suggest 24 MB of cache. Using lmbench suggests the L3 cache attached to first four cores is 16MB and the next groups have 8MB likely together even though topology above makes them separate.

This hybrid SOC shows up in the following coremark scaling comparison as shown in the graph below. There are several different regions

  • From 1 to 4 cores we compare Zen4 cores against Zen5 cores. The coremark value for 4 cores is ~12% ahead.
  • From 5 to 8 cores, we now have Zen5 + Zen5C cores against Zen4 cores. The coremark value for 8 cores is ~7% behind
  • From 9 to 12 cores, we use all the cores on HX 370 and start using SMT for the 7840. The coremark value for 12 cores is 6% ahead
  • From 13 to 16 cores we go to using SMT for all the Zen5 cores and not-SMT for Zen5C cores. The 7840 moves to fully SMT. The coremark value for 16 cores is 11% ahead
  • From 17 to 24 cores, we go to adding SMT for Zen5C cores. The overall coremark using all cores (24 vs 16) is 21% ahead.

This suggests for coremark and other workloads there will be different regions where combinations of SMT and Zen5 vs Zen5C cores will create interesting comparisons between the systems.

The tabular version of coremark including performance counters is shown below.

CoresCoremark HX 370Coremark 7840Scaling HX 370Scaling 7840Retiring HX 370Frontend HX 370Backend HX 370Speculation HX 370SMT-contention HX 370Retiring 7840Frontend 7840Backend 7840Speculation 7840SMT-contention 7840
14824543881100%100%44.2%25.2%62.0%2.0%0.0%43.9%12.4%43.0%0.7%0.0%
29610685758100%98%44.0%25.5%61.8%2.0%0.0%43.9%12.4%43.1%0.7%0.0%
3144147128841100%98%44.0%25.5%61.8%2.0%0.0%43.6%13.0%42.7%0.7%0.0%
4192537171061100%97%44.1%25.4%61.9%2.0%0.0%43.9%12.3%43.1%0.7%0.0%
521422321036889%96%44.0%25.5%61.8%2.0%0.0%43.9%12.3%43.1%0.7%0.0%
622753225170579%96%44.0%25.4%61.9%2.0%0.0%43.2%12.9%43.2%0.7%0.0%
726081128136977%92%44.0%25.7%61.7%2.0%0.0%43.3%12.2%43.7%0.7%0.0%
829700231909877%91%44.1%25.3%61.9%2.0%0.0%42.7%12.8%43.8%0.7%0.0%
932541733460275%85%44.1%25.3%62.0%2.0%0.0%40.2%15.9%36.3%0.6%7.1%
1034763634724672%79%44.0%25.3%61.9%2.0%0.0%38.4%17.8%30.2%0.5%13.1%
1138058735940272%74%44.0%25.5%61.8%2.0%0.0%36.9%19.6%25.3%0.5%17.8%
1241357536328871%69%44.0%25.4%61.9%2.0%0.0%35.5%21.1%21.6%0.4%21.3%
1342612336214468%63%42.1%28.2%52.9%1.8%8.3%34.4%22.4%18.5%0.4%24.3%
1444637937776766%61%40.5%30.6%45.6%1.6%15.1%33.1%24.4%15.2%0.4%26.9%
1545213439714562%60%39.5%32.2%40.6%1.4%19.7%32.2%25.3%12.0%0.3%30.2%
1646443141846260%60%38.3%33.7%35.8%1.3%24.2%31.1%26.0%9.5%0.3%33.1%
1747641658%37.9%34.4%33.5%1.2%26.3%
1848900156%37.2%35.0%31.2%1.2%28.7%
1948465553%36.6%35.4%29.2%1.1%30.9%
2049582651%36.5%36.5%26.3%1.0%33.1%
2150145749%35.7%37.3%23.9%1.0%35.5%
2251094648%35.1%37.7%22.0%0.9%37.6%
2354489549%34.7%38.5%19.5%0.8%39.8%
2456347749%34.0%38.2%19.4%0.8%40.9%

I also measured stream and it looks ~15% faster than my 7840 system.

-------------------------------------------------------------
STREAM version $Revision: 5.10 $
-------------------------------------------------------------
This system uses 8 bytes per array element.
-------------------------------------------------------------
Array size = 100000000 (elements), Offset = 0 (elements)
Memory per array = 762.9 MiB (= 0.7 GiB).
Total memory required = 2288.8 MiB (= 2.2 GiB).
Each kernel will be executed 100 times.
 The *best* time for each kernel (excluding the first iteration)
 will be used to compute the reported bandwidth.
-------------------------------------------------------------
Number of Threads requested = 2
Number of Threads counted = 2
-------------------------------------------------------------
Your clock granularity/precision appears to be 1 microseconds.
Each test below will take on the order of 31409 microseconds.
   (= 31409 clock ticks)
Increase the size of the arrays if this shows that
you are not getting at least 20 clock ticks per test.
-------------------------------------------------------------
WARNING -- The above is only a rough guideline.
For best results, please be sure you know the
precision of your system timer.
-------------------------------------------------------------
Function    Best Rate MB/s  Avg time     Min time     Max time
Copy:           86725.2     0.018665     0.018449     0.021070
Scale:          86626.7     0.018713     0.018470     0.020643
Add:            88192.8     0.027540     0.027213     0.031095
Triad:          87655.3     0.027729     0.027380     0.031028
-------------------------------------------------------------
Solution Validates: avg error less than 1.000000e-13 on all three arrays
-------------------------------------------------------------

Here is a phoronix article comparing Ryzen AI 9 HX 370 with a variety of laptop systems. The overall geomean is ~10% but there is a wider variety between tests. Can be interesting to puzzle out why some of the differences. It is also likely that the power points used for the laptop comparisons in the phoronix article are less since I see lower scores e.g. coremark or different gaps than what I see with the same benchmark. So will need to puzzle out some of the SOC/power choices.

Posted in experiment, hardware | Tagged 7840HS, coremark, Ryzen AI 9 HX 370, stream, Zen5 | Leave a reply

Stream, experiments

Performance analysis, tools and experiments Posted on December 16, 2023 by mevDecember 17, 2023

I copied Stream from https://www.cs.virginia.edu/stream/ and put a copy in https://github.com/cycletourist/perf. This suggested the following compilation flags On my system with a Ryzen 7 7800X3D this results in the following performance: The question is what is the sensitivity of various … Continue reading →

Posted in experiment | Tagged stream | Leave a reply

Meta

  • Log in
  • Entries feed
  • Comments feed
  • WordPress.org

Archives

  • July 2026
  • November 2024
  • October 2024
  • September 2024
  • July 2024
  • June 2024
  • March 2024
  • February 2024
  • January 2024
  • December 2023
  • February 2023

Tags

7840HS Aarch64 bad data benchmarks cachyos cluster compiler coremark cpu2017 data fabric getrusage gnuplot i5-13500H icache ipc kernel l3 metrics namd opcache perf performance counters perf_event_open phoronix Ryzen AI 9 HX 370 Ryzen AI 365 scaling stream threshold topdown tree virtualization website wsl Zen5

Recent Posts

  • Performance Measurements on four systems
  • New scripts with Ryzen AI 9 HX 370
  • ARM box, inventory script
  • Virtualization comparisons
  • Updating to a new kernel and graphics driver
©2026 - Performance analysis, tools and experiments - Weaver Xtreme Theme
↑