↓
 

Performance analysis, tools and experiments

An eclectic collection

  • Overview
  • Blog
  • Workloads
    • cpu2017
      • 500.perlbench_r
      • 502.gcc_r
      • 503.bwaves_r
      • 505.mcf_r
      • 507.cactuBSSN_r
      • 508.namd_r
      • 510.parest_r
      • 511.povray_r
      • 519.lbm_r
      • 520.omnetpp_r
      • 521.wrf_r
      • 523.xalancbmk_r
      • 525.x264_r
      • 526.blender_r
      • 527.cam4_r
      • 531.deepsjeng_r
      • 538.imagick_r
      • 541.leela_r
      • 544.nab_r
      • 548.exchange2_r
      • 549.fotonik3d_r
      • 554.roms_r
      • 557.xz_r
    • geekbench
    • lmbench
    • passmark
    • pbbs
    • phoronix
      • ai-benchmark
      • aircrack-ng
      • amg
      • aobench
      • aom-av1
      • apache
      • apache-iotdb
      • appleseed
      • arrayfire
      • askap
      • asmfish
      • astcenc
      • avifenc
      • basis
      • blake2
      • blogbench
      • blender
      • blosc
      • bork
      • botan
      • brl-cad
      • build-apache
      • build-clash
      • build-eigen
      • build-erlang
      • build-ffmpeg
      • build-gcc
      • build-gdb
      • build-gem5
      • build-godot
      • build-imagemagick
      • build-linux-kernel
      • build-llvm
      • build-mesa
      • build-mplayer
      • build-nodejs
      • build-php
      • build-python
      • build-wasmer
      • build2
      • bullet
      • byte
      • cachebench
      • cassandra
      • clickhouse
      • clomp
      • cloverleaf
      • cockroach
      • compilebench
      • compress-7zip
      • compress-gzip
      • compress-lz4
      • compress-pbzip2
      • compress-rar
      • compress-xz
      • compress-zstd
      • core-latency
      • coremark
      • cp2k
      • cpp-perf-bench
      • cpuminer-opt
      • crafty
      • c-ray
      • cryptopp
      • cryptsetup
      • ctx-clock
      • cython-bench
      • dacapobench
      • daphne
      • darktable
      • dav1d
      • dbench
      • deepsparse
      • deepspeech
      • dolfyn
      • draco
      • dragonflydb
      • duckdb
      • easywave
      • ebizzy
      • embree
      • encode-flac
      • encode-mp3
      • encode-opus
      • encode-wavpack
      • espeak
      • etcpak
      • faiss
      • fast-cli
      • ffmpeg
      • ffte
      • fftw
      • fhourstones
      • financebench
      • furmark
      • gcrypt
      • gegl
      • gimp
      • git
      • glibc-bench
      • gmpbench
      • gnupg
      • gnuradio
      • go-benchmark
      • gpaw
      • graph500
      • graphics-magick
      • gromacs
      • hackbench
      • hadoop
      • heffte
      • helsing
      • himeno
      • hmmer
      • hpcg
      • incompact3d
      • indigobench
      • inkscape
      • ipc-benchmark
      • java-jmh
      • java-scimark2
      • john-the-ripper
      • jpegxl
      • jpegxl-decode
      • kvazaar
      • kripke
      • lammps
      • lczero
      • libraw
      • libreoffice
      • libxsmm
      • liquid-dsp
      • llama.cpp
      • llamafile
      • lulesh
      • lzbench
      • mbw
      • memcached
      • minibude
      • minife
      • mnn
      • mpcbench
      • m-queens
      • mrbayes
      • mutex
      • namd
      • mt-dgemm
      • ncnn
      • neat
      • nettle
      • nginx
      • ngspice
      • node-octane
      • node-web-tooling
      • npb
      • n-queens
      • numpy
      • nwchem
      • oidn
      • onednn
      • octave-benchmark
      • onnx
      • opencv
      • openfoam
      • openjpeg
      • openssl
      • openradioss
      • openscad
      • openvino
      • openvkl
      • ospray
      • ospray-studio
      • palabos
      • parboil
      • pennant
      • perl-benchmark
      • pgbench
      • phpbench
      • pjsip
      • polybench-c
      • polyhedron
      • povray
      • primesieve
      • pybench
      • pyhpc
      • pyperformance
      • pytorch
      • quadray
      • qe
      • qmcpack
      • quantlib
      • quicksilver
      • ramspeed
      • rav1e
      • rawtherapee
      • rbenchmark
      • redis
      • renaissance
      • rnnoise
      • rocksdb
      • rodinia
      • rsvg
      • schbench
      • scikit-learn
      • scimark2
      • scylladb
      • securemark
      • selenium
      • simdjson
      • smallpt
      • smhasher
      • spark
      • spark-tpcds
      • speedb
      • specfem3d
      • sqlite
      • srsran
      • stargate
      • stockfish
      • stream
      • stress-ng
      • svt-av1
      • svt-hevc
      • svt-vp9
      • sudokut
      • synthmark
      • sysbench
      • tensorflow
      • tensorflow-lite
      • tesseract
      • tjbench
      • tnn
      • toybrot
      • tscp
      • ttsiod-renderer
      • tungsten
      • uvg266
      • vkpeak
      • vpxenc
      • v-ray
      • vvenc
      • webp
      • webp2
      • whisper.cpp
      • whisperfile
      • wireguard
      • x264
      • x265
      • xmrig
      • xnnpack
      • y-cruncher
      • z3
    • stream
  • Tools
    • Compilers
    • likwid
    • perf
    • trace-cmd and kernelshark
    • wspy
  • Experiments
    • Histograms
    • clustering
    • Adding summary statistics for all benchmarks
  • Home
  • Blog
  • Workloads
    • cpu2017
      • 500.perlbench_r
      • 502.gcc_r
      • 503.bwaves_r
      • 505.mcf_r
      • 507.cactuBSSN_r
      • 508.namd_r
      • 510.parest_r
      • 511.povray_r
      • 519.lbm_r
      • 520.omnetpp_r
      • 521.wrf_r
      • 523.xalancbmk_r
      • 525.x264_r
      • 526.blender_r
      • 527.cam4_r
      • 531.deepsjeng_r
      • 538.imagick_r
      • 541.leela_r
      • 544.nab_r
      • 548.exchange2_r
      • 549.fotonik3d_r
      • 554.roms_r
      • 557.xz_r
    • geekbench
    • lmbench
    • passmark
    • pbbs
    • phoronix
      • ai-benchmark
      • aircrack-ng
      • amg
      • aobench
      • aom-av1
      • apache
      • apache-iotdb
      • appleseed
      • arrayfire
      • askap
      • asmfish
      • astcenc
      • avifenc
      • b
      • basis
      • blake2
      • blender
      • blogbench
      • blosc
      • bork
      • botan
      • brl-cad
      • build-apache
      • build-clash
      • build-eigen
      • build-erlang
      • build-ffmpeg
      • build-gcc
      • build-gdb
      • build-gem5
      • build-godot
      • build-imagemagick
      • build-linux-kernel
      • build-llvm
      • build-mesa
      • build-mplayer
      • build-nodejs
      • build-php
      • build-python
      • build-wasmer
      • build2
      • bullet
      • byte
      • c-ray
      • cachebench
      • cassandra
      • clickhouse
      • clomp
      • cloverleaf
      • cockroach
      • compilebench
      • compress-7zip
      • compress-gzip
      • compress-lz4
      • compress-pbzip2
      • compress-rar
      • compress-xz
      • compress-zstd
      • core-latency
      • coremark
      • cp2k
      • cpp-perf-bench
      • cpuminer-opt
      • crafty
      • cryptopp
      • cryptsetup
      • ctx-clock
      • cython-bench
      • dacapobench
      • daphne
      • darktable
      • dav1d
      • dbench
      • deepsparse
      • deepspeech
      • dolfyn
      • draco
      • dragonflydb
      • duckdb
      • easywave
      • ebizzy
      • embree
      • encode-flac
      • encode-mp3
      • encode-opus
      • encode-wavpack
      • espeak
      • etcpak
      • faiss
      • fast-cli
      • ffmpeg
      • ffte
      • fftw
      • fhourstones
      • financebench
      • furmark
      • gcrypt
      • gegl
      • gimp
      • git
      • glibc-bench
      • gmpbench
      • gnupg
      • gnuradio
      • go-benchmark
      • gpaw
      • graph500
      • graphics-magick
      • gromacs
      • hackbench
      • hadoop
      • heffte
      • helsing
      • himeno
      • hmmer
      • hpcg
      • incompact3d
      • indigobench
      • inkscape
      • ipc-benchmark
      • java-jmh
      • java-scimark2
      • john-the-ripper
      • jpegxl
      • jpegxl-decode
      • kripke
      • kvazaar
      • lammps
      • lczero
      • libraw
      • libreoffice
      • libxsmm
      • liquid-dsp
      • llama.cpp
      • llamafile
      • lulesh
      • lzbench
      • m-queens
      • mbw
      • memcached
      • minibude
      • minife
      • mnn
      • mpcbench
      • mrbayes
      • mt-dgemm
      • mutex
      • n-queens
      • namd
      • ncnn
      • neat
      • nettle
      • nginx
      • ngspice
      • node-octane
      • node-web-tooling
      • npb
      • numpy
      • nwchem
      • octave-benchmark
      • oidn
      • onednn
      • onnx
      • opencv
      • openfoam
      • openjpeg
      • openradioss
      • openscad
      • openssl
      • openvino
      • openvkl
      • ospray
      • ospray-studio
      • palabos
      • parboil
      • pennant
      • perl-benchmark
      • pgbench
      • phpbench
      • pjsip
      • polybench-c
      • polyhedron
      • povray
      • primesieve
      • pybench
      • pyhpc
      • pyperformance
      • pytorch
      • qe
      • qmcpack
      • quadray
      • quantlib
      • quicksilver
      • ramspeed
      • rav1e
      • rawtherapee
      • rays1bench
      • rbenchmark
      • redis
      • renaissance
      • rnnoise
      • rocksdb
      • rodinia
      • rsvg
      • schbench
      • scikit-learn
      • scimark2
      • scylladb
      • securemark
      • selenium
      • simdjson
      • smallpt
      • smhasher
      • spark
      • spark-tpcds
      • specfem3d
      • speedb
      • sqlite
      • srsran
      • stargate
      • stockfish
      • stream
      • stress-ng
      • sudokut
      • svt-av1
      • svt-hevc
      • svt-vp9
      • synthmark
      • sysbench
      • tensorflow
      • tensorflow-lite
      • tesseract
      • tjbench
      • tnn
      • toybrot
      • tscp
      • ttsiod-renderer
      • tungsten
      • uvg266
      • v-ray
      • vkpeak
      • vpxenc
      • vvenc
      • webp
      • webp2
      • whisper.cpp
      • whisperfile
      • wireguard
      • x264
      • x265
      • xmrig
      • xnnpack
      • y-cruncher
      • z3
    • stream
  • Tools
    • Compilers
    • likwid
    • perf
    • trace-cmd and kernelshark
    • wspy
  • Experiments
Home→Tags coremark

Tag Archives: coremark

New scripts with Ryzen AI 9 HX 370

Performance analysis, tools and experiments Posted on July 12, 2026 by mevJuly 12, 2026

With my new topology aware scripts I went back to my Ryzen AI 9 HX 370 system to see how things mapped.

Below is output of the topology script

================================================================================
CPU CORE & CACHE HIERARCHY MAP (X86_64)
================================================================================
Total Cores: 24

Core Details Table:
Core   Core Model                   Stepping   Max Freq     L1I/L1D Cache    L2 Cache  
----------------------------------------------------------------------------------------
cpu0   AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 5158 MHz     32K / 48K        1024K     
cpu1   AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 5158 MHz     32K / 48K        1024K     
cpu2   AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 5158 MHz     32K / 48K        1024K     
cpu3   AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 5158 MHz     32K / 48K        1024K     
cpu4   AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 3289 MHz     32K / 48K        1024K     
cpu5   AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 3289 MHz     32K / 48K        1024K     
cpu6   AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 3289 MHz     32K / 48K        1024K     
cpu7   AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 3289 MHz     32K / 48K        1024K     
cpu8   AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 3289 MHz     32K / 48K        1024K     
cpu9   AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 3289 MHz     32K / 48K        1024K     
cpu10  AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 3289 MHz     32K / 48K        1024K     
cpu11  AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 3289 MHz     32K / 48K        1024K     
cpu12  AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 5158 MHz     32K / 48K        1024K     
cpu13  AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 5158 MHz     32K / 48K        1024K     
cpu14  AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 5158 MHz     32K / 48K        1024K     
cpu15  AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 5158 MHz     32K / 48K        1024K     
cpu16  AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 3289 MHz     32K / 48K        1024K     
cpu17  AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 3289 MHz     32K / 48K        1024K     
cpu18  AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 3289 MHz     32K / 48K        1024K     
cpu19  AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 3289 MHz     32K / 48K        1024K     
cpu20  AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 3289 MHz     32K / 48K        1024K     
cpu21  AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 3289 MHz     32K / 48K        1024K     
cpu22  AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 3289 MHz     32K / 48K        1024K     
cpu23  AMD Ryzen AI 9 HX 370 w/ Radeon 890M stepping 0, ucode 0xb20401b 3289 MHz     32K / 48K        1024K     

System-wide Cache Capacity Summary:
 - L1 Data        Cache: 576 KiB      (Total across 12x 48K)
 - L1 Instruction Cache: 384 KiB      (Total across 12x 32K)
 - L2 Unified     Cache: 12.0 MiB     (Total across 12x 1024K)
 - L3 Unified     Cache: 24.0 MiB     (Total across 1x 16384K, 1x 8192K)

================================================================================
TOPOLOGY PLUMBING TREE
================================================================================
L3 Unified Cache [16384K]
├── L2 Unified Cache [1024K]
│   ├── L1 Data Cache [48K]
│   │   ├── Core 15: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 5.16 GHz
│   │   └── Core 3: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 5.16 GHz
│   └── L1 Instruction Cache [32K]
├── L2 Unified Cache [1024K]
│   ├── L1 Data Cache [48K]
│   │   ├── Core 14: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 5.16 GHz
│   │   └── Core 2: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 5.16 GHz
│   └── L1 Instruction Cache [32K]
├── L2 Unified Cache [1024K]
│   ├── L1 Data Cache [48K]
│   │   ├── Core 13: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 5.16 GHz
│   │   └── Core 1: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 5.16 GHz
│   └── L1 Instruction Cache [32K]
└── L2 Unified Cache [1024K]
    ├── L1 Data Cache [48K]
    │   ├── Core 12: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 5.16 GHz
    │   └── Core 0: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 5.16 GHz
    └── L1 Instruction Cache [32K]
L3 Unified Cache [8192K]
├── L2 Unified Cache [1024K]
│   ├── L1 Data Cache [48K]
│   │   ├── Core 23: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 3.29 GHz
│   │   └── Core 11: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 3.29 GHz
│   └── L1 Instruction Cache [32K]
├── L2 Unified Cache [1024K]
│   ├── L1 Data Cache [48K]
│   │   ├── Core 22: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 3.29 GHz
│   │   └── Core 10: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 3.29 GHz
│   └── L1 Instruction Cache [32K]
├── L2 Unified Cache [1024K]
│   ├── L1 Data Cache [48K]
│   │   ├── Core 21: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 3.29 GHz
│   │   └── Core 9: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 3.29 GHz
│   └── L1 Instruction Cache [32K]
├── L2 Unified Cache [1024K]
│   ├── L1 Data Cache [48K]
│   │   ├── Core 20: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 3.29 GHz
│   │   └── Core 8: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 3.29 GHz
│   └── L1 Instruction Cache [32K]
├── L2 Unified Cache [1024K]
│   ├── L1 Data Cache [48K]
│   │   ├── Core 19: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 3.29 GHz
│   │   └── Core 7: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 3.29 GHz
│   └── L1 Instruction Cache [32K]
├── L2 Unified Cache [1024K]
│   ├── L1 Data Cache [48K]
│   │   ├── Core 18: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 3.29 GHz
│   │   └── Core 6: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 3.29 GHz
│   └── L1 Instruction Cache [32K]
├── L2 Unified Cache [1024K]
│   ├── L1 Data Cache [48K]
│   │   ├── Core 17: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 3.29 GHz
│   │   └── Core 5: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 3.29 GHz
│   └── L1 Instruction Cache [32K]
└── L2 Unified Cache [1024K]
    ├── L1 Data Cache [48K]
    │   ├── Core 16: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 3.29 GHz
    │   └── Core 4: AMD Ryzen AI 9 HX 370 w/ Radeon 890M (stepping 0, ucode 0xb20401b) @ 3.29 GHz
    └── L1 Instruction Cache [32K]

One thing that surprised me was previously likwid-topology had given a different mapping of L3 topology. Apparently, there are two core complexes. One has 4 Zen5 cores (8 threads) and 16MB of L3 and the other has 8 Zen5 compact cores (16 threads) and only 8MB of L3. That is consistent with specs from AMD so a spot where likwid-topology isn’t quite complete. I can probably update the topology report above to reflect hyperthreading but otherwise useful additional information.

Similar to previous ARM experiment, I tried a stream sweep to find best configuration

triad_MBps	opt_level	threads	strategy	domain	cpus	run
74213.600000	O2	2	spread_l3	rr	0,4	1
72749.400000	Ofast	2	spread_l3	rr	0,4	1
72687.900000	O3	2	spread_l3	rr	0,4	1
71942.400000	O2	3	spread_l3	rr	0,4,1	1
71505.000000	O2	4	spread_l3	rr	0,4,1,5	1
71500.900000	Ofast	3	spread_l3	rr	0,4,1	1
71464.400000	O3	3	spread_l3	rr	0,4,1	1
71460.800000	Ofast	4	spread_l3	rr	0,4,1,5	1
71403.000000	O3	4	spread_l3	rr	0,4,1,5	1
71041.700000	O2	2	local_l3	0	0,1	1
69966.800000	O2	2	local_l3	1	4,5	1
69621.300000	O2	3	local_l3	0	0,1,2	1
69486.700000	O3	2	local_l3	0	0,1	1
69432.500000	Ofast	2	local_l3	0	0,1	1
69296.300000	O3	3	local_l3	0	0,1,2	1
69168.200000	Ofast	3	local_l3	0	0,1,2	1
69074.700000	O2	3	local_l3	1	4,5,6	1
68989.500000	O2	4	local_l3	0	0,1,2,3	1
68979.600000	O2	4	local_l3	1	4,5,6,7	1
68943.700000	Ofast	3	local_l3	1	4,5,6	1

In this case, the fastest triad comes from using one thread from the Zen5 cores and one from the Zen5c cores, each with their own L3.

coremark_average	scenario	domain	threads	cpus	run
641556.737750	all_logical	all	24	0,1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23	1
465524.800224	thread0_only	all	12	0,1,2,3,4,5,6,7,8,9,10,11	1
385663.999579	capability_group	1	16	4,5,6,7,8,9,10,11,16,17,18,19,20,21,22,23	1
381649.327648	l3_complex	1	16	4,5,6,7,8,9,10,11,16,17,18,19,20,21,22,23	1
307810.051852	l3_complex	0	8	0,1,2,3,12,13,14,15	1
305829.027379	capability_group	0	8	0,1,2,3,12,13,14,15	1

Looking at coremark across the different types of cores isn’t surprising.

  • Running on all cores gives greatest throughput.
  • Running only one thread per core (hyperthreading) gets about 72.5% of the performance.
  • The “capability group” (Zen5 vs Zen5c) and “l3 complex” (those with the 16MB L3 and those with the 8 MB L3) are the same sets so results are the same. The performance of 8 ZenC cores (16 threads) still more than performance of 4 Zen cores (8 threads).

Posted in experiment, hardware | Tagged coremark, Ryzen AI 9 HX 370, stream | Leave a reply

ARM box, inventory script

Performance analysis, tools and experiments Posted on July 12, 2026 by mevJuly 12, 2026

I have a Minisforum MS-R1 system that will let me experiment further with Aarch 64 cores on Linux. To help me with this discovery, I extended several scripts to inventory and examine the system so will also show those scripts outputs on this system and others to calibrate.

The first script generates the topology of both cores and cache on the system from /sys entries:

================================================================================
CPU CORE & CACHE HIERARCHY MAP (AARCH64)
================================================================================
Total Cores: 12

Core Details Table:
Core   Core Model                   Stepping   Max Freq     L1I/L1D Cache    L2 Cache  
----------------------------------------------------------------------------------------
cpu0   ARM Cortex-A720              r0p1       2600 MHz     64K / 64K        512K      
cpu1   ARM Cortex-A720              r0p1       2600 MHz     64K / 64K        512K      
cpu2   ARM Cortex-A520              r0p1       1800 MHz     32K / 32K        N/A       
cpu3   ARM Cortex-A520              r0p1       1800 MHz     32K / 32K        N/A       
cpu4   ARM Cortex-A520              r0p1       1800 MHz     32K / 32K        N/A       
cpu5   ARM Cortex-A520              r0p1       1800 MHz     32K / 32K        N/A       
cpu6   ARM Cortex-A720              r0p1       2300 MHz     64K / 64K        512K      
cpu7   ARM Cortex-A720              r0p1       2300 MHz     64K / 64K        512K      
cpu8   ARM Cortex-A720              r0p1       2200 MHz     64K / 64K        512K      
cpu9   ARM Cortex-A720              r0p1       2200 MHz     64K / 64K        512K      
cpu10  ARM Cortex-A720              r0p1       2500 MHz     64K / 64K        512K      
cpu11  ARM Cortex-A720              r0p1       2500 MHz     64K / 64K        512K      

System-wide Cache Capacity Summary:
 - L1 Data        Cache: 640 KiB      (Total across 4x 32K, 8x 64K)
 - L1 Instruction Cache: 640 KiB      (Total across 4x 32K, 8x 64K)
 - L2 Unified     Cache: 4.0 MiB      (Total across 8x 512K, 4x unreported)
 - L3 Unified     Cache: 12.0 MiB     (Total across 1x 12288K)

================================================================================
TOPOLOGY PLUMBING TREE
================================================================================
L3 Unified Cache [12288K]
├── Core 11: ARM Cortex-A720 (r0p1) @ 2.50 GHz
│   └─ Private Caches: L1 Data (64K), L1 Instruction (64K), L2 Unified (512K)
├── Core 10: ARM Cortex-A720 (r0p1) @ 2.50 GHz
│   └─ Private Caches: L1 Data (64K), L1 Instruction (64K), L2 Unified (512K)
├── Core 9: ARM Cortex-A720 (r0p1) @ 2.20 GHz
│   └─ Private Caches: L1 Data (64K), L1 Instruction (64K), L2 Unified (512K)
├── Core 8: ARM Cortex-A720 (r0p1) @ 2.20 GHz
│   └─ Private Caches: L1 Data (64K), L1 Instruction (64K), L2 Unified (512K)
├── Core 7: ARM Cortex-A720 (r0p1) @ 2.30 GHz
│   └─ Private Caches: L1 Data (64K), L1 Instruction (64K), L2 Unified (512K)
├── Core 6: ARM Cortex-A720 (r0p1) @ 2.30 GHz
│   └─ Private Caches: L1 Data (64K), L1 Instruction (64K), L2 Unified (512K)
├── Core 5: ARM Cortex-A520 (r0p1) @ 1.80 GHz
│   └─ Private Caches: L1 Data (32K), L1 Instruction (32K)
├── Core 4: ARM Cortex-A520 (r0p1) @ 1.80 GHz
│   └─ Private Caches: L1 Data (32K), L1 Instruction (32K)
├── Core 3: ARM Cortex-A520 (r0p1) @ 1.80 GHz
│   └─ Private Caches: L1 Data (32K), L1 Instruction (32K)
├── Core 2: ARM Cortex-A520 (r0p1) @ 1.80 GHz
│   └─ Private Caches: L1 Data (32K), L1 Instruction (32K)
├── Core 1: ARM Cortex-A720 (r0p1) @ 2.60 GHz
│   └─ Private Caches: L1 Data (64K), L1 Instruction (64K), L2 Unified (512K)
└── Core 0: ARM Cortex-A720 (r0p1) @ 2.60 GHz
    └─ Private Caches: L1 Data (64K), L1 Instruction (64K), L2 Unified (512K)

This shows this SOC has a mix of both Cortex A720 cores and Cortex A520 cores and even the Cortex A720 cores have different maximum frequencies reflecting possible different between “high” and “mid” cores.

The first thing we do on this new system is run a sweep of stream pinned to different cores. I updated a script to consider the core capability, the optimization level and the number of threads to create the following measurements:

triad_MBps	opt_level	threads	strategy	domain	cpus	run
39170.300000	O2	4	local_l3	0	0,1,10,11	1
38441.200000	O2	2	local_l3	0	0,1	1
37068.000000	O2	2	local_l3_cc	0	0,10	1
36531.400000	O2	3	local_l3	0	0,1,10	1
35612.500000	O2	2	local_l3_cc	0	0,6	1
34266.600000	O2	2	local_l3_cc	0	0,11	1
33733.900000	O2	2	local_l3_cc	0	0,7	1
26004.700000	O2	1	local_l3	0	0	1

The highest stream triad comes from running four threads and pinning to the four highest frequency Cortex 720 cores. Just using the two fastest cores comes close.

The next thing we do is try coremark for each individual core

cpu	run	status	coremark_average	log_file
0	1	OK	25512.373511	/home/mev/source/perf/results/coremark_each_core/coremark.cpu0.run1.txt
1	1	OK	25511.108133	/home/mev/source/perf/results/coremark_each_core/coremark.cpu1.run1.txt
2	1	OK	8182.294456	/home/mev/source/perf/results/coremark_each_core/coremark.cpu2.run1.txt
3	1	OK	8168.957901	/home/mev/source/perf/results/coremark_each_core/coremark.cpu3.run1.txt
4	1	OK	8172.910173	/home/mev/source/perf/results/coremark_each_core/coremark.cpu4.run1.txt
5	1	OK	8176.058957	/home/mev/source/perf/results/coremark_each_core/coremark.cpu5.run1.txt
6	1	OK	22462.563490	/home/mev/source/perf/results/coremark_each_core/coremark.cpu6.run1.txt
7	1	OK	22442.589501	/home/mev/source/perf/results/coremark_each_core/coremark.cpu7.run1.txt
8	1	OK	21470.071023	/home/mev/source/perf/results/coremark_each_core/coremark.cpu8.run1.txt
9	1	OK	21462.437238	/home/mev/source/perf/results/coremark_each_core/coremark.cpu9.run1.txt
10	1	OK	24509.139361	/home/mev/source/perf/results/coremark_each_core/coremark.cpu10.run1.txt
11	1	OK	24504.245930	/home/mev/source/perf/results/coremark_each_core/coremark.cpu11.run1.txt

This shows us results that are consistent with the cores and their maximum frequencies. It surprises me how much quicker a Cortex 720 is than a Cortex 520, so some strategy that pins to the 8 Cortex 720 cores might make sense in some situations.

coremark_average	scenario	domain	threads	cpus	run
172518.777163	thread0_only	all	12	0,1,2,3,4,5,6,7,8,9,10,11	1
168110.086124	all_logical	all	12	0,1,2,3,4,5,6,7,8,9,10,11	1
51064.558269	capability_group	0	2	0,1	1
49082.209232	capability_group	1	2	10,11	1
44976.394529	capability_group	2	2	6,7	1
43020.862299	capability_group	3	2	8,9	1
33342.065373	capability_group	4	4	2,3,4,5	1

I also tried some aggregation of running on all cores, on all thread 0 (hyperthread) and for each pair of cores that have the same capabilities. The first two numbers are close because they are the same run. The pairs of Cortex cores by themselves have coremark scores according to frequencies. The four Cortex 520 cores by themselves are slower than any pair of Cortex 720 cores.

As a follow on post, I will also document what these same scripts have shown with a Ryzen AI 9 370 also show.

Posted in experiment, hardware, Tools | Tagged Aarch64, coremark, stream | Leave a reply

Virtualization comparisons

Performance analysis, tools and experiments Posted on November 13, 2024 by mevNovember 13, 2024

Installing and reinstalling Operating Systems can be easier to do if I maintain several virtual machines with each configuration. While this lets me compare VMs and OSs against each other, there is also a question on how the virtual environment compares against the host environment. So I’ve created a few configurations I can use for these comparisons. In particular:

NameThreadsMemoryNotes
boulder1632GBHost: 7840HS, Zen4; ubuntu 24.04
niwot2432GBHost: RX 370, Zen5; ubuntu 24.04
boulder-ubuntu816GBubuntu 24.04 guest
boulder-cachyos816GBcachyos guest
boulder “constrainted”832 GBhost with taskset –cpulist
niwot-ubuntu1216GBubuntu 24.04 guest
niwot-cachyos1216GBcachyos guest
niwot “constrained”1232GBhost with taskset –cpulist

Since I can’t dedicate the entire machine to the VM, I instead bind the VM to run with one thread bound per (hyper-threaded host) core. I also define half the memory. I can then compare this against a host “constrained” configuration that also runs on those same cores, e.g.

taskset --cpu-list 0-15:2 phoronix-test-suite batch-run coremark

The first benchmark I pick for such a comparison is coremark.

NameThreadsScore
boulder16412415
niwot24563857
boulder-ubuntu8296640
boulder-cachyos8310674
boulder “constrained”8317576
niwot-ubuntu12356810
niwot-cachyos12369503
niwot “constrained”12401518

First thing to note is that the 7840 “constrained” configuration runs at 77% of the full host configuration (317576/412415) while the 370 “constrained” configuration runs at 71% (401518/563857) so running half the cores isn’t quite as large for the 370.

Next thing to notice is the Ubuntu virtual machine performance of 7840 is 93% of constrained while 370 is 88% of constrained. The net effect is the host only benchmark is 1.37x faster on 370 than 7840 but the virtual machine is only 1.20x faster. CachyOS is faster and hence it is 98% of host on 7840 and 93% of host on 370.

This is only one benchmark so will also be useful to cross-check how much these trends also apply to other workloads. I can probably also separate this to see how much the “constrained” matches the full system and then see what the virtualization overhead as two separate comparisons.

Posted in experiment | Tagged coremark, virtualization | Leave a reply

New Ryzen AI 9 HX 370 machine

Performance analysis, tools and experiments Posted on October 8, 2024 by mevOctober 10, 2024

I have a new AMD performance machine for experiments. The processor is a Ryzen AI 9 HX 370 in a Beelink SER9 mini-PC. Following are some of the major parameters.in comparison with my Ryzen 7840HS comparison machine. Following are the … Continue reading →

Posted in experiment, hardware | Tagged 7840HS, coremark, Ryzen AI 9 HX 370, stream, Zen5 | Leave a reply

Coremark scaling 7840HS

Performance analysis, tools and experiments Posted on September 27, 2024 by mevSeptember 27, 2024

The following chart shows the Phoronix test suite coremark value when running from 1 to 16 cores. Graphically it looks as follows The question is what causes the inflection points on the graph? The scaling from 1-8 cores decreases only … Continue reading →

Posted in experiment | Tagged 7840HS, coremark, scaling | Leave a reply

Meta

  • Log in
  • Entries feed
  • Comments feed
  • WordPress.org

Archives

  • July 2026
  • November 2024
  • October 2024
  • September 2024
  • July 2024
  • June 2024
  • March 2024
  • February 2024
  • January 2024
  • December 2023
  • February 2023

Tags

7840HS Aarch64 bad data benchmarks cachyos cluster compiler coremark cpu2017 data fabric getrusage gnuplot i5-13500H icache ipc kernel l3 metrics namd opcache perf performance counters perf_event_open phoronix Ryzen AI 9 HX 370 Ryzen AI 365 scaling stream threshold topdown tree virtualization website wsl Zen5

Recent Posts

  • Performance Measurements on four systems
  • New scripts with Ryzen AI 9 HX 370
  • ARM box, inventory script
  • Virtualization comparisons
  • Updating to a new kernel and graphics driver
©2026 - Performance analysis, tools and experiments - Weaver Xtreme Theme
↑