↓
 

Performance analysis, tools and experiments

An eclectic collection

  • Overview
  • Blog
  • Workloads
    • cpu2017
      • 500.perlbench_r
      • 502.gcc_r
      • 503.bwaves_r
      • 505.mcf_r
      • 507.cactuBSSN_r
      • 508.namd_r
      • 510.parest_r
      • 511.povray_r
      • 519.lbm_r
      • 520.omnetpp_r
      • 521.wrf_r
      • 523.xalancbmk_r
      • 525.x264_r
      • 526.blender_r
      • 527.cam4_r
      • 531.deepsjeng_r
      • 538.imagick_r
      • 541.leela_r
      • 544.nab_r
      • 548.exchange2_r
      • 549.fotonik3d_r
      • 554.roms_r
      • 557.xz_r
    • geekbench
    • lmbench
    • passmark
    • pbbs
    • phoronix
      • ai-benchmark
      • aircrack-ng
      • amg
      • aobench
      • aom-av1
      • apache
      • apache-iotdb
      • appleseed
      • arrayfire
      • askap
      • asmfish
      • astcenc
      • avifenc
      • basis
      • blake2
      • blogbench
      • blender
      • blosc
      • bork
      • botan
      • brl-cad
      • build-apache
      • build-clash
      • build-eigen
      • build-erlang
      • build-ffmpeg
      • build-gcc
      • build-gdb
      • build-gem5
      • build-godot
      • build-imagemagick
      • build-linux-kernel
      • build-llvm
      • build-mesa
      • build-mplayer
      • build-nodejs
      • build-php
      • build-python
      • build-wasmer
      • build2
      • bullet
      • byte
      • cachebench
      • cassandra
      • clickhouse
      • clomp
      • cloverleaf
      • cockroach
      • compilebench
      • compress-7zip
      • compress-gzip
      • compress-lz4
      • compress-pbzip2
      • compress-rar
      • compress-xz
      • compress-zstd
      • core-latency
      • coremark
      • cp2k
      • cpp-perf-bench
      • cpuminer-opt
      • crafty
      • c-ray
      • cryptopp
      • cryptsetup
      • ctx-clock
      • cython-bench
      • dacapobench
      • daphne
      • darktable
      • dav1d
      • dbench
      • deepsparse
      • deepspeech
      • dolfyn
      • draco
      • dragonflydb
      • duckdb
      • easywave
      • ebizzy
      • embree
      • encode-flac
      • encode-mp3
      • encode-opus
      • encode-wavpack
      • espeak
      • etcpak
      • faiss
      • fast-cli
      • ffmpeg
      • ffte
      • fftw
      • fhourstones
      • financebench
      • furmark
      • gcrypt
      • gegl
      • gimp
      • git
      • glibc-bench
      • gmpbench
      • gnupg
      • gnuradio
      • go-benchmark
      • gpaw
      • graph500
      • graphics-magick
      • gromacs
      • hackbench
      • hadoop
      • heffte
      • helsing
      • himeno
      • hmmer
      • hpcg
      • incompact3d
      • indigobench
      • inkscape
      • ipc-benchmark
      • java-jmh
      • java-scimark2
      • john-the-ripper
      • jpegxl
      • jpegxl-decode
      • kvazaar
      • kripke
      • lammps
      • lczero
      • libraw
      • libreoffice
      • libxsmm
      • liquid-dsp
      • llama.cpp
      • llamafile
      • lulesh
      • lzbench
      • mbw
      • memcached
      • minibude
      • minife
      • mnn
      • mpcbench
      • m-queens
      • mrbayes
      • mutex
      • namd
      • mt-dgemm
      • ncnn
      • neat
      • nettle
      • nginx
      • ngspice
      • node-octane
      • node-web-tooling
      • npb
      • n-queens
      • numpy
      • nwchem
      • oidn
      • onednn
      • octave-benchmark
      • onnx
      • opencv
      • openfoam
      • openjpeg
      • openssl
      • openradioss
      • openscad
      • openvino
      • openvkl
      • ospray
      • ospray-studio
      • palabos
      • parboil
      • pennant
      • perl-benchmark
      • pgbench
      • phpbench
      • pjsip
      • polybench-c
      • polyhedron
      • povray
      • primesieve
      • pybench
      • pyhpc
      • pyperformance
      • pytorch
      • quadray
      • qe
      • qmcpack
      • quantlib
      • quicksilver
      • ramspeed
      • rav1e
      • rawtherapee
      • rbenchmark
      • redis
      • renaissance
      • rnnoise
      • rocksdb
      • rodinia
      • rsvg
      • schbench
      • scikit-learn
      • scimark2
      • scylladb
      • securemark
      • selenium
      • simdjson
      • smallpt
      • smhasher
      • spark
      • spark-tpcds
      • speedb
      • specfem3d
      • sqlite
      • srsran
      • stargate
      • stockfish
      • stream
      • stress-ng
      • svt-av1
      • svt-hevc
      • svt-vp9
      • sudokut
      • synthmark
      • sysbench
      • tensorflow
      • tensorflow-lite
      • tesseract
      • tjbench
      • tnn
      • toybrot
      • tscp
      • ttsiod-renderer
      • tungsten
      • uvg266
      • vkpeak
      • vpxenc
      • v-ray
      • vvenc
      • webp
      • webp2
      • whisper.cpp
      • whisperfile
      • wireguard
      • x264
      • x265
      • xmrig
      • xnnpack
      • y-cruncher
      • z3
    • stream
  • Tools
    • Compilers
    • likwid
    • perf
    • trace-cmd and kernelshark
    • wspy
  • Experiments
    • Histograms
    • clustering
    • Adding summary statistics for all benchmarks
  • Home
  • Blog
  • Workloads
    • cpu2017
      • 500.perlbench_r
      • 502.gcc_r
      • 503.bwaves_r
      • 505.mcf_r
      • 507.cactuBSSN_r
      • 508.namd_r
      • 510.parest_r
      • 511.povray_r
      • 519.lbm_r
      • 520.omnetpp_r
      • 521.wrf_r
      • 523.xalancbmk_r
      • 525.x264_r
      • 526.blender_r
      • 527.cam4_r
      • 531.deepsjeng_r
      • 538.imagick_r
      • 541.leela_r
      • 544.nab_r
      • 548.exchange2_r
      • 549.fotonik3d_r
      • 554.roms_r
      • 557.xz_r
    • geekbench
    • lmbench
    • passmark
    • pbbs
    • phoronix
      • ai-benchmark
      • aircrack-ng
      • amg
      • aobench
      • aom-av1
      • apache
      • apache-iotdb
      • appleseed
      • arrayfire
      • askap
      • asmfish
      • astcenc
      • avifenc
      • b
      • basis
      • blake2
      • blender
      • blogbench
      • blosc
      • bork
      • botan
      • brl-cad
      • build-apache
      • build-clash
      • build-eigen
      • build-erlang
      • build-ffmpeg
      • build-gcc
      • build-gdb
      • build-gem5
      • build-godot
      • build-imagemagick
      • build-linux-kernel
      • build-llvm
      • build-mesa
      • build-mplayer
      • build-nodejs
      • build-php
      • build-python
      • build-wasmer
      • build2
      • bullet
      • byte
      • c-ray
      • cachebench
      • cassandra
      • clickhouse
      • clomp
      • cloverleaf
      • cockroach
      • compilebench
      • compress-7zip
      • compress-gzip
      • compress-lz4
      • compress-pbzip2
      • compress-rar
      • compress-xz
      • compress-zstd
      • core-latency
      • coremark
      • cp2k
      • cpp-perf-bench
      • cpuminer-opt
      • crafty
      • cryptopp
      • cryptsetup
      • ctx-clock
      • cython-bench
      • dacapobench
      • daphne
      • darktable
      • dav1d
      • dbench
      • deepsparse
      • deepspeech
      • dolfyn
      • draco
      • dragonflydb
      • duckdb
      • easywave
      • ebizzy
      • embree
      • encode-flac
      • encode-mp3
      • encode-opus
      • encode-wavpack
      • espeak
      • etcpak
      • faiss
      • fast-cli
      • ffmpeg
      • ffte
      • fftw
      • fhourstones
      • financebench
      • furmark
      • gcrypt
      • gegl
      • gimp
      • git
      • glibc-bench
      • gmpbench
      • gnupg
      • gnuradio
      • go-benchmark
      • gpaw
      • graph500
      • graphics-magick
      • gromacs
      • hackbench
      • hadoop
      • heffte
      • helsing
      • himeno
      • hmmer
      • hpcg
      • incompact3d
      • indigobench
      • inkscape
      • ipc-benchmark
      • java-jmh
      • java-scimark2
      • john-the-ripper
      • jpegxl
      • jpegxl-decode
      • kripke
      • kvazaar
      • lammps
      • lczero
      • libraw
      • libreoffice
      • libxsmm
      • liquid-dsp
      • llama.cpp
      • llamafile
      • lulesh
      • lzbench
      • m-queens
      • mbw
      • memcached
      • minibude
      • minife
      • mnn
      • mpcbench
      • mrbayes
      • mt-dgemm
      • mutex
      • n-queens
      • namd
      • ncnn
      • neat
      • nettle
      • nginx
      • ngspice
      • node-octane
      • node-web-tooling
      • npb
      • numpy
      • nwchem
      • octave-benchmark
      • oidn
      • onednn
      • onnx
      • opencv
      • openfoam
      • openjpeg
      • openradioss
      • openscad
      • openssl
      • openvino
      • openvkl
      • ospray
      • ospray-studio
      • palabos
      • parboil
      • pennant
      • perl-benchmark
      • pgbench
      • phpbench
      • pjsip
      • polybench-c
      • polyhedron
      • povray
      • primesieve
      • pybench
      • pyhpc
      • pyperformance
      • pytorch
      • qe
      • qmcpack
      • quadray
      • quantlib
      • quicksilver
      • ramspeed
      • rav1e
      • rawtherapee
      • rays1bench
      • rbenchmark
      • redis
      • renaissance
      • rnnoise
      • rocksdb
      • rodinia
      • rsvg
      • schbench
      • scikit-learn
      • scimark2
      • scylladb
      • securemark
      • selenium
      • simdjson
      • smallpt
      • smhasher
      • spark
      • spark-tpcds
      • specfem3d
      • speedb
      • sqlite
      • srsran
      • stargate
      • stockfish
      • stream
      • stress-ng
      • sudokut
      • svt-av1
      • svt-hevc
      • svt-vp9
      • synthmark
      • sysbench
      • tensorflow
      • tensorflow-lite
      • tesseract
      • tjbench
      • tnn
      • toybrot
      • tscp
      • ttsiod-renderer
      • tungsten
      • uvg266
      • v-ray
      • vkpeak
      • vpxenc
      • vvenc
      • webp
      • webp2
      • whisper.cpp
      • whisperfile
      • wireguard
      • x264
      • x265
      • xmrig
      • xnnpack
      • y-cruncher
      • z3
    • stream
  • Tools
    • Compilers
    • likwid
    • perf
    • trace-cmd and kernelshark
    • wspy
  • Experiments
Home→Categories Tools

Category Archives: Tools

ARM box, inventory script

Performance analysis, tools and experiments Posted on July 12, 2026 by mevJuly 12, 2026

I have a Minisforum MS-R1 system that will let me experiment further with Aarch 64 cores on Linux. To help me with this discovery, I extended several scripts to inventory and examine the system so will also show those scripts outputs on this system and others to calibrate.

The first script generates the topology of both cores and cache on the system from /sys entries:

================================================================================
CPU CORE & CACHE HIERARCHY MAP (AARCH64)
================================================================================
Total Cores: 12

Core Details Table:
Core   Core Model                   Stepping   Max Freq     L1I/L1D Cache    L2 Cache  
----------------------------------------------------------------------------------------
cpu0   ARM Cortex-A720              r0p1       2600 MHz     64K / 64K        512K      
cpu1   ARM Cortex-A720              r0p1       2600 MHz     64K / 64K        512K      
cpu2   ARM Cortex-A520              r0p1       1800 MHz     32K / 32K        N/A       
cpu3   ARM Cortex-A520              r0p1       1800 MHz     32K / 32K        N/A       
cpu4   ARM Cortex-A520              r0p1       1800 MHz     32K / 32K        N/A       
cpu5   ARM Cortex-A520              r0p1       1800 MHz     32K / 32K        N/A       
cpu6   ARM Cortex-A720              r0p1       2300 MHz     64K / 64K        512K      
cpu7   ARM Cortex-A720              r0p1       2300 MHz     64K / 64K        512K      
cpu8   ARM Cortex-A720              r0p1       2200 MHz     64K / 64K        512K      
cpu9   ARM Cortex-A720              r0p1       2200 MHz     64K / 64K        512K      
cpu10  ARM Cortex-A720              r0p1       2500 MHz     64K / 64K        512K      
cpu11  ARM Cortex-A720              r0p1       2500 MHz     64K / 64K        512K      

System-wide Cache Capacity Summary:
 - L1 Data        Cache: 640 KiB      (Total across 4x 32K, 8x 64K)
 - L1 Instruction Cache: 640 KiB      (Total across 4x 32K, 8x 64K)
 - L2 Unified     Cache: 4.0 MiB      (Total across 8x 512K, 4x unreported)
 - L3 Unified     Cache: 12.0 MiB     (Total across 1x 12288K)

================================================================================
TOPOLOGY PLUMBING TREE
================================================================================
L3 Unified Cache [12288K]
├── Core 11: ARM Cortex-A720 (r0p1) @ 2.50 GHz
│   └─ Private Caches: L1 Data (64K), L1 Instruction (64K), L2 Unified (512K)
├── Core 10: ARM Cortex-A720 (r0p1) @ 2.50 GHz
│   └─ Private Caches: L1 Data (64K), L1 Instruction (64K), L2 Unified (512K)
├── Core 9: ARM Cortex-A720 (r0p1) @ 2.20 GHz
│   └─ Private Caches: L1 Data (64K), L1 Instruction (64K), L2 Unified (512K)
├── Core 8: ARM Cortex-A720 (r0p1) @ 2.20 GHz
│   └─ Private Caches: L1 Data (64K), L1 Instruction (64K), L2 Unified (512K)
├── Core 7: ARM Cortex-A720 (r0p1) @ 2.30 GHz
│   └─ Private Caches: L1 Data (64K), L1 Instruction (64K), L2 Unified (512K)
├── Core 6: ARM Cortex-A720 (r0p1) @ 2.30 GHz
│   └─ Private Caches: L1 Data (64K), L1 Instruction (64K), L2 Unified (512K)
├── Core 5: ARM Cortex-A520 (r0p1) @ 1.80 GHz
│   └─ Private Caches: L1 Data (32K), L1 Instruction (32K)
├── Core 4: ARM Cortex-A520 (r0p1) @ 1.80 GHz
│   └─ Private Caches: L1 Data (32K), L1 Instruction (32K)
├── Core 3: ARM Cortex-A520 (r0p1) @ 1.80 GHz
│   └─ Private Caches: L1 Data (32K), L1 Instruction (32K)
├── Core 2: ARM Cortex-A520 (r0p1) @ 1.80 GHz
│   └─ Private Caches: L1 Data (32K), L1 Instruction (32K)
├── Core 1: ARM Cortex-A720 (r0p1) @ 2.60 GHz
│   └─ Private Caches: L1 Data (64K), L1 Instruction (64K), L2 Unified (512K)
└── Core 0: ARM Cortex-A720 (r0p1) @ 2.60 GHz
    └─ Private Caches: L1 Data (64K), L1 Instruction (64K), L2 Unified (512K)

This shows this SOC has a mix of both Cortex A720 cores and Cortex A520 cores and even the Cortex A720 cores have different maximum frequencies reflecting possible different between “high” and “mid” cores.

The first thing we do on this new system is run a sweep of stream pinned to different cores. I updated a script to consider the core capability, the optimization level and the number of threads to create the following measurements:

triad_MBps	opt_level	threads	strategy	domain	cpus	run
39170.300000	O2	4	local_l3	0	0,1,10,11	1
38441.200000	O2	2	local_l3	0	0,1	1
37068.000000	O2	2	local_l3_cc	0	0,10	1
36531.400000	O2	3	local_l3	0	0,1,10	1
35612.500000	O2	2	local_l3_cc	0	0,6	1
34266.600000	O2	2	local_l3_cc	0	0,11	1
33733.900000	O2	2	local_l3_cc	0	0,7	1
26004.700000	O2	1	local_l3	0	0	1

The highest stream triad comes from running four threads and pinning to the four highest frequency Cortex 720 cores. Just using the two fastest cores comes close.

The next thing we do is try coremark for each individual core

cpu	run	status	coremark_average	log_file
0	1	OK	25512.373511	/home/mev/source/perf/results/coremark_each_core/coremark.cpu0.run1.txt
1	1	OK	25511.108133	/home/mev/source/perf/results/coremark_each_core/coremark.cpu1.run1.txt
2	1	OK	8182.294456	/home/mev/source/perf/results/coremark_each_core/coremark.cpu2.run1.txt
3	1	OK	8168.957901	/home/mev/source/perf/results/coremark_each_core/coremark.cpu3.run1.txt
4	1	OK	8172.910173	/home/mev/source/perf/results/coremark_each_core/coremark.cpu4.run1.txt
5	1	OK	8176.058957	/home/mev/source/perf/results/coremark_each_core/coremark.cpu5.run1.txt
6	1	OK	22462.563490	/home/mev/source/perf/results/coremark_each_core/coremark.cpu6.run1.txt
7	1	OK	22442.589501	/home/mev/source/perf/results/coremark_each_core/coremark.cpu7.run1.txt
8	1	OK	21470.071023	/home/mev/source/perf/results/coremark_each_core/coremark.cpu8.run1.txt
9	1	OK	21462.437238	/home/mev/source/perf/results/coremark_each_core/coremark.cpu9.run1.txt
10	1	OK	24509.139361	/home/mev/source/perf/results/coremark_each_core/coremark.cpu10.run1.txt
11	1	OK	24504.245930	/home/mev/source/perf/results/coremark_each_core/coremark.cpu11.run1.txt

This shows us results that are consistent with the cores and their maximum frequencies. It surprises me how much quicker a Cortex 720 is than a Cortex 520, so some strategy that pins to the 8 Cortex 720 cores might make sense in some situations.

coremark_average	scenario	domain	threads	cpus	run
172518.777163	thread0_only	all	12	0,1,2,3,4,5,6,7,8,9,10,11	1
168110.086124	all_logical	all	12	0,1,2,3,4,5,6,7,8,9,10,11	1
51064.558269	capability_group	0	2	0,1	1
49082.209232	capability_group	1	2	10,11	1
44976.394529	capability_group	2	2	6,7	1
43020.862299	capability_group	3	2	8,9	1
33342.065373	capability_group	4	4	2,3,4,5	1

I also tried some aggregation of running on all cores, on all thread 0 (hyperthread) and for each pair of cores that have the same capabilities. The first two numbers are close because they are the same run. The pairs of Cortex cores by themselves have coremark scores according to frequencies. The four Cortex 520 cores by themselves are slower than any pair of Cortex 720 cores.

As a follow on post, I will also document what these same scripts have shown with a Ryzen AI 9 370 also show.

Posted in experiment, hardware, Tools | Tagged Aarch64, coremark, stream | Leave a reply

Close to 100 workloads, adding thresholds

Performance analysis, tools and experiments Posted on January 26, 2024 by mevJanuary 26, 2024

I am now close to 100 overall phoronix tests added. Recent articles still include a number of new benchmarks, typically I have ~2/3 of the ones in an article and then need to add the remaining ones. However, over time have to get closer to having all the ones as articles come out.

As I have this number of workloads, I can now start to set more precise thresholds on what it means to be “high” or “low” on a metric. These could be slightly different between my AMD and Intel CPU –

  • IPC reported by Intel is slightly higher
  • Retirement rate reported by Intel is slightly higher
  • Frontend and backend stalls reported by Intel are slightly lower
  • Speculation misses reported by Intel are higher

Some of this might be because of differences in how the metrics are defined/counted and some due to the processors themselves. However, for now I’ve hard coded some thresholds into the tool that are same for both since these are mostly guidance and over time if my workload mix shifts or I find different values on other processors, I might adjust. Following are the initial guidelines added:

MetricHighLow
IPC3.00.7
retiring54%14%
frontend45%5%
backend70%18%
retiring10%1%
Posted in experiment, Tools | Tagged ipc, threshold, topdown | Leave a reply

topdown – adding process trees and statistics

Performance analysis, tools and experiments Posted on January 1, 2024 by mevJanuary 1, 2024

I have not enhanced the topdown tool with ability to print process trees. This enables the key features of my previous “wspy” command.

The interfaces is as follows. I added the following options to topdown to record process information:

	--tree <file>             - create CSV of processes
	--tree-cmdline            - record full command lines

The –tree option uses strace(2) to record fork/exec/exit events and save information to the file for later processing. An example of some information saved is as follows:

0.000 14119 root
0.002 14119 fork 14120
0.017 14120 comm cc1
0.017 14120 cmdline /usr/lib/gcc/x86_64-linux-gnu/11/cc1 -quiet -imultiarch x86_64-linux-gnu hello.c -quiet -dumpbase hello.c -dumpbase-ext .c -mtune=generic -march=x86-64 -fasynchronous-unwind-tables -fstack-protector-strong -Wformat -Wformat-security -fstack-clash-protection -fcf-protection -o /tmp/ccIyphnx.s
0.017 14120 exit 14120 (cc1) t 14119 14118 13158 34819 14118 1077936128 1221 0 0 0 1 0 0 0 20 0 1 0 1313576 46571520 3835 18446744073709551615 5890048 21573621 140722169364960 0 0 0 0 0 1256 1 0 0 17 18 0 0 0 0 0 30095936 30148584 59514880 140722169373230 140722169373523 140722169373523 140722169376723 0
0.018 14119 fork 14121
0.021 14121 comm as
0.021 14121 cmdline as --64 -o /tmp/cceku4R5.o /tmp/ccIyphnx.s
0.021 14121 exit 14121 (as) t 14119 14118 13158 34819 14118 1077936128 441 0 0 0 0 0 0 0 20 0 1 0 1313578 12435456 1332 18446744073709551615 94138168020992 94138168333961 140723056042896 0 0 0 0 0 1256 1 0 0 17 11 0 0 0 0 0 94138168430896 94138168453272 94138176069632 140723056051009 140723056051052 140723056051052 140723056054252 0
0.022 14119 fork 14124
0.023 14124 fork 14125
0.040 14125 comm ld
0.040 14125 cmdline /usr/bin/ld -plugin /usr/lib/gcc/x86_64-linux-gnu/11/liblto_plugin.so -plugin-opt=/usr/lib/gcc/x86_64-linux-gnu/11/lto-wrapper -plugin-opt=-fresolution=/tmp/ccWwmx4O.res -plugin-opt=-pass-through=-lgcc -plugin-opt=-pass-through=-lgcc_s -plugin-opt=-pass-through=-lc -plugin-opt=-pass-through=-lgcc -plugin-opt=-pass-through=-lgcc_s --build-id --eh-frame-hdr -m elf_x86_64 --hash-style=gnu --as-needed -dynamic-linker /lib64/ld-linux-x86-64.so.2 -pie -z now -z relro -o hello /usr/lib/gcc/x86_64-linux-gnu/11/../../../x86_64-linux-gnu/Scrt1.o /usr/lib/gcc/x86_64-linux-gnu/11/../../../x86_64-linux-gnu/crti.o /usr/lib/gcc/x86_64-linux-gnu/11/crtbeginS.o -L/usr/lib/gcc/x86_64-linux-gnu/11 -L/usr/lib/gcc/x86_64-linux-gnu/11/../../../x86_64-linux-gnu -L/usr/lib/gcc/x86_64-linux-gnu/11/../../../../lib -L/lib/x86_64-linux-gnu -L/lib/../lib -L/usr/lib/x86_64-linux-gnu -L/usr/lib/../lib -L/usr/lib/gcc/x86_64-linux-gnu/11/../../.. /tmp/cceku4R5.o -lgcc --push-state --as-needed -lgcc_s --pop-state -lc -lgcc --push-state --as-needed -lgcc_s --pop-state /usr/lib/gcc/x86_64-linux-gnu/11/crtendS.o /usr/lib/gcc/x86_64-linux-gnu/11/../../../x86_64-linux-gnu/crtn.o
0.040 14125 exit 14125 (ld) t 14124 14118 13158 34819 14118 1077936128 1732 0 0 0 1 0 0 0 20 0 1 0 1313578 16846848 2276 18446744073709551615 94366230806528 94366231102693 140724178835088 0 0 0 0 0 0 1 0 0 17 20 0 0 0 0 0 94366232461104 94366232495352 94366259900416 140724178836727 140724178837886 140724178837886 140724178841580 0
0.040 14124 comm collect2
0.040 14124 cmdline /usr/lib/gcc/x86_64-linux-gnu/11/collect2 -plugin /usr/lib/gcc/x86_64-linux-gnu/11/liblto_plugin.so -plugin-opt=/usr/lib/gcc/x86_64-linux-gnu/11/lto-wrapper -plugin-opt=-fresolution=/tmp/ccWwmx4O.res -plugin-opt=-pass-through=-lgcc -plugin-opt=-pass-through=-lgcc_s -plugin-opt=-pass-through=-lc -plugin-opt=-pass-through=-lgcc -plugin-opt=-pass-through=-lgcc_s --build-id --eh-frame-hdr -m elf_x86_64 --hash-style=gnu --as-needed -dynamic-linker /lib64/ld-linux-x86-64.so.2 -pie -z now -z relro -o hello /usr/lib/gcc/x86_64-linux-gnu/11/../../../x86_64-linux-gnu/Scrt1.o /usr/lib/gcc/x86_64-linux-gnu/11/../../../x86_64-linux-gnu/crti.o /usr/lib/gcc/x86_64-linux-gnu/11/crtbeginS.o -L/usr/lib/gcc/x86_64-linux-gnu/11 -L/usr/lib/gcc/x86_64-linux-gnu/11/../../../x86_64-linux-gnu -L/usr/lib/gcc/x86_64-linux-gnu/11/../../../../lib -L/lib/x86_64-linux-gnu -L/lib/../lib -L/usr/lib/x86_64-linux-gnu -L/usr/lib/../lib -L/usr/lib/gcc/x86_64-linux-gnu/11/../../.. /tmp/cceku4R5.o -lgcc --push-state --as-needed -lgcc_s --pop-state -lc -lgcc --push-state --as-needed -lgcc_s --pop-state /usr/lib/gcc/x86_64-linux-gnu/11/crtendS.o /usr/lib/gcc/x86_64-linux-gnu/11/../../../x86_64-linux-gnu/crtn.o
0.040 14124 exit 14124 (collect2) t 14119 14118 13158 34819 14118 1077936128 85 1732 0 0 0 0 1 0 20 0 1 0 1313578 8839168 250 18446744073709551615 4202496 4414097 140733621318048 0 0 0 0 0 9287 1 0 0 17 19 0 0 0 0 0 4488704 4494640 30085120 140733621320891 140733621322080 140733621322080 140733621325774 0
0.041 14119 comm gcc
0.041 14119 cmdline gcc -o hello hello.c
0.041 14119 exit 14119 (gcc) t 14118 14118 13158 34819 14118 1077936128 135 3479 0 0 0 0 2 1 20 0 1 0 1313576 9617408 251 18446744073709551615 4206592 4563953 140732373361024 0 0 0 0 0 20483 1 0 0 17 14 0 0 0 0 0 5114624 5124176 36139008 140732373369904 140732373369925 140732373369925 140732373372907 0

The “exit” event captures the contents of /proc/<pid>/stat when the process exits. I am not sure if this is reliable for hundreds of thousands of processes but for smaller several hundred examples it works find. If the –tree-cmdline option is given then we also capture /proc/<pid>/cmdline when the process exits.

This data file can then be processed with the proctree program with the following options

./source/wspy/proctree: fatal error: usage: ./source/wspy/proctree -[CcFfSsTtuv][-w width] file
	-C	turn on longer command line
	-c	turn on abbreviated command (default)
	-F	urn on start/finish info (default)
	-f	turn off start/finish info
	-S	turn on summary output
	-s	turn off summary output (default)
	-T	turn on tree output (default)
	-t	turn off tree output
	-U	turn off utime in tree
	-u	turn on utime in tree
	-v	verbose messages
	-w width	set command width

Here is a basic output with both summary statistics and tree information

5 processes
	  1 cc1                      0.01     0.00
	  1 ld                       0.01     0.00
	  1 as                       0.00     0.00
	  1 collect2                 0.00     0.00
	  1 gcc                      0.00     0.00
0 processes running
3 maximum processes

14119) gcc start=0.00  finish=0.04 
  14120) cc1 start=0.00  finish=0.02 
  14121) as start=0.02  finish=0.02 
  14124) collect2 start=0.02  finish=0.04 
    14125) ld start=0.02  finish=0.04 

We can see more of the command line by adding the -C switch and also increasing the -w width

5 processes
	  1 cc1                      0.01     0.00
	  1 ld                       0.01     0.00
	  1 as                       0.00     0.00
	  1 collect2                 0.00     0.00
	  1 gcc                      0.00     0.00
0 processes running
3 maximum processes

14119) gcc -o hello hello.c start=0.00  finish=0.04 
  14120) /usr/lib/gcc/x86_64-linux-gnu/11/cc1 -quiet -imultiarch x86_64-linux-gnu hello.c -quiet -dumpbase hello.c -dum start=0.00  finish=0.02 
  14121) as --64 -o /tmp/cceku4R5.o /tmp/ccIyphnx.s start=0.02  finish=0.02 
  14124) /usr/lib/gcc/x86_64-linux-gnu/11/collect2 -plugin /usr/lib/gcc/x86_64-linux-gnu/11/liblto_plugin.so -plugin-op start=0.02  finish=0.04 
    14125) /usr/bin/ld -plugin /usr/lib/gcc/x86_64-linux-gnu/11/liblto_plugin.so -plugin-opt=/usr/lib/gcc/x86_64-linux- start=0.02  finish=0.04 

Overall, this is a useful tool that helps me get more of the process overview e.g. single-threaded vs multi-threaded as well as summarizing processes that take the most time. As needed I also have a mechanism to decorate with additional instrumentation. Two examples might be (a) checking for particular syscalls e.g. file open events (b) investigating more of a process drill down not to the initial parent but to multiple sub-runs.

However, for now I have a basic topdown tool with both periodic output and a process tree to examine different workloads.

Posted in Tools | Tagged topdown, tree | Leave a reply

Creating basic metrics and adding topdown plots

Performance analysis, tools and experiments Posted on December 31, 2023 by mevDecember 31, 2023

I have made several enhancements to the topdown tool. I also have some fragile things I still need to sort out along the way. The net combination is best seen below where I include both a topdown metrics summary (created … Continue reading →

Posted in Tools | Tagged gnuplot, performance counters, topdown | Leave a reply

Potential interface and potential counter groups for topdown tool

Performance analysis, tools and experiments Posted on December 25, 2023 by mevDecember 25, 2023

I have looked through the Family 19h PPR reference, output from “perf list -v –detail” and also some likwid counter groups to figure out combinations of counters I might be able to add as instrumentation options for a topdown command. … Continue reading →

Posted in Tools | Tagged performance counters, topdown | Leave a reply

topdown – updated tool and metrics

Performance analysis, tools and experiments Posted on December 23, 2023 by mevDecember 23, 2023

I have updated and enhanced the topdown tool and also used this as an occasion to explore Zen4 topdown performance counters, Intel hybrid CPU while building something to compare Intel i5-13500H and Ryzen 7940 processor metrics. The interface might change, … Continue reading →

Posted in Tools | Tagged getrusage, perf_event_open, performance counters | Leave a reply

perf – new performance counters with Linux 6.2

Performance analysis, tools and experiments Posted on February 22, 2023 by mevFebruary 22, 2023

It looks like there are many new capabilities in the linux perf command run on a Zen4 core under Linux 6.2 when compared with Zen1 core under Linux 5.4. I compared the “perf list” output between: Ubuntu 20.04, Linux 5.4, … Continue reading →

Posted in Tools | Tagged perf | Leave a reply

Meta

  • Log in
  • Entries feed
  • Comments feed
  • WordPress.org

Archives

  • July 2026
  • November 2024
  • October 2024
  • September 2024
  • July 2024
  • June 2024
  • March 2024
  • February 2024
  • January 2024
  • December 2023
  • February 2023

Tags

7840HS Aarch64 bad data benchmarks cachyos cluster compiler coremark cpu2017 data fabric getrusage gnuplot i5-13500H icache ipc kernel l3 metrics namd opcache perf performance counters perf_event_open phoronix Ryzen AI 9 HX 370 Ryzen AI 365 scaling stream threshold topdown tree virtualization website wsl Zen5

Recent Posts

  • Performance Measurements on four systems
  • New scripts with Ryzen AI 9 HX 370
  • ARM box, inventory script
  • Virtualization comparisons
  • Updating to a new kernel and graphics driver
©2026 - Performance analysis, tools and experiments - Weaver Xtreme Theme
↑