A multiplatform benchmark designed to provide a holistic, detailed, and close-to-hardware view of memory system performance through bandwidth-latency curves.
This is an update to the original Mess Benchmark (now deprecated), focused on improved usability and portability.
- Documentation
- Motivation
- Tools Included
- Architecture Support
- Installation
- Quick Start
- Graphic User Interface
- Common Workflows
- Mess Profiler
- Plotter-Parser
- Learning Resources
- Troubleshooting
- Contributors
- Citation
- References
Project documentation is available in the GitHub wiki:
- Wiki Home
- Installation
- Understanding CLI Arguments
- Mess Benchmark
- Mess Profiler
- Plotter-Parser
- Architecture Support
- FAQ
Traditional memory benchmarks report isolated metrics such as peak bandwidth or idle latency, which often fail to capture how memory systems behave under realistic workloads. Mess (Memory Stress) addresses this limitation by characterizing memory performance through bandwidth-latency curves that cover the full range of memory traffic intensity, from unloaded to fully saturated.
This approach reveals critical insights:
- Memory writes degrade performance significantly compared to reads
- Systems typically saturate at 70-90% of theoretical maximum bandwidth
- Latency ranges from 85-130ns when idle to 200-600ns+ under saturation
These ranges are illustrative; measured values depend on the system and memory configuration.
Mess provides a holistic, close-to-hardware view of memory system behavior, enabling researchers and engineers to understand real-world performance characteristics that standard benchmarks miss.
MICRO 2024 Best Paper Runner-Up: The Mess methodology was published at the 57th IEEE/ACM International Symposium on Microarchitecture.
For a detailed explanation of the benchmark methodology, see the Memory BSC Tools page.
Mess provides an integrated workflow for memory system characterization, from benchmarking to application profiling:
- Mess Benchmark: Characterizes your memory system by generating bandwidth-latency curves that reveal how it behaves under varying load.
- Mess Profiler: Automates counter discovery and runs profiling tools (
perf,likwid, etc.) with the correct configuration, ensuring application measurements align with benchmark data. - Plotter-Parser: Generates publication-quality plots as well as CSV and JSON files containing the parsed bandwidth-latency curves.
- Mess GUI: Desktop interface for configuring benchmark runs and inspecting results on Linux, macOS and Windows.
- Traffic Generator: The low-level engine that generates precise memory traffic patterns at the assembly level. Can also be used independently for custom microbenchmarks.
Mess can characterize memory systems in consumer laptops and desktops, workstations, and servers. Availability depends on both the operating system and access to bandwidth counters for the processor. ISA support alone does not guarantee that a complete bandwidth-latency curve can be measured.
| System | Availability | Bandwidth backend / requirements |
|---|---|---|
| Linux PCs, workstations and servers | Available on supported CPUs | perf, LIKWID, Intel PCM or VTune; requires accessible counters for the CPU and memory system |
| Apple Silicon Macs (M-series) | Available | IOReport reads Apple memory-controller counters; Xcode Command Line Tools required |
| Intel Macs running macOS | Unavailable for bandwidth-latency curves | No macOS bandwidth-counter backend for Intel CPUs |
| Windows PCs and workstations (Intel/AMD x86-64) | Available with vendor tools | Intel VTune or the Intel client IMC backend, or AMD uProf; Administrator access required |
| Linux RISC-V systems | Partial / WIP | Kernel generation and latency support are implemented; bandwidth measurement remains pending |
The Mess GUI provides a desktop interface for Linux, macOS (Apple Silicon) and Windows. It uses the same benchmark engine and has the same measurement requirements as the CLI.
Platform-specific limits:
- Linux: Instruction-latency sampling requires a supported PEBS/SPE CPU and
access to its sampling events. NUMA memory binding uses
--bind. - macOS:
--bind,--inst-latand--add-countersare unavailable. - Windows:
--bind,--inst-latand--add-countersare unavailable. Backends are implemented, but native Windows driver and curve validation remains pending.
Available modes depend on the hardware. For implementation details, see Architecture Support.
| Architecture | ISA / API | Notes |
|---|---|---|
| x86-64 CPUs | SCALAR, SSE2, AVX, AVX2, AVX-512 | Intel and AMD; vector modes require the corresponding CPU features |
| ARM CPUs | SCALAR, NEON, SVE | Apple Silicon, Neoverse, Graviton, A64FX and other supported ARM systems; SVE modes require SVE-capable hardware |
| Power-PC CPUs | VSX, VMX | Power8 and newer on Linux |
| RISC-V CPUs | SCALAR, RVV | Partial / WIP; bandwidth measurement is not yet available |
| CUDA-capable NVIDIA GPUs | CUDA | Run with --gpu; requires the NVIDIA driver and CUDA Toolkit 12.6+ (nvcc, CUPTI) |
| Other GPU vendors | — | Backend support is not yet available |
See full instructions in the Installation wiki page.
For the desktop interface and installer information, see Mess GUI. The instructions below cover building and running the CLI.
git clone --recursive https://github.com/bsc-mem/Mess.git
cd MessImportant: The --recursive flag is required to download submodules that Mess depends on.
If you forget --recursive, initialize submodules later:
git submodule update --init --recursiveCore requirements:
- C++17 compiler: GCC 9+ (recommended), Clang 10+, Intel OneAPI (ICX), AOCC
- numactl: Required for NUMA memory binding
- taskset: Core pinning (preferred, part of
util-linux) - perf: Recommended counter backend (
linux-tools-common) - CMake 3.21+: Required for CMake builds; the Linux
makebuild does not need it - Python 3: Plotting utilities
To run the NVIDIA GPU benchmark with --gpu, install the NVIDIA driver, CUDA
Toolkit 12.6 or newer (including nvcc and CUPTI). Mess compiles the GPU
kernels at runtime and reads DRAM counters in-process through CUPTI.
Ensure perf access:
cat /proc/sys/kernel/perf_event_paranoid
echo 0 | sudo tee /proc/sys/kernel/perf_event_paranoidHuge pages are recommended for more accurate measurements:
echo 1024 | sudo tee /proc/sys/vm/nr_hugepagesEven without huge pages, Mess automatically compensates for page walk latency. See Huge-Memory-Pages for details.
Core requirements:
- Apple Silicon: M-series Mac required; bandwidth is read from the Apple memory controller through IOReport, and Intel Macs have no counter backend
- Xcode Command Line Tools: Apple Clang (C++17) plus the CoreFoundation and IOKit frameworks
- CMake 3.21+: Only for the
macos-cmakepreset, themakebuild does not need it - Python 3: Plotting utilities
xcode-select --install
brew install cmake python3No counter setup is needed: Mess reads /usr/lib/libIOReport.dylib directly, so
there is no driver to install and no elevated prompt to open.
Core requirements:
- MSYS2: UCRT64 toolchain (msys2.org)
- CMake 3.21+: Required by CMake and the Make compatibility commands
- Python 3: Plotting utilities
- Ninja: Generator used by that preset
- Intel VTune, AMD uProf, or the Intel PCM driver: Counter backend (see below)
Install the toolchain from an MSYS2 shell:
pacman -S mingw-w64-ucrt-x86_64-gcc mingw-w64-ucrt-x86_64-cmake ninjaWindows exposes no memory-bandwidth counter to normal programs, so Mess reads one through a vendor tool and its kernel driver. Install the tool for your CPU, then always run Mess from an Administrator prompt.
| Your CPU | Tool to install |
|---|---|
| AMD | AMD uProf, with bandwidth collection support for your CPU |
| Intel | Intel VTune Profiler, where memory-bandwidth counters are supported |
| Intel client CPUs (IMC fallback) | Intel PCM msr.sys driver with physical-memory mapping support; Mess uses --measurer=imc |
Mess automatically selects the available backend with --measurer=auto.
The Intel client IMC backend does not support Xeon/server memory controllers.
The --measurer=pcm CLI backend is unavailable on Windows.
make
make installBinaries are generated in build/bin/:
mess— Core benchmarkmess-profiler— Memory bandwidth profilergenerate_code— Kernel generator;make installruns it to emit the traffic generator
Generated traffic-generator binaries are placed alongside mess in build/bin/:
traffic_gen_multiseq.x— Standalone traffic generation tool
Optional PATH setup:
export PATH=$PATH:$(pwd)/build/bin$env:PATH="C:\msys64\ucrt64\bin;$env:PATH"
cmake --preset windows-mingw
cmake --build --preset windows-mingw
build\cmake\windows-mingw\bin\Release\generate_code.exeWith this preset, binaries are generated in build/cmake/windows-mingw/bin/Release/:
mess.exe— Core benchmarkmess-profiler.exe— Memory bandwidth profilergenerate_code.exe— Kernel generator; run it as shown above to emit the traffic generator
Keep C:\msys64\ucrt64\bin first on PATH. --bind, --inst-lat
and --add-counters are unavailable on Windows.
Linux and macOS (make build):
./build/bin/mess --version
./build/bin/mess --dry-run --verbose=2Windows (PowerShell, CMake preset):
.\build\cmake\windows-mingw\bin\Release\mess.exe --version
.\build\cmake\windows-mingw\bin\Release\mess.exe --dry-run --verbose=2The examples below use the Linux/macOS make build. On Windows, use
.\build\cmake\windows-mingw\bin\Release\mess.exe with the same supported
options. See System Availability for platform limits.
Profiling is enabled by default: CPU and GPU runs save measurement files under
measuring/ (or --folder=DIR). Use --no-profile to discard them after the run.
./build/bin/mess --dry-run --verbose=2
./build/bin/mess
./build/bin/mess --no-profile
./build/bin/mess --tier=liteCommon options:
| Option | Description | Example |
|---|---|---|
--gpu |
Run the NVIDIA GPU bandwidth-latency benchmark (requires CUDA Toolkit 12.6+) | --gpu |
--ratio=N[,N...] |
Issued load ratio(s) in % | --ratio=100,75,50 |
--pause=N[,N...] |
Pause bubble values | --pause=0,10,100,1000 |
--tier=TIER |
Adaptive pause discovery preset (lite/standard/detailed) |
--tier=lite |
--point-count=N |
Custom adaptive point budget | --point-count=75 |
--no-profile |
Discard measurement files after the run (saved by default) | --no-profile |
--verbose=N |
Verbosity level 0-4 |
--verbose=3 |
--measurer=TYPE |
Counter backend (auto/perf/likwid/pcm/vtune/ioreport/imc/uprof) |
--measurer=perf |
--inst-lat |
Linux: use supported PEBS/SPE instruction sampling for latency | --inst-lat |
--add-counters=LIST |
Measure extra counters alongside bandwidth | --add-counters=cycles,instructions |
--bind=LIST |
Linux: NUMA memory-node binding | --bind=0 |
--cores=LIST |
Explicit traffic-generator cores | --cores=0-15 |
--total-cores=N |
Number of traffic-generator cores | --total-cores=16 |
--repetitions=N |
Number of point repetitions (default: 3) | --repetitions=5 |
Some options are platform-specific. For the complete set your build supports,
use ./build/bin/mess --help or see Understanding-CLI-Arguments.
The Mess GUI is a desktop front-end for the same benchmark binary the CLI drives: set up quick runs, launch full bandwidth-latency curves, and inspect the results.
Installers for Linux, macOS (Apple Silicon) and Windows will be available on the Mess GUI download page. Choose the package for your operating system and follow the installation instructions provided there.
The GUI needs the build tools and counter backend for your platform to compile Mess and collect measurements.
The GUI asks where Mess lives:
| Choice | When to pick it | What it does |
|---|---|---|
| Locate existing Mess | You already cloned this repository | Points the GUI at your existing checkout |
| Install Mess | Nothing cloned yet | Clones this repository for you |
Either way, the GUI then compiles Mess and uses the resulting binary as the backend for every measurement.
./build/bin/mess --ratio=100 --pause=0 --verbose=3 --repetitions=1
./build/bin/mess --ratio=0 --pause=0 --verbose=3 --repetitions=1./build/bin/mess --tier=standard
./build/bin/mess --point-count=75Linux only, on CPUs with supported PEBS/SPE sampling events:
./build/bin/mess --inst-lat --tier=liteLinux examples; available backends vary by platform:
./build/bin/mess --measurer=likwid --bind=0
./build/bin/mess --measurer=pcm --bind=0
./build/bin/mess --measurer=vtune --tier=lite
./build/bin/mess --measurer=perf --add-counters=cycles,instructionsLinux only, on systems with multiple memory nodes:
./build/bin/mess --bind=0 --folder=numa0
./build/bin/mess --bind=1 --folder=numa1for c in 2 4 8 16; do
./build/bin/mess --total-cores=$c --folder=cores_$c
donemess-profiler reuses Mess counter discovery to profile applications with consistent output.
./build/bin/mess-profiler --dry-run
./build/bin/mess-profiler -s 100ms -o app_profile.csv ./my_appProfiler docs: Mess-Profiler
Visualization utilities in utils/:
plotter.py: generates memory-curve plots and processed CSV/JSONapp_plotter.py: overlays application profile points on curves
cd utils
python3 -m venv venv
source venv/bin/activate
pip install -r requirements.txtPlotter docs: Plotter-Parser
- Wiki: Mess Wiki
- Tutorials and Slides: mess.bsc.es/tutorials
- Detailed Methodology: memory.bsc.es/tools/mess-benchmark
- Results Dataset: Mess Results (GitHub) — bandwidth-latency curves for a variety of system architectures
- Permission errors on counters: set
perf_event_paranoidto0 - Missing counters/backend mismatch: check
./build/bin/mess-profiler --dry-run - Unstable measurements: increase
--repetitionsand use--verbose=3 - Windows: no bandwidth counters — run as Administrator and install the vendor tool for your CPU (see Windows)
- macOS: bandwidth measurement requires Apple Silicon;
--bind,--inst-latand--add-countersare not supported (see macOS)
More: FAQ and Iterative-Debugging
Found a bug? Open an issue on GitHub or email mess@bsc.es.
Mess is developed by the Memory Systems Team at the Barcelona Supercomputing Center (BSC).
|
Victor Xirau Guardans Main Mess developer victor.xirau@bsc.es |
Mariana Carmin Mess developer mcarmin@bsc.es |
|
Pau Díaz Mess developer pau.diazcuesta@bsc.es |
Javier Beiro Mess GPU developer javier.beiro@bsc.es |
Mess Paper author: Pouya Esmaili Dokht (pouya.esmaili@bsc.es)
Or email: mess@bsc.es
If you use Mess in research, please cite:
@inproceedings{esmaili2024mess,
title = {A Mess of Memory System Benchmarking, Simulation and Application Profiling},
author = {Esmaili-Dokht, Pouya and Sgherzi, Francesco and Girelli, Valeria Soldera
and Boixaderas, Isaac and Carmin, Mariana and Monemi, Alireza
and Armejach, Adria and Mercadal, Estanislao and Llort, German
and Radojkovi{\'c}, Petar and Moreto, Miquel and Gim{\'e}nez, Judit
and Martorell, Xavier and Ayguad{\'e}, Eduard and Labarta, Jesus
and Confalonieri, Emanuele and Dubey, Rishabh and Adlard, Joshua},
booktitle = {Proceedings of the 57th IEEE/ACM International Symposium on Microarchitecture (MICRO)},
pages = {136--152},
year = {2024},
publisher = {IEEE}
}- Mess Benchmark — The original implementation of the Mess benchmark.
- Mess Simulator — Analytical memory model using bandwidth-latency curves.
- Mess-Paraver — Integration with Paraver for visualization.
- Mess Results — Collection of bandwidth-latency curves for various system architectures; a dataset users can browse and compare.
- Mess Paper — Esmaili-Dokht, P., Sgherzi, F., Girelli, V. S., Boixaderas, I., Carmin, M., Monemi, A., Armejach, A., Mercadal, E., Llort, G., Radojković, P., Moreto, M., Giménez, J., Martorell, X., Ayguadé, E., Labarta, J., Confalonieri, E., Dubey, R., & Adlard, J. (2024). A mess of memory system benchmarking, simulation and application profiling. In Proceedings of the 57th IEEE/ACM International Symposium on Microarchitecture (MICRO) (pp. 136-152). IEEE.
Mess Benchmark is released under the BSD 3-Clause License