Repository navigation
Mess Benchmark
Mess (Memory Stress) is a comprehensive benchmarking framework that characterizes memory system performance through bandwidth-latency curves covering the full range of memory traffic intensity, from unloaded to fully saturated.
Traditional memory benchmarks report isolated metrics:
- Peak bandwidth (e.g., STREAM)
- Idle latency (e.g., pointer-chase benchmarks)
These metrics fail to capture how memory systems behave under realistic workloads, where applications operate between these extremes.
Mess provides the complete picture: a family of bandwidth-latency curves showing how latency degrades as bandwidth increases, across different read/write ratios.
Example of a single bandwidth-latency curve
Mess operates through four systematic phases:
Multiple CPU cores execute tight assembly loops that generate memory pressure:
- Configurable load/store ratios: From 100% reads to 100% writes
- Adjustable intensity: Pause cycles between operations control bandwidth saturation
- SIMD instructions: Maximize memory throughput per core
┌─────────────────────────────────────────────────────────────┐
│ Core 0 Core 1 Core 2 ... Core N-1 │
│ ↓ ↓ ↓ ↓ │
│ [Traffic] [Traffic] [Traffic] [Traffic] │
│ ↓ ↓ ↓ ↓ │
│ ═══════════════════════════════════════════════════════ |
│ Memory Controller │
│ ═══════════════════════════════════════════════════════ │
│ ↕ │
│ DRAM │
└─────────────────────────────────────────────────────────────┘
While traffic generators saturate bandwidth, a dedicated core runs pointer-chase operations:
- Follows a pointer-chase linked list through memory
- Measures access time under load
- Provides application-representative latency values
Hardware performance counters track actual memory controller traffic:
- Read bytes: Data read from DRAM
- Write bytes: Data written to DRAM
- Uses
perf,likwid, or other instrumenting tools depending on system
This captures architecture-level bandwidth including cache writeback traffic and prefetcher activity.
By varying two parameters, Mess constructs a family of bandwidth-latency curves:
| Parameter | Effect |
|---|---|
| Pause cycles | Controls bandwidth saturation level |
| Issued Load/store ratio | Changes read vs write mix |
Each bandwidth-latency curve shows how latency degrades as bandwidth increases for a specific read/write ratio.
Mess kernels are written at the assembly level to:
- Minimize compiler and OS overhead
- Provide precise control over memory access patterns
- Ensure reproducible, hardware-representative results
Native support for major CPU architectures:
| Architecture | SIMD | Notes |
|---|---|---|
| x86-64 | AVX2, AVX-512 | Intel and AMD |
| ARM | NEON, SVE | Neoverse, Graviton |
| Power | VSX | Power8+ |
| RISC-V | RVV 1.0 | Vector extension |
See Architecture Support for details.
By default, Mess uses temporal stores for write operations.
See Temporal vs Non-Temporal stores for technical details.
Research using Mess has revealed critical insights about memory systems:
Memory writes significantly degrade performance compared to reads. On most systems, write bandwidth is 30-50% lower than read bandwidth due to:
- Read-for-ownership protocol overhead
- Memory controller scheduling priorities
- DRAM timing constraints
Systems typically saturate at 70-90% of theoretical maximum bandwidth. Beyond this point:
- Latency increases sharply
- Additional cores provide diminishing returns
- Applications become memory-bound
Typical latency values across the bandwidth-latency curve:
- Idle (unloaded): 85-130ns
- Moderate load: 150-300ns
- Near saturation: 200-600ns+
After Installation, run the benchmark:
# Basic run (measurement files are saved by default)
./build/bin/mess
Results are saved to the measuring/ directory and can be visualized using Plotter.
When --pause is omitted, Mess uses Adaptive Curve-Guided Pause Discovery to choose pause values automatically. The default standard tier measures up to 50 points per enabled execution mode, focusing measurements around curve transitions and knees.
For detailed command reference, see:
- Understanding CLI arguments - Experimenting with configurations
- Adaptive Curve-Guided Pause Discovery - Automatic pause selection
- Iterative Debugging - Step-by-step validation
Mess generates bandwidth and latency measurements for each configuration. Use the Plotter utilities to:
- Generate bandwidth-latency curve plots: Visualize the bandwidth-latency relationship
- Compare systems: Overlay bandwidth-latency curves from different machines
- Profile applications: Map your workload onto the bandwidth-latency curves using Mess Profiler
- Mess Website - Project website with tutorials, publications, and updates
- Mess Paper - Esmaili-Dokht, P., Sgherzi, F., Girelli, V. S., Boixaderas, I., Carmin, M., Monemi, A., Armejach, A., Mercadal, E., Llort, G., Radojković, P., Moreto, M., Giménez, J., Martorell, X., Ayguadé, E., Labarta, J., Confalonieri, E., Dubey, R., & Adlard, J. (2024). A mess of memory system benchmarking, simulation and application profiling. In Proceedings of the 57th IEEE/ACM International Symposium on Microarchitecture (MICRO) (pp. 136-152). IEEE.
- Mess Profiler - Profile application memory bandwidth
- Plotter - Visualize benchmark results
- Traffic Generator - How traffic generation works
- Understand output - Output file formats