Skip to content

Mess Benchmark

Victor Xirau Guardans edited this page Sep 29, 2026 · 3 revisions

Mess Benchmark

Mess (Memory Stress) is a comprehensive benchmarking framework that characterizes memory system performance through bandwidth-latency curves covering the full range of memory traffic intensity, from unloaded to fully saturated.


Why Bandwidth-Latency Curves?

Traditional memory benchmarks report isolated metrics:

  • Peak bandwidth (e.g., STREAM)
  • Idle latency (e.g., pointer-chase benchmarks)

These metrics fail to capture how memory systems behave under realistic workloads, where applications operate between these extremes.

Mess provides the complete picture: a family of bandwidth-latency curves showing how latency degrades as bandwidth increases, across different read/write ratios.

Bandwidth-Latency Curve Example
Example of a single bandwidth-latency curve


Measurement Methodology

Mess operates through four systematic phases:

1. Traffic Generation

Multiple CPU cores execute tight assembly loops that generate memory pressure:

  • Configurable load/store ratios: From 100% reads to 100% writes
  • Adjustable intensity: Pause cycles between operations control bandwidth saturation
  • SIMD instructions: Maximize memory throughput per core
┌─────────────────────────────────────────────────────────────┐
│      Core 0    Core 1    Core 2    ...    Core N-1          │
│        ↓         ↓         ↓                 ↓              │
│      [Traffic] [Traffic] [Traffic]       [Traffic]          │
│        ↓         ↓         ↓                 ↓              │
│  ═══════════════════════════════════════════════════════    |
│                   Memory Controller                         │
│  ═══════════════════════════════════════════════════════    │
│                         ↕                                   │
│                       DRAM                                  │
└─────────────────────────────────────────────────────────────┘

2. Latency Measurement

While traffic generators saturate bandwidth, a dedicated core runs pointer-chase operations:

  • Follows a pointer-chase linked list through memory
  • Measures access time under load
  • Provides application-representative latency values

3. Bandwidth Monitoring

Hardware performance counters track actual memory controller traffic:

  • Read bytes: Data read from DRAM
  • Write bytes: Data written to DRAM
  • Uses perf, likwid, or other instrumenting tools depending on system

This captures architecture-level bandwidth including cache writeback traffic and prefetcher activity.

4. Bandwidth-Latency Curve Construction

By varying two parameters, Mess constructs a family of bandwidth-latency curves:

Parameter Effect
Pause cycles Controls bandwidth saturation level
Issued Load/store ratio Changes read vs write mix

Each bandwidth-latency curve shows how latency degrades as bandwidth increases for a specific read/write ratio.


Key Technical Features

Assembly-Level Implementation

Mess kernels are written at the assembly level to:

  • Minimize compiler and OS overhead
  • Provide precise control over memory access patterns
  • Ensure reproducible, hardware-representative results

Multi-Architecture Support

Native support for major CPU architectures:

Architecture SIMD Notes
x86-64 AVX2, AVX-512 Intel and AMD
ARM NEON, SVE Neoverse, Graviton
Power VSX Power8+
RISC-V RVV 1.0 Vector extension

See Architecture Support for details.

Non-Temporal Stores

By default, Mess uses temporal stores for write operations.

See Temporal vs Non-Temporal stores for technical details.


What the Bandwidth-Latency Curves Reveal

Research using Mess has revealed critical insights about memory systems:

Write Penalty

Memory writes significantly degrade performance compared to reads. On most systems, write bandwidth is 30-50% lower than read bandwidth due to:

  • Read-for-ownership protocol overhead
  • Memory controller scheduling priorities
  • DRAM timing constraints

Saturation Point

Systems typically saturate at 70-90% of theoretical maximum bandwidth. Beyond this point:

  • Latency increases sharply
  • Additional cores provide diminishing returns
  • Applications become memory-bound

Latency Range

Typical latency values across the bandwidth-latency curve:

  • Idle (unloaded): 85-130ns
  • Moderate load: 150-300ns
  • Near saturation: 200-600ns+

Running Mess

After Installation, run the benchmark:

# Basic run (measurement files are saved by default)
./build/bin/mess

Results are saved to the measuring/ directory and can be visualized using Plotter.

When --pause is omitted, Mess uses Adaptive Curve-Guided Pause Discovery to choose pause values automatically. The default standard tier measures up to 50 points per enabled execution mode, focusing measurements around curve transitions and knees.

For detailed command reference, see:


Output and Visualization

Mess generates bandwidth and latency measurements for each configuration. Use the Plotter utilities to:

  1. Generate bandwidth-latency curve plots: Visualize the bandwidth-latency relationship
  2. Compare systems: Overlay bandwidth-latency curves from different machines
  3. Profile applications: Map your workload onto the bandwidth-latency curves using Mess Profiler

Resources

  • Mess Website - Project website with tutorials, publications, and updates
  • Mess Paper - Esmaili-Dokht, P., Sgherzi, F., Girelli, V. S., Boixaderas, I., Carmin, M., Monemi, A., Armejach, A., Mercadal, E., Llort, G., Radojković, P., Moreto, M., Giménez, J., Martorell, X., Ayguadé, E., Labarta, J., Confalonieri, E., Dubey, R., & Adlard, J. (2024). A mess of memory system benchmarking, simulation and application profiling. In Proceedings of the 57th IEEE/ACM International Symposium on Microarchitecture (MICRO) (pp. 136-152). IEEE.

See Also

Clone this wiki locally