Skip to content

Instruction Sampling

Victor Xirau Guardans edited this page Sep 29, 2026 · 2 revisions

This page documents instruction sampling for latency measurement, an alternative to the default pointer-chase method.


Overview

Mess supports two methods for measuring memory latency under load:

Method How It Works Cores Used for Traffic
Pointer chase (default) A dedicated core runs a pointer-chase loop to measure latency N - 1 cores
Instruction sampling (--inst-lat) Hardware samples latency directly from traffic generator loads All N cores

The Pointer-Chase Proxy

In the default mode, one core is reserved for a pointer-chase linked-list traversal while the remaining cores generate memory traffic. The pointer-chase latency serves as a proxy for the load-to-use latency experienced by actual memory operations under that traffic pressure.

This approach is well-established and works across all architectures, but it has a limitation: the latency is measured by a separate operation (pointer chase), not by the traffic generator instructions themselves. Additionally, one core is consumed by the measurement and does not contribute to traffic generation.

Instruction Sampling as an Alternative

With --inst-lat, Mess uses hardware instruction sampling to measure latency directly from the traffic generator loads. The CPU's sampling hardware tags load instructions with their observed load-to-use latency in cycles. Mess collects these samples, filters them to retain only main-memory accesses, and computes latency statistics.

Because latency is measured from the traffic generators themselves, all cores run traffic generation — no core is reserved for pointer chasing.


How It Works

1. Sampling Setup

When --inst-lat is specified, Mess creates an instruction sampler backend appropriate for the detected hardware. Auto-detection currently tries Intel PEBS first and ARM SPE second. The sampler is configured with:

  • Latency threshold: Minimum load latency to record (default: 50 cycles). This filters out cache hits early at the hardware level.
  • Sample rate/period: Controls how frequently loads are sampled.

2. Asynchronous Collection

Sampling runs asynchronously alongside the traffic generators:

  1. Traffic generators start on all selected cores.
  2. During bandwidth stabilization, after a few bandwidth samples have been collected, the sampler launches perf in the background, targeting the traffic generator cores and PIDs.
  3. perf records sampled load events with their latency and data source metadata.
  4. When bandwidth fully stabilizes, sampling stops and the collected data is processed.

3. Filtering

Not every sampled load is relevant. Mess filters the raw samples to isolate main-memory load-to-use latency:

  • RAM-only: Only loads that hit main memory (DRAM) are kept. Cache hits (L1/L2/L3) are discarded, since we are measuring memory latency.
  • Valid TLB translations: Only loads where the TLB translation was served from L1 or L2 TLB are kept. This excludes accesses inflated by TLB miss penalties (page walks), matching the same goal as the huge pages strategy used in pointer-chase mode.
  • Non-zero latency: Samples with a zero weight field are discarded.

4. Statistics

After filtering, Mess computes a statistical breakdown from the remaining samples:

Statistic Description
Samples Number of valid RAM-hit samples collected
Mean Average latency (cycles and ns)
IQR mean Mean after keeping the interquartile range
Min Minimum observed latency
Median (p50) 50th percentile latency
p90 90th percentile latency
p95 95th percentile latency
p99 99th percentile latency
p99.9 99.9th percentile latency
Max Maximum observed latency

Cycle counts are converted to nanoseconds using the CPU frequency selected for the run. The bandwidth-latency curve uses the IQR mean when it is available, otherwise the full mean.


Intel PEBS

Processor Event-Based Sampling (PEBS) is Intel's hardware mechanism for precise instruction sampling. PEBS captures architectural state at the point of instruction retirement, including the load-to-use latency for memory load operations.

How Mess Uses PEBS

Mess uses perf mem record with the load-latency (ldlat) facility:

perf mem record -t load --ldlat <threshold> -F <freq> -C <cores>
  • -t load: Record only load operations.
  • --ldlat <threshold>: Hardware-level filter — only loads with latency >= threshold cycles are recorded. Default: 50 cycles.
  • -F <freq>: Sampling frequency. Default: 997 Hz.
  • -C <cores>: Target only the traffic generator cores.

PMU Event Discovery

Different Intel generations expose load-latency events under different names:

Intel Generation PMU Device Event
Sapphire Rapids and newer cpu_core mem-loads-aux
Older generations cpu mem-loads

Mess automatically discovers the correct PMU and event by scanning /sys/bus/event_source/devices/.

Data Source Decoding

After collection, Mess runs perf script to extract per-sample fields:

  • CPU/core selection: PEBS collection is constrained to the selected traffic generator cores with perf mem record -C.
  • Weight: The load-to-use latency in cycles.
  • Data source (data_src): A bitmask encoding the memory hierarchy level and TLB status of each access.

The data_src field is decoded to check:

  • Memory level bits: Whether the load was served from RAM (not L1/L2/L3 cache).
  • TLB bits: Whether the address translation was served from L1/L2 TLB (no page walk penalty).

Only samples passing both filters contribute to the latency statistics. If perf script does not expose usable samples, Mess falls back to perf report --mem-mode and keeps RAM-hit samples reported there.


ARM SPE

Statistical Profiling Extension (SPE) is ARM's equivalent to Intel PEBS. SPE is an architectural feature (ARMv8.2+) that provides non-invasive statistical sampling of instructions, including memory operations with their observed latency.

How Mess Uses SPE

Mess detects SPE availability by scanning /sys/bus/event_source/devices/ for an arm_spe device, then launches perf with the SPE PMU:

  • Sample period: Default 4096 (one sample every ~4096 matching operations).
  • Minimum latency threshold: Default 50 cycles.
  • Core/PID limit: SPE sampling is capped at 8 cores by default. If core-based collection is not available, or if MESS_SPE_PREFER_PIDS=1 is set, Mess can instead attach to up to 8 traffic generator PIDs. This cap keeps the number of perf events and file descriptors bounded while still sampling representative workers.

SPE vs PEBS

Aspect Intel PEBS ARM SPE
Availability Intel Core/Xeon (Nehalem+) ARMv8.2+ with SPE
Sampling trigger Frequency-based (Hz) Period-based (every N ops)
Latency field weight via ldlat Latency field in SPE record
Core limit None 8 cores by default, or 8 PIDs when PID attach is selected
Event discovery /sys/bus/event_source/devices/cpu[_core] /sys/bus/event_source/devices/arm_spe*

Backend Auto-Detection

When --inst-lat is specified, Mess automatically selects the appropriate backend:

  1. Check for Intel PEBS support (try PEBS PMU discovery).
  2. If not available, check for ARM SPE support.
  3. If neither is available, Mess warns that no sampler backend is available.

This is handled by the InstructionSamplerFactory with a SamplerBackend::AUTO policy.

Note: Instruction sampling is currently supported on Intel (PEBS) and ARM (SPE) only. Power and RISC-V architectures must use the default pointer-chase approach.


Usage

Enable instruction sampling with the --inst-lat flag:

./build/bin/mess --inst-lat

This replaces pointer-chase latency measurement with hardware instruction sampling. All other Mess options (ratios, pauses, output, etc.) work as usual.

Instruction sampling can also be combined with adaptive pause discovery:

./build/bin/mess --inst-lat --tier=standard
./build/bin/mess --inst-lat --ratio=100 --point-count=75

Output

When running with --inst-lat, Mess creates a sampler/ directory alongside the usual bw/ and lat/ directories:

measuring/
├── bw/                        # Raw bandwidth measurements
├── lat/                       # Raw latency measurements (from sampler stats)
└── sampler/                   # Raw sampler data
    ├── sampler_100_0.csv      # All valid latency samples for Ratio 100%, Pause 0
    ├── sampler_100_10.csv     # All valid latency samples for Ratio 100%, Pause 10
    └── ...

Each sampler_<ratio>_<pause>.csv file contains all the valid (filtered, RAM-hit) latency values in cycles that were measured for that configuration. These raw samples can be used for custom analysis and plotting beyond the summary statistics.

The regular lat/lat_<ratio>_<pause>.txt files contain the summarized statistics used by Mess:

  • samples
  • mean_ns, iqr_mean_ns, min_ns, median_ns, p90_ns, p95_ns, p99_ns, p99_9_ns, max_ns
  • matching cycle fields such as mean_cycles and p99_cycles

iqr_mean_ns is populated when the backend computes it. The current PEBS path computes it from the interquartile sample range; the curve value uses it when it is greater than zero, otherwise it falls back to mean_ns.


Curve Construction

Each measurement point in the bandwidth-latency curve produces a distribution of latency samples rather than only one value. Mess uses iqr_mean_ns as the representative curve value when it is available, and falls back to mean_ns otherwise.

Exploring how to best leverage the richer statistical breakdown (percentiles, distribution shape) for curve construction and visualization is an area of ongoing work. The latencyPlot.py utility in the utils/ directory provides initial tools for this analysis — it can generate per-ratio latency CDFs, histograms, tail latency plots, and IQR-filtered curves from the sampler/ data. See utils/README.md for usage details.


Source Code References

For manual tuning or understanding the exact perf commands and filtering logic, the relevant implementation files are:

File Contents
src/measurement/instruction_samplers/IntelPebsSampler.cpp PEBS perf mem record command construction, perf script parsing, data_src filtering logic
src/measurement/instruction_samplers/ArmSpeSampler.cpp SPE perf command construction, core subsampling, SPE record parsing
src/measurement/InstructionSamplerFactory.cpp Auto-detection logic and default parameter values
include/measurement/instruction_samplers/IntelPebsSampler.h PEBS configurable parameters (threshold, frequency)
include/measurement/instruction_samplers/ArmSpeSampler.h SPE configurable parameters (period, threshold, core cap)

See Also

Clone this wiki locally