Repository navigation
Instruction Sampling
This page documents instruction sampling for latency measurement, an alternative to the default pointer-chase method.
Mess supports two methods for measuring memory latency under load:
| Method | How It Works | Cores Used for Traffic |
|---|---|---|
| Pointer chase (default) | A dedicated core runs a pointer-chase loop to measure latency | N - 1 cores |
Instruction sampling (--inst-lat) |
Hardware samples latency directly from traffic generator loads | All N cores |
In the default mode, one core is reserved for a pointer-chase linked-list traversal while the remaining cores generate memory traffic. The pointer-chase latency serves as a proxy for the load-to-use latency experienced by actual memory operations under that traffic pressure.
This approach is well-established and works across all architectures, but it has a limitation: the latency is measured by a separate operation (pointer chase), not by the traffic generator instructions themselves. Additionally, one core is consumed by the measurement and does not contribute to traffic generation.
With --inst-lat, Mess uses hardware instruction sampling to measure latency directly from the traffic generator loads. The CPU's sampling hardware tags load instructions with their observed load-to-use latency in cycles. Mess collects these samples, filters them to retain only main-memory accesses, and computes latency statistics.
Because latency is measured from the traffic generators themselves, all cores run traffic generation — no core is reserved for pointer chasing.
When --inst-lat is specified, Mess creates an instruction sampler backend appropriate for the detected hardware. Auto-detection currently tries Intel PEBS first and ARM SPE second. The sampler is configured with:
- Latency threshold: Minimum load latency to record (default: 50 cycles). This filters out cache hits early at the hardware level.
- Sample rate/period: Controls how frequently loads are sampled.
Sampling runs asynchronously alongside the traffic generators:
- Traffic generators start on all selected cores.
- During bandwidth stabilization, after a few bandwidth samples have been collected, the sampler launches
perfin the background, targeting the traffic generator cores and PIDs. -
perfrecords sampled load events with their latency and data source metadata. - When bandwidth fully stabilizes, sampling stops and the collected data is processed.
Not every sampled load is relevant. Mess filters the raw samples to isolate main-memory load-to-use latency:
- RAM-only: Only loads that hit main memory (DRAM) are kept. Cache hits (L1/L2/L3) are discarded, since we are measuring memory latency.
- Valid TLB translations: Only loads where the TLB translation was served from L1 or L2 TLB are kept. This excludes accesses inflated by TLB miss penalties (page walks), matching the same goal as the huge pages strategy used in pointer-chase mode.
- Non-zero latency: Samples with a zero weight field are discarded.
After filtering, Mess computes a statistical breakdown from the remaining samples:
| Statistic | Description |
|---|---|
| Samples | Number of valid RAM-hit samples collected |
| Mean | Average latency (cycles and ns) |
| IQR mean | Mean after keeping the interquartile range |
| Min | Minimum observed latency |
| Median (p50) | 50th percentile latency |
| p90 | 90th percentile latency |
| p95 | 95th percentile latency |
| p99 | 99th percentile latency |
| p99.9 | 99.9th percentile latency |
| Max | Maximum observed latency |
Cycle counts are converted to nanoseconds using the CPU frequency selected for the run. The bandwidth-latency curve uses the IQR mean when it is available, otherwise the full mean.
Processor Event-Based Sampling (PEBS) is Intel's hardware mechanism for precise instruction sampling. PEBS captures architectural state at the point of instruction retirement, including the load-to-use latency for memory load operations.
Mess uses perf mem record with the load-latency (ldlat) facility:
perf mem record -t load --ldlat <threshold> -F <freq> -C <cores>
-
-t load: Record only load operations. -
--ldlat <threshold>: Hardware-level filter — only loads with latency >= threshold cycles are recorded. Default: 50 cycles. -
-F <freq>: Sampling frequency. Default: 997 Hz. -
-C <cores>: Target only the traffic generator cores.
Different Intel generations expose load-latency events under different names:
| Intel Generation | PMU Device | Event |
|---|---|---|
| Sapphire Rapids and newer | cpu_core |
mem-loads-aux |
| Older generations | cpu |
mem-loads |
Mess automatically discovers the correct PMU and event by scanning /sys/bus/event_source/devices/.
After collection, Mess runs perf script to extract per-sample fields:
-
CPU/core selection: PEBS collection is constrained to the selected traffic generator cores with
perf mem record -C. - Weight: The load-to-use latency in cycles.
-
Data source (
data_src): A bitmask encoding the memory hierarchy level and TLB status of each access.
The data_src field is decoded to check:
- Memory level bits: Whether the load was served from RAM (not L1/L2/L3 cache).
- TLB bits: Whether the address translation was served from L1/L2 TLB (no page walk penalty).
Only samples passing both filters contribute to the latency statistics. If perf script does not expose usable samples, Mess falls back to perf report --mem-mode and keeps RAM-hit samples reported there.
Statistical Profiling Extension (SPE) is ARM's equivalent to Intel PEBS. SPE is an architectural feature (ARMv8.2+) that provides non-invasive statistical sampling of instructions, including memory operations with their observed latency.
Mess detects SPE availability by scanning /sys/bus/event_source/devices/ for an arm_spe device, then launches perf with the SPE PMU:
- Sample period: Default 4096 (one sample every ~4096 matching operations).
- Minimum latency threshold: Default 50 cycles.
-
Core/PID limit: SPE sampling is capped at 8 cores by default. If core-based collection is not available, or if
MESS_SPE_PREFER_PIDS=1is set, Mess can instead attach to up to 8 traffic generator PIDs. This cap keeps the number of perf events and file descriptors bounded while still sampling representative workers.
| Aspect | Intel PEBS | ARM SPE |
|---|---|---|
| Availability | Intel Core/Xeon (Nehalem+) | ARMv8.2+ with SPE |
| Sampling trigger | Frequency-based (Hz) | Period-based (every N ops) |
| Latency field |
weight via ldlat
|
Latency field in SPE record |
| Core limit | None | 8 cores by default, or 8 PIDs when PID attach is selected |
| Event discovery | /sys/bus/event_source/devices/cpu[_core] |
/sys/bus/event_source/devices/arm_spe* |
When --inst-lat is specified, Mess automatically selects the appropriate backend:
- Check for Intel PEBS support (try PEBS PMU discovery).
- If not available, check for ARM SPE support.
- If neither is available, Mess warns that no sampler backend is available.
This is handled by the InstructionSamplerFactory with a SamplerBackend::AUTO policy.
Note: Instruction sampling is currently supported on Intel (PEBS) and ARM (SPE) only. Power and RISC-V architectures must use the default pointer-chase approach.
Enable instruction sampling with the --inst-lat flag:
./build/bin/mess --inst-latThis replaces pointer-chase latency measurement with hardware instruction sampling. All other Mess options (ratios, pauses, output, etc.) work as usual.
Instruction sampling can also be combined with adaptive pause discovery:
./build/bin/mess --inst-lat --tier=standard
./build/bin/mess --inst-lat --ratio=100 --point-count=75When running with --inst-lat, Mess creates a sampler/ directory alongside the usual bw/ and lat/ directories:
measuring/
├── bw/ # Raw bandwidth measurements
├── lat/ # Raw latency measurements (from sampler stats)
└── sampler/ # Raw sampler data
├── sampler_100_0.csv # All valid latency samples for Ratio 100%, Pause 0
├── sampler_100_10.csv # All valid latency samples for Ratio 100%, Pause 10
└── ...
Each sampler_<ratio>_<pause>.csv file contains all the valid (filtered, RAM-hit) latency values in cycles that were measured for that configuration. These raw samples can be used for custom analysis and plotting beyond the summary statistics.
The regular lat/lat_<ratio>_<pause>.txt files contain the summarized statistics used by Mess:
samples-
mean_ns,iqr_mean_ns,min_ns,median_ns,p90_ns,p95_ns,p99_ns,p99_9_ns,max_ns - matching cycle fields such as
mean_cyclesandp99_cycles
iqr_mean_ns is populated when the backend computes it. The current PEBS path computes it from the interquartile sample range; the curve value uses it when it is greater than zero, otherwise it falls back to mean_ns.
Each measurement point in the bandwidth-latency curve produces a distribution of latency samples rather than only one value. Mess uses iqr_mean_ns as the representative curve value when it is available, and falls back to mean_ns otherwise.
Exploring how to best leverage the richer statistical breakdown (percentiles, distribution shape) for curve construction and visualization is an area of ongoing work. The latencyPlot.py utility in the utils/ directory provides initial tools for this analysis — it can generate per-ratio latency CDFs, histograms, tail latency plots, and IQR-filtered curves from the sampler/ data. See utils/README.md for usage details.
For manual tuning or understanding the exact perf commands and filtering logic, the relevant implementation files are:
| File | Contents |
|---|---|
src/measurement/instruction_samplers/IntelPebsSampler.cpp |
PEBS perf mem record command construction, perf script parsing, data_src filtering logic |
src/measurement/instruction_samplers/ArmSpeSampler.cpp |
SPE perf command construction, core subsampling, SPE record parsing |
src/measurement/InstructionSamplerFactory.cpp |
Auto-detection logic and default parameter values |
include/measurement/instruction_samplers/IntelPebsSampler.h |
PEBS configurable parameters (threshold, frequency) |
include/measurement/instruction_samplers/ArmSpeSampler.h |
SPE configurable parameters (period, threshold, core cap) |
- Mess Benchmark - Measurement methodology overview
- Adaptive Curve-Guided Pause Discovery - Automatic pause selection
- Huge Memory Pages - TLB miss mitigation (relevant to sample filtering)
- Load-Store vs Read-Write - Issued vs measured traffic
- Traffic Generator - Traffic generation internals
- Understand output - Output file formats