Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 3 additions & 3 deletions arch/x86/events/amd/uncore.c
Original file line number Diff line number Diff line change
Expand Up @@ -639,7 +639,7 @@ void amd_uncore_df_ctx_scan(struct amd_uncore *uncore, unsigned int cpu)
info.split.aux_data = 0;
info.split.num_pmcs = NUM_COUNTERS_NB;
info.split.gid = 0;
info.split.cid = topology_logical_package_id(cpu);
info.split.cid = topology_amd_node_id(cpu);

if (pmu_version >= 2) {
ebx.full = cpuid_ebx(EXT_PERFMON_DEBUG_FEATURES);
Expand Down Expand Up @@ -899,8 +899,8 @@ void amd_uncore_umc_ctx_scan(struct amd_uncore *uncore, unsigned int cpu)
cpuid(EXT_PERFMON_DEBUG_FEATURES, &eax, &ebx.full, &ecx, &edx);
info.split.aux_data = ecx; /* stash active mask */
info.split.num_pmcs = ebx.split.num_umc_pmc;
info.split.gid = topology_logical_package_id(cpu);
info.split.cid = topology_logical_package_id(cpu);
info.split.gid = topology_amd_node_id(cpu);
info.split.cid = topology_amd_node_id(cpu);
*per_cpu_ptr(uncore->info, cpu) = info;
}

Expand Down
7 changes: 5 additions & 2 deletions tools/arch/x86/include/asm/amd/ibs.h
Original file line number Diff line number Diff line change
Expand Up @@ -67,7 +67,8 @@ union ibs_op_ctl {
opmaxcnt_ext:7, /* 20-26: upper 7 bits of periodic op maximum count */
reserved0:5, /* 27-31: reserved */
opcurcnt:27, /* 32-58: periodic op counter current count */
reserved1:5; /* 59-63: reserved */
ldlat_thrsh:4, /* 59-62: Load Latency threshold */
ldlat_en:1; /* 63: Load Latency enabled */
};
};

Expand Down Expand Up @@ -98,7 +99,9 @@ union ibs_op_data2 {
rmt_node:1, /* 4: destination node */
cache_hit_st:1, /* 5: cache hit state */
data_src_hi:2, /* 6-7: data source high */
reserved1:56; /* 8-63: reserved */
strm_st:1, /* 8: streaming store */
rmt_socket:1, /* 9: remote socket */
reserved1:54; /* 10-63: reserved */
};
};

Expand Down
237 changes: 237 additions & 0 deletions tools/perf/Documentation/perf-amd-ibs.txt
Original file line number Diff line number Diff line change
@@ -0,0 +1,237 @@
perf-amd-ibs(1)
===============

NAME
----
perf-amd-ibs - Support for AMD Instruction-Based Sampling (IBS) with perf tool

SYNOPSIS
--------
[verse]
'perf record' -e ibs_op//
'perf record' -e ibs_fetch//

DESCRIPTION
-----------

Instruction-Based Sampling (IBS) provides precise Instruction Pointer (IP)
profiling support on AMD platforms. IBS has two independent components: IBS
Op and IBS Fetch. IBS Op sampling provides information about instruction
execution (micro-op execution to be precise) with details like d-cache
hit/miss, d-TLB hit/miss, cache miss latency, load/store data source, branch
behavior etc. IBS Fetch sampling provides information about instruction fetch
with details like i-cache hit/miss, i-TLB hit/miss, fetch latency etc. IBS is
per-smt-thread i.e. each SMT hardware thread contains standalone IBS units.

Both, IBS Op and IBS Fetch, are exposed as PMUs by Linux and can be exploited
using the Linux perf utility. The following files will be created at boot time
if IBS is supported by the hardware and kernel.

/sys/bus/event_source/devices/ibs_op/
/sys/bus/event_source/devices/ibs_fetch/

IBS Op PMU supports two events: cycles and micro ops. IBS Fetch PMU supports
one event: fetch ops.

IBS PMUs do not have user/kernel filtering capability and thus it requires
CAP_SYS_ADMIN or CAP_PERFMON privilege.

IBS VS. REGULAR CORE PMU
------------------------

IBS gives samples with precise IP, i.e. the IP recorded with IBS sample has
no skid. Whereas the IP recorded by regular core PMU will have some skid
(sample was generated at IP X but perf would record it at IP X+n). Hence,
regular core PMU might not help for profiling with instruction level
precision. Further, IBS provides additional information about the sample in
question. On the other hand, regular core PMU has it's own advantages like
plethora of events, counting mode (less interference), up to 6 parallel
counters, event grouping support, filtering capabilities etc.

Three regular core PMU events are internally forwarded to IBS Op PMU when
precise_ip attribute is set:

-e cpu-cycles:p becomes -e ibs_op//
-e r076:p becomes -e ibs_op//
-e r0C1:p becomes -e ibs_op/cnt_ctl=1/

EXAMPLES
--------

IBS Op PMU
~~~~~~~~~~

System-wide profile, cycles event, sampling period: 100000

# perf record -e ibs_op// -c 100000 -a

Per-cpu profile (cpu10), cycles event, sampling period: 100000

# perf record -e ibs_op// -c 100000 -C 10

Userspace only, per-cpu profile (cpu10), cycles event, sampling period: 100000

Zen6 onward (See NOTES):
# perf record -e ibs_op//u -c 100000 -C 10

Until Zen5:
# perf record -e ibs_op/swfilt=1/u -c 100000 -C 10

Per-cpu profile (cpu10), cycles event, sampling freq: 1000

# perf record -e ibs_op// -F 1000 -C 10

System-wide profile, uOps event, sampling period: 100000

# perf record -e ibs_op/cnt_ctl=1/ -c 100000 -a

Same command, but also capture IBS register raw dump along with perf sample:

# perf record -e ibs_op/cnt_ctl=1/ -c 100000 -a --raw-samples

System-wide profile, uOps event, sampling period: 100000, L3MissOnly (Zen4 onward)

# perf record -e ibs_op/cnt_ctl=1,l3missonly=1/ -c 100000 -a

System-wide profile, cycles event, sampling period: 100000, LdLat filtering (Zen5
onward)

# perf record -e ibs_op/ldlat=128/ -c 100000 -a

Supported load latency threshold values are 128 to 2048 (both inclusive).
Latency value which is a multiple of 128 incurs a little less profiling
overhead compared to other values.

System-wide profile, cycles event, sampling period: 100000, streaming store
filter (Zen6 onward)

# perf record -e ibs_op/strmst=1/ -c 100000 -a

Per process(upstream v6.2 onward), uOps event, sampling period: 100000

# perf record -e ibs_op/cnt_ctl=1/ -c 100000 -p 1234

Per process(upstream v6.2 onward), uOps event, sampling period: 100000

# perf record -e ibs_op/cnt_ctl=1/ -c 100000 -- ls

To analyse recorded profile in aggregate mode

# perf report
/* Select a line and press 'a' to drill down at instruction level. */

To go over each sample

# perf script

Raw dump of IBS registers when profiled with --raw-samples

# perf report -D
/* Look for PERF_RECORD_SAMPLE */

Example register raw dump:

ibs_op_ctl: 000002c30006186a MaxCnt 100000 L3MissOnly 0 En 1
Val 1 CntCtl 0=cycles CurCnt 707
IbsOpRip: ffffffff8204aea7
ibs_op_data: 0000010002550001 CompToRetCtr 1 TagToRetCtr 597
BrnRet 0 RipInvalid 0 BrnFuse 0 Microcode 1
ibs_op_data2: 0000000000000013 RmtNode 1 DataSrc 3=DRAM
ibs_op_data3: 0000000031960092 LdOp 0 StOp 1 DcL1TlbMiss 0
DcL2TlbMiss 0 DcL1TlbHit2M 1 DcL1TlbHit1G 0 DcL2TlbHit2M 0
DcMiss 1 DcMisAcc 0 DcWcMemAcc 0 DcUcMemAcc 0 DcLockedOp 0
DcMissNoMabAlloc 0 DcLinAddrValid 1 DcPhyAddrValid 1
DcL2TlbHit1G 0 L2Miss 1 SwPf 0 OpMemWidth 32 bytes
OpDcMissOpenMemReqs 12 DcMissLat 0 TlbRefillLat 0
IbsDCLinAd: ff110008a5398920
IbsDCPhysAd: 00000008a5398920

IBS applied in a real world usecase

~90% regression was observed in tbench with specific scheduler hint
which was counter intuitive. IBS profile of good and bad run captured
using perf helped in identifying exact cause of the problem:

https://lore.kernel.org/r/20220921063638.2489-1-kprateek.nayak@amd.com

IBS Fetch PMU
~~~~~~~~~~~~~

Similar commands can be used with Fetch PMU as well.

System-wide profile, fetch ops event, sampling period: 100000

# perf record -e ibs_fetch// -c 100000 -a

Userspace only, system-wide profile, fetch ops event, sampling period: 100000

Zen6 onward (See NOTES):
# perf record -e ibs_fetch//u -c 100000 -a

Until Zen5:
# perf record -e ibs_fetch/swfilt=1/u -c 100000 -a

System-wide profile, fetch ops event, sampling period: 100000, Random enable

# perf record -e ibs_fetch/rand_en=1/ -c 100000 -a

Random enable adds small degree of variability to sample period. This
helps in cases like long running loops where PMU is tagging the same
instruction over and over because of fixed sample period.

System-wide profile, fetch ops event, sampling period: 10000, fetch latency
filter (Zen6 onward)

# perf record -e ibs_fetch/fetchlat=128/ -c 10000 -a

Supported fetch latency threshold values are 128 to 1920 (both inclusive).
Latency value which is a multiple of 128 incurs a little less profiling
overhead compared to other values.

etc.

PERF MEM AND PERF C2C
---------------------

perf mem is a memory access profiler tool and perf c2c is a shared data
cacheline analyser tool. Both of them internally uses IBS Op PMU on AMD.
Below is a simple example of the perf mem tool.

# perf mem record -c 100000 -- make
# perf mem report

A normal perf mem report output will provide detailed memory access profile.
However, it can also be aggregated based on output fields. For example:

# perf mem report -F mem,sample,snoop
Samples: 3M of event 'ibs_op//', Event count (approx.): 23524876
Memory access Samples Snoop
N/A 1903343 N/A
L1 hit 1056754 N/A
L2 hit 75231 N/A
L3 hit 9496 HitM
L3 hit 2270 N/A
RAM hit 8710 N/A
Remote node, same socket RAM hit 3241 N/A
Remote core, same node Any cache hit 1572 HitM
Remote core, same node Any cache hit 514 N/A
Remote node, same socket Any cache hit 1216 HitM
Remote node, same socket Any cache hit 350 N/A
Uncached hit 18 N/A

Please refer to their man page for more detail.

NOTES
-----
Hardware privilege filtering uses bit 63 to distinguish between kernel
and userspace addresses. Hardware privilege filtering is not supported
on 32-bit systems. Also, the bit 63 convention is not universal and can
fail in specific environments, such as, using 64-bit host IBS to profile
a 32-bit guest, using 64-bit host IBS to profile non-Linux 64-bit guests
that do not adhere to the bit 63 privilege standard etc.

SEE ALSO
--------

linkperf:perf-record[1], linkperf:perf-script[1], linkperf:perf-report[1],
linkperf:perf-mem[1], linkperf:perf-c2c[1]
3 changes: 2 additions & 1 deletion tools/perf/Documentation/perf.txt
Original file line number Diff line number Diff line change
Expand Up @@ -77,7 +77,8 @@ linkperf:perf-stat[1], linkperf:perf-top[1],
linkperf:perf-record[1], linkperf:perf-report[1],
linkperf:perf-list[1]

linkperf:perf-annotate[1],linkperf:perf-archive[1],linkperf:perf-arm-spe[1],
linkperf:perf-amd-ibs[1], linkperf:perf-annotate[1],
linkperf:perf-archive[1], linkperf:perf-arm-spe[1],
linkperf:perf-bench[1], linkperf:perf-buildid-cache[1],
linkperf:perf-buildid-list[1], linkperf:perf-c2c[1],
linkperf:perf-config[1], linkperf:perf-data[1], linkperf:perf-diff[1],
Expand Down
34 changes: 29 additions & 5 deletions tools/perf/pmu-events/arch/x86/amdzen6/floating-point.json
Original file line number Diff line number Diff line change
Expand Up @@ -212,7 +212,7 @@
{
"EventName": "fp_ops_ret_by_type.scalar_logical",
"EventCode": "0x0a",
"BriefDescription": "Retired scalar floating-point move uops.",
"BriefDescription": "Retired scalar floating-point logical uops.",
"UMask": "0x0d"
},
{
Expand Down Expand Up @@ -665,6 +665,12 @@
"BriefDescription": "Retired 256-bit packed floating-point shuffle uops (may include instructions not necessarily thought of as including shuffles e.g. horizontal add, dot product, and certain MOV instructions).",
"UMask": "0xb0"
},
{
"EventName": "fp_pack_ops_ret.fp256_bfloat",
"EventCode": "0x0c",
"BriefDescription": "Retired 256-bit packed floating-point bfloat uops.",
"UMask": "0xc0"
},
{
"EventName": "fp_pack_ops_ret.fp256_logical",
"EventCode": "0x0c",
Expand Down Expand Up @@ -758,7 +764,7 @@
{
"EventName": "fp_pack_int_ops_ret.int128_vnni",
"EventCode": "0x0d",
"BriefDescription": "Retired 128-bit packed integer VNNI ops.",
"BriefDescription": "Retired 128-bit packed integer VNNI uops.",
"UMask": "0x0c"
},
{
Expand Down Expand Up @@ -803,12 +809,30 @@
"BriefDescription": "Retired 256-bit packed integer multiply-accumulate uops.",
"UMask": "0x40"
},
{
"EventName": "fp_pack_int_ops_ret.int256_aes",
"EventCode": "0x0d",
"BriefDescription": "Retired 256-bit packed integer AES uops.",
"UMask": "0x50"
},
{
"EventName": "fp_pack_int_ops_ret.int256_sha",
"EventCode": "0x0d",
"BriefDescription": "Retired 256-bit packed integer SHA uops.",
"UMask": "0x60"
},
{
"EventName": "fp_pack_int_ops_ret.int256_cmp",
"EventCode": "0x0d",
"BriefDescription": "Retired 256-bit packed integer compare uops.",
"UMask": "0x70"
},
{
"EventName": "fp_pack_int_ops_ret.int256_cvt",
"EventCode": "0x0d",
"BriefDescription": "Retired 256-bit packed integer convert or pack uops.",
"UMask": "0x80"
},
{
"EventName": "fp_pack_int_ops_ret.int256_shift",
"EventCode": "0x0d",
Expand Down Expand Up @@ -1083,19 +1107,19 @@
"EventName": "fp_nsq_read_stalls.fp_prf",
"EventCode": "0x13",
"BriefDescription": "Cycles when reads of the NSQ and writes to the floating-point or SIMD schedulers are stalled due to insufficient free physical register file (FP-PRF) entries.",
"UMask": "0x0e"
"UMask": "0x02"
},
{
"EventName": "fp_nsq_read_stalls.k_prf",
"EventCode": "0x13",
"BriefDescription": "Cycles when reads of the NSQ and writes to the floating-point or SIMD schedulers are stalled due to insufficient free mask physical register file (K-PRF) entries.",
"UMask": "0x0e"
"UMask": "0x04"
},
{
"EventName": "fp_nsq_read_stalls.fp_sq",
"EventCode": "0x13",
"BriefDescription": "Cycles when reads of the NSQ and writes to the floating-point or SIMD schedulers are stalled due to insufficient free scheduler entries.",
"UMask": "0x0e"
"UMask": "0x08"
},
{
"EventName": "fp_nsq_read_stalls.all",
Expand Down
Loading