Libcpucycles is a public-domain microlibrary for counting CPU cycles
cpucycles.cr.yp.to
cpucycles.cr.yp.to
https://cpucycles.cr.yp.to/libcpucycles-20230105/cpucycles/d...
memset(&attr,0,sizeof attr);
attr.type = PERF_TYPE_HARDWARE;
attr.size = sizeof(struct perf_event_attr);
attr.config = PERF_COUNT_HW_CPU_CYCLES;
attr.disabled = 1;
attr.exclude_kernel = 1;
attr.exclude_hv = 1;
fddev = syscall(__NR_perf_event_open,&attr,0,-1,-1,0);
It needs to read time enabled / time running and scale returned values: attr.read_format = PERF_FORMAT_TOTAL_TIME_ENABLED | PERF_FORMAT_TOTAL_TIME_RUNNING;I imagine you'd see some weird effects if the CPU changed its frequency during a benchmark, or if the process moved between cores since (as I understand it) the RAM would get faster relative to the CPU when the CPU is clocked lower. So some operations would take fewer cycles when the CPU is running slower.
Or am I misunderstanding what this is measuring?
- default-perfevent: the kernel will keep track of counter values on core moves
- amd64-pmc: Accesses a cycle counter through RDPMC: best to pin the benchmark binary to a specific core to avoid measurement issues when the task moves between cores: "taskset -c 1 <benchmark-binary>"
- amd64-tsc, amd64-tscasm: RDTSC is off-core counter - not influenced by cross-core moves ("On current CPUs, this is an off-core clock rather than a cycle counter, but it is typically a very fast off-core clock, making it adequate for seeing cycle counts if overclocking and underclocking are disabled.")
Yes, you must disable CPU frequency scaling in your BIOS if you're doing this kind of work (i.e. building cryptographic primitives that don't leak information via timing).
Generally speaking, just because modern x86 CPUs use numerous abstractions for high-level instruction set features does not mean they don't use cycles. (Clock) cycles are an inherent feature of synchronous logic, and all CPUs (modern as well as ancient) use synchronous logic. Yes, there might be some esoteric outliers, but 99.9% is synchronous.
The difference to ancient CPUs is that modern CPUs don't necessarily retire (essentially, "execute") an instruction in every clock cycle. In any given clock cycle, a CPU might retire none, just one or even multiple instructions.
A typical application (where clock cycles are important) is looking at the number of retired instructions per cycle (IPC), which can give a rough overview of whether a program is frontend-bound or backend-bound. You can try this yourself using "perf stat" (part of the Linux perf tools): Try running different applications using perf stat and look for "insn per cycle" in the report. You will generally find that interpreted programs (Python, Node etc.) will have a mediocre IPC of well below 1 (due to poor cache utilization and branch mispredictions) while compiled, optimized programs might reach an IPC of 2. A high IPC value is not necessarily better though, it's complicated^TM.
[0] https://bitbucket.org/icl/papi/wiki/PAPI-Overview.md#markdow...