Trying to figure out how this would be of use to anyone else, in practice. You need a hacked up runtime and a privileged process to use the PMU. Compared to just using perf, it seems like a hassle.
I'm happy to rebuild the binary when I'm stuck on a performance problem. Obviously, the best case is having all the instrumentation on the production system that is experiencing problems, of course. I've run into problems that only manifest themselves after running for several hours, so the debugging cycle is slow when you aren't collecting what you need to debug the issue. In other cases, I've had bugs where a test case can reproduce them immediately, but the profiler still provides more targeted guesses than randomly changing something to see if it improves the outcome (even if the edit/test cycle is fast enough to support this). That's the case where this sounds helpful.
Yeah, but you know what also works? Running perf and running that through normal pprof...
The examples seem to run fine and report counters without any special privileges. Building the toolchain and then building your own code with the toolchain is trivial too, I was up and running in 10 minutes, much of which was waiting for the builds.
That suggests your system has a liberal setting of /proc/sys/kernel/perf_event_paranoid or your process is privileged, running as root or with either CAP_SYS_ADMIN or CAP_PERFMON. Generally unprivileged processes cannot access the PMU, because the PMU can be used to take over the system.